Multimedia data generation method, audio data generation method and electronic equipment
By responding to target trigger events in the application and generating matching multimedia data, the problem of fixed interaction patterns of existing applications is solved, and a richer user experience and personalized needs are achieved.
Patent Information
- Application Number
- CN202411978356.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-09
AI Technical Summary
The operating mode of existing applications in the entertainment and social fields is highly fixed, which is difficult to meet users' personalized needs, and there are shortcomings in terms of fun and interactivity, resulting in limited depth and breadth of user experience.
By responding to the target trigger event, the characteristic data of the target application is obtained and matching multimedia data is generated based on the data, increasing the flexibility and dynamic interaction between the application and the user.
It achieves a richer user experience, meets user personalized needs, and improves the attractiveness and user stickiness of applications in a fiercely competitive market.
Smart Images

Figure CN119961465A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a multimedia data generating method, an audio data generating method and an electronic device. Background Art
[0002] In the digital age, various applications are widely popular in the fields of entertainment and social interaction. However, the existing application operation mode has obvious limitations. At present, when users use applications for entertainment and social activities, the application's input and output feedback mechanism for users is highly fixed. In most cases, special audio and video effects can only be triggered when specific preset conditions are met. This fixed mode makes the interaction between the target application and the user lack flexibility and dynamism. This not only makes it difficult to meet the growing personalized needs of users, but also has major deficiencies in fun and interactivity, resulting in limited depth and breadth of user experience, and unable to fully tap the potential of applications in enhancing user participation and immersion, which is not conducive to applications continuing to attract users and maintain user stickiness in a highly competitive market environment. Summary of the invention
[0003] According to a first aspect of the present disclosure, there is provided a method for generating multimedia data, comprising:
[0004] In response to a target triggering event, obtaining target feature data of a target application;
[0005] Generate target multimedia data matching the target trigger event based on the target feature data;
[0006] The target trigger event is an event that can cause the operating parameters of the target application to change.
[0007] According to a second aspect of the present disclosure, there is provided a method for generating audio data, comprising:
[0008] In response to a target operation acting on a target game application, obtaining target application data of the target game application;
[0009] generating target audio data matching the target operation based on the target application data;
[0010] The target operation can trigger changes in operating parameters of the target game application.
[0011] According to a third aspect of the present disclosure, an electronic device includes a processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following:
[0012] In response to a target triggering event, obtaining target feature data of a target application;
[0013] Generate target multimedia data matching the target trigger event based on the target feature data;
[0014] The target trigger event is an event that can cause the operating parameters of the target application to change.
[0015] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which:
[0017] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0018] Figure 1 A first optional flow chart of the multimedia data generating method provided by an embodiment of the present disclosure is shown;
[0019] Figure 2 A second optional flow chart of the multimedia data generating method provided by the embodiment of the present disclosure is shown;
[0020] Figure 3 A third optional flow chart of the multimedia data generating method provided by the embodiment of the present disclosure is shown;
[0021] Figure 4 A fourth optional flow chart of the multimedia data generating method provided by an embodiment of the present disclosure is shown;
[0022] Figure 5 A fifth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown;
[0023] Figure 6 A sixth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown;
[0024] Figure 7 A seventh optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown;
[0025] Figure 8 An eighth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown;
[0026] Fig. 9 A first optional flow chart of the method for generating audio data provided by an embodiment of the present disclosure is shown;
[0027] Fig.10 A second optional flow chart of the method for generating audio data provided by an embodiment of the present disclosure is shown;
[0028] Fig.11 An optional structural diagram of a multimedia data generating device provided in an embodiment of the present disclosure is shown;
[0029] Fig.12 An optional structural diagram of an audio data generating device provided by an embodiment of the present disclosure is shown;
[0030] Fig.13 A schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0031] In order to make the purpose, features, and advantages of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0032] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0033] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.
[0034] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as those commonly understood by those skilled in the art to which this disclosure belongs. The terms used in this disclosure are only for the purpose of describing the embodiments of this disclosure and are not intended to limit this disclosure.
[0035] It should be understood that in the various embodiments of the present disclosure, the size of the serial number of each implementation process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0036] In the digital age, various applications are widely popular in the fields of entertainment and social interaction. However, the existing application operation mode has obvious limitations. At present, when users use applications for entertainment and social activities, the input and output feedback mechanism of the application for users is highly fixed. In most cases, special audio and video effects can only be triggered when specific preset conditions are met. This fixed mode makes the interaction between applications and users lack flexibility and dynamism. For example, in some social chat applications, after users send messages or perform operations, the response form is single, and the audio and video presentation is undifferentiated; in entertainment applications such as mini-games, if the established level goals or score requirements are not achieved, the unique audio-visual feedback cannot be experienced. This not only makes it difficult to meet the growing personalized needs of users, but also has major deficiencies in fun and interactivity, resulting in limited depth and breadth of user experience, and unable to fully tap the potential of applications in enhancing user participation and immersion, which is not conducive to the application continuing to attract users and maintain user stickiness in a highly competitive market environment.
[0037] In view of the defects existing in the related art, the embodiments of the present disclosure provide a multimedia data generating method, an audio data generating method and an electronic device to at least solve all or part of the above-mentioned technical problems.
[0038] Figure 1 A first optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0039] Step S101, in response to a target triggering event, obtaining target feature data of a target application.
[0040] In some embodiments, the target application may include a game type application, a conference type application, an audio-visual type application, and a live broadcast type application.
[0041] In some embodiments, the target trigger event is an event that can cause the operating parameters of the target application to change, such as an event or behavior that causes the operating state or operating environment of the target application to change.
[0042] In some embodiments, if the target application is a game-type application, after entering the game from the waiting state, the CPU usage and memory usage in the target application's operating parameters will increase; after entering the game, when the scene with fewer characters is switched to the scene with more characters and multiple people release skills, the CPU usage and memory usage will increase; after the game ends and the settlement page is entered, the CPU usage and memory usage will decrease; specifically, the target trigger event may include the start of the target application, the first target object corresponding to the user in the target application releases skills, the first target object switches game scenes, the first target object is attacked, the first target object achieves achievements, and the target application's operating state changes. Among them, the target application's operating state changes may include the end of the game, the first target object is temporarily unable to fight, etc.
[0043] Further, in some embodiments, if the target application is a game-type application, the target feature data may include the character data of the first target object, the skill data of the first target object releasing skills, the scene data (such as map data) before the first target object switches the game scene and the scene data (such as map data) after the switch, the character data of the first target object when the first target object is attacked and the character data and skill data of other characters attacking the first target object, the video data before the first target object achieves an achievement, the character data and skill data of other characters in the same screen as the first target object, and the running state change data of the target application. The running state change data may include the end of the game, the start of the game, the temporary inability of the first target object to fight, etc.
[0044] In some embodiments, if the target application is a conference type application, then entering the meeting, the remote host performing a series of operations during the meeting, the near-end computer user performing a series of operations, and exiting the meeting will cause the CPU usage, memory usage, and the target application's operating parameters to change due to the opening or closing of other applications. Specifically, the target triggering event may include: the remote host's scene switching (from camera capture to screen projection), host switching, screen projection content switching, screen projection device switching, other remote users speaking, and other remote users liking; wherein the other remote users and the remote host are in the same scene or in different scenes; the near-end user starts screen recording, ends screen recording (when recording, the recording module of the target application switches from standby mode to running mode, which will occupy the CPU and memory), the near-end user speaks, the near-end user turns on the camera, the near-end user turns off the camera, the near-end user likes, the near-end user inputs data (such as answering questions, asking questions, explaining other people's questions, etc.), the near-end user turns on the hang-up mode, the near-end user turns on the focus mode, etc. Among them, the hang-up mode includes that the near-end user temporarily leaves the screen, and the auxiliary function helps the near-end user to record and respond to other people's interactions; the focus mode includes that the user highlights the video and audio displayed by the target application.
[0045] Further, in some embodiments, if the target application is a conference-type application, the target characteristic data may include content data displayed on the application screen when receiving the target trigger event, displayed item data, discussed content data, and host information.
[0046] In some embodiments, if the target application is an audio-visual type application, then entering the video or audio playback page, playing the video or audio, pausing the video or audio, exiting the video or audio playback page, and a series of operations performed by the user will cause the CPU occupancy, memory occupancy, and the operating parameters of the target application caused by the opening or closing of other applications. Specifically, the target trigger event may include entering the video or audio playback page, playing the video or audio, pausing the video or audio, exiting the video or audio playback page, turning on the barrage, turning off the barrage, turning on comments, turning off comments, liking at least one barrage, disliking at least one barrage, taking screenshots, recording the screen, sharing video or audio, reporting at least one barrage, posting a barrage, commenting, replying to a comment, liking at least one comment, and disliking at least one comment.
[0047] Further, in some embodiments, if the target application is an audio-visual type application, the target feature data may include content data, actor data, scene data, action data, barrage data and comment data played by the application when receiving the target trigger event.
[0048] In some embodiments, if the target application is a live broadcast type application, then entering the live broadcast room, exiting the live broadcast room, a series of operations performed by the anchor, and a series of operations performed by the user will cause the operating parameters of the target application to change. Specifically, the target triggering event may include entering the live broadcast room, exiting the live broadcast room, the anchor displaying the product, changing the anchor, opening the live broadcast room, closing the live broadcast room, users posting comments, and users giving likes.
[0049] Further, in some embodiments, if the target application is a live broadcast type application, the target characteristic data may include content data, anchor data, scene data, product data and comment data played by the application when receiving the target trigger event.
[0050] Step S102: generating target multimedia data matching the target triggering event based on the target feature data.
[0051] In some embodiments, the carrier implementing the multimedia data generation method (hereinafter referred to as the carrier) can call different types of large models to generate target multimedia data according to the type of target application, the target feature data or the target trigger event. The carrier can be a computer program, electronic circuit, database, mobile application, electronic device, cloud computing platform, distributed system, artificial intelligence framework, mathematical model, automation tool and microcontroller, etc., which can realize the software or hardware of the algorithm and method flow.
[0052] Specifically, the large model may include an image generation model, a text generation model, a video generation model, and an audio generation model, or a large language model with more functions; the target application may directly call the large model to generate target multimedia data; if the target application cannot directly call the large model, the target multimedia data may be generated by calling the corresponding large model through an application program that establishes an associated relationship. The application may be an artificial intelligence type application, such as Lenovo Xiaotian or AI NOW.
[0053] In some embodiments, the target multimedia data matches the target trigger event; the large model can generate the target multimedia data in combination with the target trigger event and the target feature data.
[0054] In specific implementation, if the target application is a game-type application, if the target trigger event is to start the target application, the target multimedia data includes the name of the target application and the identifier of the target application; if the target trigger event is to release a skill, the target multimedia data includes the audio data, video data (or effect data) of the skill, evaluation data of other users and the name of the character; if the target trigger event is to switch scenes, the target multimedia data includes scene switching prompt data, scene switching video data and scene switching audio data; if the target trigger event is to pause the game or to host, the target multimedia data includes a pause identifier and data of the hosting assistant controlling the operation of the first target object; if the target trigger event is the end of the game, the target multimedia data includes settlement data, game highlight data and other user evaluation data; if the target trigger event is to receive an attack, the target multimedia data includes attack receipt prompt data, character status prompt data, audio data of skills released by other characters, video data of skills released by other characters, and data of other characters that cause the highest damage.
[0055] In a specific implementation, if the target application is a conference-type application, if the target trigger event is entering a conference, the target multimedia data includes a conference entry identifier, a camera status identifier (whether the camera status is turned on), and a microphone status identifier (whether the microphone status is turned on); if the target trigger event is a scene switch, the target multimedia data includes a scene switch prompt and an evaluation prompt of the new scene; if the target trigger event is a change of host, the target multimedia data includes a change of host prompt and an evaluation prompt of the new host; if the target trigger event is a change of display items or content, the target multimedia data includes a new display item prompt and evaluation, or new content prompt or evaluation; if the target trigger event is a remote user speaking, liking, or evaluating, the target multimedia data includes the user information of the remote user, the behavior information of the remote user, and the evaluation prompt for the remote user; if the target trigger event is the near-end user turning on the focus mode, the target multimedia data includes the focus mode logo, the brush logo, and displaying related logos according to the subsequent operations of the near-end user (such as bolding / highlighting the text of the user logo); if the target trigger event is the near-end user turning on the hang-up mode, the target multimedia data includes user behavior data learned based on user historical data (such as responding to remote user questions and marking content that the user is interested in).
[0056] In specific implementation, if the target application is an audio-visual type application, if the target trigger event is entering a video or audio playback page, playing a video or audio, pausing the playback of a video or audio, exiting a video or audio playback page, turning on a barrage, turning off a barrage, turning on comments, or turning off comments, then the target multimedia data includes relevant prompt identifiers; if the target trigger event includes liking at least one barrage, disliking at least one barrage, taking a screenshot, recording a screen, sharing a video or audio, reporting at least one barrage, posting a barrage, commenting, replying to a comment, liking at least one comment, and disliking at least one comment, then the target multimedia data includes relevant prompt identifiers, graphic data generated by combining the video playback content with user behavior data, audio data, and video data.
[0057] In specific implementation, if the target application is a live broadcast type application, if the target trigger event is one of entering a live broadcast room, exiting a live broadcast room, starting a live broadcast, and closing a live broadcast, then the target multimedia data includes a prompt mark; if the target trigger event is changing the anchor, then the target multimedia data may include anchor information, anchor evaluation information, and evaluation prompts for the anchor; if the target trigger time is changing the displayed goods, then the target multimedia data may include new displayed goods information, price information, evaluation information of the new displayed goods, and evaluation prompts for the new displayed goods obtained in combination with user historical behavior data.
[0058] In some embodiments, after the target multimedia data is generated, the target multimedia data is displayed in a display screen of the target application.
[0059] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, matching target multimedia data is generated and displayed by using the target feature data obtained when a target trigger event is triggered, thereby increasing the flexibility and dynamism of the target application in the interaction between users, satisfying the user's personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0060] Figure 2 A second optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0061] Step S201, obtaining category information of a target triggering event, and obtaining target feature data of the target application based on the category information.
[0062] In some embodiments, the category information of the target trigger event may include the type of the target trigger event, and different category information corresponds to different data acquisition methods and different target feature data types.
[0063] In some embodiments, if the target application is a game type application, the category information may include: application startup, the first target object releasing skills, the first target object switching game scenes, the first target object being attacked, the first target object achieving achievements, and the running state changing. The carrier determines the corresponding data acquisition method according to the type information of different target triggering events, and then obtains the target feature data of the target application.
[0064] For example, if the target trigger event is the start of a game, the game name and game identifier are obtained based on the category information; if the target trigger event is the first target object releasing a skill, the character data of the first target object and the skill data of the skill released by the first target object are obtained based on the category information; if the target trigger event is the first target object being attacked, the character data of the first target object and the character data and skill data of other characters attacking the first target object are obtained based on the category information; if the target trigger event is switching game scenes, the video content before and after the switch and the current background audio are obtained to produce a transition video or a transition special effect.
[0065] In some embodiments, if the target application is a conference-type application, the category information may include: scene switching by the remote host, host switching, screen projection content switching, screen projection device switching, other remote users speaking, other remote users giving likes, the near-end user starting screen recording, ending screen recording, the near-end user speaking, the near-end user turning on the camera, the near-end user turning off the camera, the near-end user giving likes, the near-end user inputting data, the near-end user turning on the hang-up mode, and the near-end user turning on the focus mode.
[0066] For example, if the type information of the target trigger event is that the remote host switches scenes, hosts, screen projection content, screen projection devices, other remote users speak, or other remote users like, then the remote user's data is obtained, such as scene data, host data, screen projection content data, other remote users' speeches, likes data, etc.; if the target trigger event is that the near-end user starts screen recording, ends screen recording, speaks, turns on the camera, turns off the camera, likes, inputs data, turns on the hang-up mode, and turns on the focus mode, then the near-end user's response information is obtained, such as receiving the near-end user's behavior data based on the near-end user's camera, microphone, and external device.
[0067] In some embodiments, if the target application is an audio and video application, the category information may include: entering the video or audio playback page, playing video or audio, pausing video or audio, exiting the video or audio playback page, turning on barrage, turning off barrage, turning on comments, turning off comments, liking at least one barrage, disliking at least one barrage, taking screenshots, recording screens, sharing video or audio, reporting at least one barrage, posting barrage, commenting, replying to comments, liking at least one comment, and disliking at least one comment.
[0068] For example, if the type information of the target trigger event is entering the video or audio playback page, playing video or audio, pausing the video or audio, exiting the video or audio playback page, turning on the barrage, turning off the barrage, turning on comments and turning off comments, then the user behavior data is obtained; if the type information of the target trigger event includes liking at least one barrage, disliking at least one barrage, taking screenshots, recording the screen, sharing video or audio, reporting at least one barrage, posting barrage, commenting, replying to comments, liking at least one comment and disliking at least one comment, then the corresponding video data (content data, actor data, scene data and action data), barrage data and comment data are obtained.
[0069] In some embodiments, if the target application is a live broadcast application, the category information may include: entering the live broadcast room, exiting the live broadcast room, the host displaying products, changing the host, opening the live broadcast room, closing the live broadcast room, users posting comments and users liking.
[0070] For example, if the type information of the target trigger event is entering the live broadcast room, exiting the live broadcast room, opening the live broadcast room and closing the live broadcast room, then the live broadcast room data and user behavior data are obtained; if the type of the target trigger event is the anchor displaying products, changing the anchor, users commenting and users liking, then the live broadcast page video data, product data and anchor data, as well as user behavior data are obtained.
[0071] Step S202: generating target multimedia data matching the target triggering event based on the target feature data.
[0072] The specific step flow of step S202 is the same as that of step S102 and will not be repeated here.
[0073] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, the target feature data is obtained through the target trigger event type information, and then the matching target multimedia data is generated and displayed based on the target feature data, thereby increasing the flexibility and dynamism of the target application in the interaction between users, meeting the user's personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0074] Figure 3 A third optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0075] Step S301: Obtain attribute information of a target application associated with the target trigger event, and obtain target feature data corresponding to the target trigger event based on the attribute information.
[0076] In some embodiments, the attribute information of the target application includes game attributes, meeting attributes, audio and video attributes, and live broadcast attributes; different attribute information can obtain different feature data when facing the same interactive behavior. For example, in the case of the same target trigger event of pausing the application, the first target object data and current game progress data of the game player are obtained for the game attributes, so as to generate a voice prompt matching the characteristics of the character or generate a hosting assistant based on the current progress; for the meeting attributes, the user's interest in the meeting and the meeting minutes are obtained to generate automatic reply response data; for the audio and video attributes, the current playing content and the user's likes or comments data are obtained, so as to subsequently generate interactive content that can interact with the user to enhance user engagement or stickiness; for the live broadcast attributes, the content that the user focuses on in the live broadcast is obtained to monitor changes in the content, give the user corresponding prompt information, and so on.
[0077] Step S302: generating target multimedia data matching the target triggering event based on the target feature data.
[0078] The specific step flow of step S302 is the same as that of step S102 and will not be repeated here.
[0079] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, the target feature data is obtained through the attribute information of the application corresponding to the target trigger event, and then the matching target multimedia data is generated and displayed based on the target feature data, thereby increasing the flexibility and dynamism of the target application in the interaction between users, meeting the user's personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0080] Figure 4 A fourth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0081] Step S401, obtaining source information of the target triggering event, and obtaining target feature data of the target application based on the source information.
[0082] In some embodiments, the source information includes a near-end user source, a far-end user source, a change in the operating state of the electronic device, or satisfaction of a time trigger condition, etc.
[0083] Specifically, the near-end user source includes the near-end user's operation on the first target object in the game application (including releasing skills, actively switching scenes, starting a game, etc.). In this scenario, the target feature data includes the near-end user's operation information, skill information, scene information and game information; it also includes the near-end user entering a meeting, leaving a meeting, liking, speaking, hosting and starting the focus mode in a meeting scenario. In this scenario, the target feature data includes meeting information, meeting video and audio information before and after triggering the target triggering event, remote user speech information and near-end user's historical behavior information; it also includes entering the video or audio playback page, playing video or audio, pausing video or audio, exiting the video or audio playback page, turning on barrage, turning off barrage, turning on comments, turning off comments, liking at least one barrage, disliking at least one barrage, taking screenshots, recording screens, sharing videos or audio, reporting at least one barrage, posting barrage, commenting, replying to comments, liking at least one comment and disliking at least one comment in a live broadcast scene. As well as entering the live broadcast room, exiting the live broadcast room, user comments and user likes in a live broadcast scene. That is, when the source information is a near-end user, the application screen information and information corresponding to the near-end user operation are obtained.
[0084] Specifically, the remote user sources include trigger events generated by other players' operations in game applications (such as attacking the first target object, invading the scene, attacking the building, etc.); remote user behaviors (speaking, liking, exiting) and host-side behaviors (changing display content, changing scenes, changing hosts, etc.) in conference applications; remote user behaviors in audio and video applications include posting bullet comments, liking bullet comments, and posting comments; remote user behaviors in live broadcast applications include host-side behaviors and other user behaviors. In this scenario, application screen information and information corresponding to remote user operations are obtained.
[0085] Specifically, the operating status of the electronic device changes or meets the time trigger conditions, including the start and end of a game in a gaming application; the arrival of a meeting time and the end of a meeting in a conference application; the arrival of a video clip playback time in an audio and video application; and the opening and closing of a live broadcast room in a live broadcast scene. In this scenario, the application screen information and operating status information are obtained.
[0086] In some embodiments, different source information obtains different target feature data. Specifically, the operation of releasing a skill (proximal user source) obtains the first target object and skill name; the end of the game (change in the operating state of the electronic device) obtains the battle result or highlight moment; the interactive operation of the user inputting text or voice obtains the behavior data or achievements of the character corresponding to the user in the game or application, etc., in order to evaluate the user's mood and polish or regenerate the interactive content to characterize the user's current state.
[0087] Step S402: generating target multimedia data matching the target triggering event based on the target feature data.
[0088] The specific steps of step S402 are the same as those of step S102 and will not be repeated here.
[0089] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, the target feature data is obtained through the source information corresponding to the target trigger event, and then the matching target multimedia data is generated and displayed based on the target feature data, thereby increasing the flexibility and dynamism of the target application in the interaction between users, meeting the user's personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0090] Figure 5 A fifth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0091] Step S501: Obtain a history log associated with the target trigger event, and obtain target feature data of the target application based on the history log.
[0092] In some embodiments, in the user pause and hosting scenario, the associated historical log is obtained to determine the target feature data of the target application corresponding to the same or similar historical target trigger event. Then, the history can be used this time or data different from the historical feature data can be obtained.
[0093] For example, in a game scenario, for pause and hosting, user behavior data can be learned from historical logs, and then the first target object can be manipulated based on the learned user behavior data; in a meeting scenario, user behavior data can be learned from historical logs to complete operations with other users, or key meeting content or content of interest to users can be marked based on the learned user behavior data; in a live broadcast scenario, user behavior data can be learned from historical logs to complete operations with the host and pre-order operations.
[0094] Step S502: generating target multimedia data matching the target triggering event based on the target feature data.
[0095] The specific step flow of step S502 is the same as that of step S102 and will not be repeated here.
[0096] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, the target feature data is obtained through the historical logs associated with the target trigger event, and then the matching target multimedia data is generated and displayed based on the target feature data, thereby increasing the flexibility and dynamism of the target application in the interaction between users, meeting the user's personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0097] Figure 6 A sixth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0098] Step S601 , in response to obtaining a target trigger event for a target application, obtaining target feature data of the target application.
[0099] In some embodiments, in response to obtaining a first operation that triggers the start of a target application, application identification data of the target application is obtained. The first operation may include single-clicking the target application identifier, double-clicking the target application identifier, voice operation, and gesture operation, the purpose of which is to start the target application. The application identification data of the target application includes the application name of the target application, the application version of the target application, the ID of the target application, the identifier (or icon) of the target application, and the type of the target application. Furthermore, if the type of the target application is different, the style of the corresponding generated target multimedia data is also different.
[0100] In some embodiments, in response to obtaining a second operation acting on the target application, the behavior data and / or configuration data of the operated object in the target application are obtained, and the second operation is an operation that can trigger the target application to provide a target function. Wherein, the second operation may include a skill release operation (skills include ordinary attacks, first skills, second skills, ultimate skills, and dodge skills), the operated object includes the first target object operated by the proximal user, the behavior data includes operation data and operation data, such as movement data, kill data, dodge data, kill data, attacked data, attack data, and skill name; the configuration data includes character data and props data corresponding to the first target object; the target function includes releasing skills. That is, in response to the user releasing a skill, the behavior data and / or configuration data of the first target object in the game application are obtained, and based on the behavior data and / or configuration data, the corresponding skill special effects are generated.
[0101] In some embodiments, in response to obtaining a third operation of switching the operated object in the target application from the first object to the second object, the configuration data and / or associated background data of the first object and the second object are obtained; wherein the first object includes one or more objects; the second object includes one or more objects, and the first object and the second object do not completely overlap, that is, the target trigger event is the user switching from operating the first object to operating the second object; in this scenario, the target feature data includes the configuration data and / or associated background data of the first object and the second object. The configuration data includes the number of first objects, the number of second objects, the props, level, identity, achievement data, combat power, etc. configured by the user for each object; the associated background data may include map data of the area where the first object and the second object are located, and / or background data and environmental data of the scene where the character is located.
[0102] In some embodiments, in response to the operated object in the target application moving from a first position to a second position, the environmental data of the first position and the second position are obtained. The moving from the first position to the second position may include moving from one coordinate to another coordinate in the same scene, or may include moving from one scene to another scene in different scenes; the environmental data of the first position and the second position include scene data, map data, and environmental data.
[0103] In some embodiments, in response to monitoring the increase in the load of the target processor, application identification data of a first application is obtained, where the first application is the application that causes the increase in the load of the target processor; wherein the target processor may include a graphics card, a CPU, and a memory, etc. The increase in the load of the target processor indicates that a new application has been started, and in this scenario, application data of the first application (i.e., the newly started application) is obtained. The application data includes name information, version information, application identification data, etc.
[0104] In some embodiments, in response to the target application switching from a first running state to a second running state, evaluation data for the first target object in the target application is obtained, wherein the resource occupancy of the target application in the second running state is less than the resource occupancy in the first running state; for example, the first running state may include a game state, and the second running state may include a game end, a waiting state, and a closed state. The target application switches from the first running state to the second running state, which may include a game in progress to a game end, at which time the CPU, memory, etc. occupancy of the target application will become lower. The target feature data may include evaluation data for the first target object. The evaluation data may include data sent by the application (such as data obtained based on the performance of the first target object in the game), and may also include data sent by other users in the same game; the evaluation data may include the achievements, operation ratings, championship and runner-up rankings, number of kills, and number of resources obtained by the first target object in the game.
[0105] In some embodiments, in response to the target application switching from a first running state to a third running state, behavior data and / or configuration data of a first target object in the target application are obtained, and the resource occupancy of the target application in the third running state is less than or greater than the resource occupancy in the first running state; wherein the third running state includes a paused state or a hosted state, then the behavior data (such as historical behavior data) and / or configuration data of the target object are obtained, target multimedia data is generated based on the historical behavior data, and the user's control of the first target object is learned and imitated.
[0106] In some embodiments, in response to obtaining interactive input with a target object in the target application, at least one of configuration data, behavior data, and evaluation data for the target object is obtained. The interactive input includes likes, comments, voice data, and text data; at least one of configuration data, behavior data, and evaluation data for the target object is obtained to determine the emotional characteristics of the user and output target multimedia data that matches the current game.
[0107] Step S602: generating target multimedia data matching the target triggering event based on the target feature data.
[0108] The specific step flow of step S602 is the same as that of step S102 and will not be repeated here.
[0109] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, target feature data is obtained by related operations on the target application or objects in the target application, and then matching target multimedia data is generated and displayed based on the target feature data, thereby increasing the flexibility and dynamism of the interaction between users in the target application, meeting the user's personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0110] Figure 7 A seventh optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0111] Step S701, in response to a target triggering event, obtaining target feature data of a target application.
[0112] The specific step flow of step S701 is the same as the step flow of any one or at least two combinations of step S101, step S201, step S301, step S401, step S501 and step S601, and will not be repeated here.
[0113] Step S702: generating target multimedia data matching the target triggering event based on the target processing model and the target feature data.
[0114] In some embodiments, the carrier can directly call the target processing model to generate and process the target feature data to obtain target multimedia data that matches the target trigger event; the target multimedia data includes at least one of target audio data, target image data, target video data, and target text data.
[0115] In other embodiments, for scenarios where the target application cannot directly call the target processing model, the target application can send the target feature data to the first application, and use the first application to call the target processing model to generate and process the target feature data to obtain at least one of the target audio data, target image data, target video data, and target text data that matches the target trigger event. The first application may include an application capable of calling a large model, such as Lenovo Xiaotian or AI NOW.
[0116] In the game scene, the target audio data may include background audio data, skill audio data, and interactive feedback audio data; the target image data may include game poster data, highlight moment image data (such as gif images); the target video data may include at least one of highlight moment video, transition video, and object switching video; the target text data may include game data statistics or slogans. Among them, the background audio data may include audio data with emotional colors generated according to the game information. For example, if the party where the first target object is located in the current game is at a disadvantage, the background audio data may have emotional colors of grief, tension, or motivation; if the party where the first target object is located in the current game is at an advantage, the background audio data may have emotional colors of excitement, relaxation, and pleasure.
[0117] In a conference scenario, the target audio data may include background audio data (such as welcome-type audio data); the target image data may include poster data of the product (item or plan) being discussed, poster data of the host information, poster data of the speaker, or poster data of the scene information; the target video data may include video data of the product, host, speaker, or scene being discussed; the target text data may include text data of the product, host, speaker, or scene being discussed.
[0118] In the audio and video scene, the target video data may include video data corresponding to the user's operation behavior; for example, if the user clicks "like", a "like" animation or a "heart" animation will appear on the screen; if the user pauses, a "pause" icon will appear on the screen. It may also include target text data pre-generated according to the current plot, so that users can quickly choose to express their own evaluation of the current plot.
[0119] In a live broadcast scenario, the target audio data may include background audio data (such as welcome-type audio data); the target image data may include poster data of the product being introduced, poster data of the anchor information, or poster data of the scene information; the target video data may include video data of the product, anchor, or scene being introduced, and may also include video data corresponding to user operation behavior, such as a like animation or a heart animation that appears on the screen for a like behavior; the target text data may include video data of the product, anchor, or scene being introduced.
[0120] In some embodiments, the carrier may also obtain interactive input data between the target user and the target application, and generate target multimedia data matching the interactive input data based on the interactive input data and the target feature data. Specifically, when the target trigger event is generated based on the interactive behavior between the user and other players in the game or interactive objects of other applications, the user's interactive input data may be further obtained and also input into the large model as a prompt to generate multimedia data matching the current scene and user needs.
[0121] Further, the carrier can predetermine the target processing model, including: the carrier determines the required processing model based on at least one of the category information of the target trigger event, the interactive input data between the target user and the target application, and the user portrait data of the target user, so as to call the determined target processing model to generate and process the target feature data.
[0122] Specifically, the carrier can determine the type of target multimedia data to be generated according to the category information of the target trigger event, and then determine the target processing model. For example, if the category information is application startup, the corresponding target processing model is an image generation model; if the category information is the first target object releasing a skill, the corresponding target processing model is a video generation model; if the category information is screen recording, the corresponding target processing model is an image generation model; if the category information is host switching, the corresponding target processing model is an image generation model or a video generation model.
[0123] In some embodiments, the carrier may also determine the required processing model based on at least one of the category information of the target trigger event, the interactive input data between the target user and the target application, and the user portrait data of the target user, so as to call the determined target processing model to generate and process the target feature data. The interactive input data may contain emotional tendencies and intentions; the user portrait data of the target user can help determine the types that the user likes and the user's emotional tendencies and intentions.
[0124] In other embodiments, the target processing model may include a first processing model and a second processing model; the carrier determines the emotional attributes or intentions of the target user represented by the target feature data based on the first processing model, and then inputs the emotional attributes and intentions, as well as the target feature data into the second processing model for generation processing to obtain target multimedia data.
[0125] Specifically, in a game scenario, the first processing model can determine the target user's emotional attributes based on the target feature data. For example, if the party to which the first target object belongs is at a disadvantage in the current game, the emotional attributes may include grief, anger, tension, or motivation; if the party to which the first target object belongs is at an advantage in the current game, the emotional attributes may include excitement or relaxation. The emotional attributes and the target feature data are then input into the second processing model to generate target multimedia data.
[0126] In conference, audio and video and live broadcast scenarios, the target user's intention and emotional attributes can be determined based on the target feature data, such as the user's intention to speak, to welcome the host, to express his or her own opinions, etc.; the intention and target data are input into the second processing model to generate target multimedia data.
[0127] In this way, through the multimedia data generation method provided by the embodiment of the present disclosure, the target feature data of the target application is obtained through the target trigger event, and then the target processing model is determined based on the target feature data, and then the matching target multimedia data is generated and displayed, thereby increasing the flexibility and dynamics of the interaction between the target application and the users, meeting the users' personalized needs while increasing the user experience, attracting users and maintaining user stickiness.
[0128] Figure 8 An eighth optional flow chart of the multimedia data generating method provided in the embodiment of the present disclosure is shown, and will be explained according to each step.
[0129] Step S901, in response to a target triggering event, obtaining target feature data of a target application.
[0130] The specific step flow of step S901 is the same as the step flow of any one or at least two combinations of step S101, step S201, step S301, step S401, step S501 and step S601, and will not be repeated here.
[0131] Step S902: Obtain template data corresponding to the target feature data, and generate corresponding target multimedia data based on the template data and the target feature data.
[0132] In some embodiments, the carrier can pre-generate template data corresponding to different target feature data; for example, when the target feature data includes target application information (such as name, version, ID), the corresponding template data can be image data, including elements such as "startup". After obtaining the target feature data, the target application information is filled into the template data to generate target multimedia data.
[0133] Alternatively, the target feature data includes information of the host, spokesperson or anchor, and the corresponding template data may be image data, including a character part and a text part. After obtaining the target feature data, the picture of the host, spokesperson or anchor is filled into the character part, and the profile of the host, spokesperson or anchor is filled into the text part.
[0134] Alternatively, the target feature data includes a skill name, and the corresponding template data can be audio data. The skill name is filled into the template data, and the audio data corresponding to the skill name is output. The emotional attribute corresponding to the target feature data can also be determined, and the corresponding template data is determined based on the emotional attribute. The skill name is filled into the corresponding template data, and audio data with emotional attributes is output.
[0135] Alternatively, the target feature data includes products discussed in the meeting or products displayed by the anchor, and the corresponding template data can be image data, including a product image part and a product text part. After obtaining the target feature data, the image of the product is filled into the image part, and the relevant information of the product (such as introduction, project name, type, price, etc.) is filled into the text part to obtain the image data.
[0136] Step S903: output the target multimedia data to the target area based on the running state of the target application and / or the user portrait data of the target user.
[0137] In some embodiments, based on the running state of the target application, the target area for outputting the target multimedia data is determined; for example, the target application is a game-type application, and if the running state of the target application includes being in a game, then the area that does not obstruct the characters, skills, and key information of the game is determined to be the target area; if the running state of the target application includes the end of the game, then the center of the screen is determined to be the target area.
[0138] In other embodiments, the carrier may also determine the target area based on user portrait data of the target user; the user portrait data is used to determine the user's habits or preferences, and the multimedia data is output to the target area according to the user's habits or preferences.
[0139] In some embodiments, the carrier may also determine the target area by comprehensively considering the running status of the target application and the user portrait data of the target user.
[0140] The target area includes a screen display area and may also include a sound output area, such as a microphone, a left channel, a right channel, etc. For example, if it is determined based on user portrait data that the user is used to wearing headphones in the left ear, the left channel is output.
[0141] Thus, through the multimedia data generation method provided by the embodiment of the present disclosure, corresponding target multimedia data can be generated based on the template data combined with the target feature data; the target area can be determined based on the running state of the target application and the user portrait data, and then the target multimedia data can be output to the target area. Generating target multimedia data based on the template can improve the processing speed; determining the target area based on the running state and the user portrait data can improve the user's experience of receiving the target multimedia data during the use of the application.
[0142] Fig. 9 A first optional flow chart of the audio data generating method provided in an embodiment of the present disclosure is shown, and will be explained according to each step.
[0143] Step S1001, in response to a target operation acting on a target game application, obtaining target application data of the target game application.
[0144] In some embodiments, the target operation may include starting the target game application (such as a stand-alone icon, double-clicking an icon, voice control, and gesture control), the first target object corresponding to the user in the target application releasing skills, the first target object switching game scenes, the first target object being attacked, the first target object achieving achievements, and the change in the running status of the target application.
[0145] In some embodiments, the target application data may include game data of the target game application (such as game name, version, game identifier), character data of the first target object, skill data of the skills released by the first target object, scene data before and after the first target object switches the game scene, character data of the first target object when the first target object is attacked and character data and skill data of other characters attacking the first target object, video data before the first target object achieves an achievement, character data and skill data of other characters in the same screen as the first target object, and running status change data of the target application.
[0146] Step S1002: Generate target audio data matching the target operation based on the target application data.
[0147] In some embodiments, the carrier may generate the target audio data based on a large model, and may also generate the target audio data based on template data.
[0148] In some embodiments, if the target operation includes the start-up operation of the target game application, the game data of the target game application is the application identification data, and the carrier inputs the game data of the target game application into the large model to generate the target audio data of the game start-up, such as "XX game start-up". Alternatively, the carrier inputs the game data of the target game application into the target data to generate the target audio data of the game start-up. The effects achieved include: the user triggers the start-up of the target game application through operations such as a single-machine icon, double-clicking an icon, voice control, and gesture control, and receives the target audio data of "XX game start-up".
[0149] In some embodiments, if the target operation includes controlling the first target object to release a skill (the skill includes a normal attack, a first skill, a second skill, a big skill, and a dodge skill), the game data of the target game application is at least one of the skill parameters of the skill, the information of the first target object, and the game application screen data, and the carrier inputs the game data of the target game application into a large model (such as an audio generation model) to generate corresponding skill audio data and / or skill effect video data. Alternatively, the carrier inputs the target application data into template data (such as audio template data) to generate corresponding skill audio data and / or skill effect video data. The effect achieved includes: the user releases a skill, and receives skill audio data and / or skill effect video data. The skill audio data may include skill effect data, such as when the skill effect is an explosion, the corresponding skill audio data is an explosion audio data; when the skill effect is rain, the corresponding skill audio data is rain audio data; when the skill effect is entanglement, the corresponding skill audio data includes entanglement audio data; or, the skill audio data may include a skill name, that is, the audio data of the skill name is output while the user releases the skill. The skill effect video data may include special effects (such as explosions, rainfall, and entanglement videos), and may also include the value of the damage caused.
[0150] In some optional embodiments, the carrier may also determine the emotional attributes of the game players based on the current game information, and generate target audio data with emotional attributes based on the emotional attributes and the target application data. For example, when the battle is going well, the generated sound effect is an excited sound effect "Attack, take them down!"; when the battle is going badly, the generated sound effect is an angry sound effect "Fight back, for the alliance!".
[0151] In some embodiments, if the target operation includes controlling the movement of the first target object or switching the game scene, the game data of the target game application includes the character data of the first target object, the scene data before the first target object switches the game scene, and the scene data after the switching. The carrier inputs the game data of the target game application into the large model to generate corresponding prompt audio data and / or prompt video data. Alternatively, the carrier inputs the target application data into the template data to generate corresponding skill audio data and / or skill effect video data. The effects achieved include: the user controls the movement of the first target object, outputs the video data of the switching scene, and the audio data of the switching scene. The video data may include text prompts, such as "Entered XX area", "Switched map", etc.; it may also include scene images before and after the switching; the audio data may include text voice prompts, such as "Entered XX area", "Switched map", and may also include background music of the scene after the switching.
[0152] In some embodiments, if the target operation includes the first target object being attacked, the game data of the target game application includes the character data of the first target object when the first target object is attacked and the character data and skill data of other characters attacking the first target object, and the carrier inputs the game data of the target game application into the large model to generate corresponding prompt audio data and / or prompt video data. Alternatively, the carrier inputs the target application data into the template data to generate corresponding skill audio data and / or skill effect video data. The effects achieved include: when the first target object is attacked, an audio and / or video warning of being attacked is output, and the skill audio and / or video of other characters causing damage to the first target object is output. The skill audio of other characters includes skill names and / or skill effect audio; the skill video of other characters includes skill special effects.
[0153] In some embodiments, if the target operation includes the first target object achieving an achievement, the target application data includes video data before the first target object achieves the achievement and feedback data from other objects in the unified game to the first target object after the first target object achieves the achievement, and the carrier inputs the game data of the target game application into the large model to generate corresponding achievement audio data and / or achievement video data. Alternatively, the carrier inputs the target application data into the template data to generate corresponding skill audio data and / or skill effect video data. The effects achieved include: after the first target object achieves the achievement, outputting highlight videos and data (including output data and kill data) when the achievement is achieved, as well as feedback data from other objects.
[0154] In some embodiments, if the target operation includes pausing the game, the target application data includes the historical behavior data of the first target object, and the carrier inputs the target application data into the large model, generates the behavior data of the first target object in combination with the game data, and controls the first target object to implement the behavior data. The effects achieved include: when the user pauses the game, the auxiliary function takes over the control of the first target object, and controls the first target object to continue playing the game according to the learned user behavior.
[0155] In some embodiments, if the target operation includes a change in running state, the target application data includes game status, game data and evaluation data, and the target application data is input into the large model to generate a settlement video and settlement audio. The effects achieved include: when the game ends, the first target object's achievements, operation ratings, championship and runner-up rankings, number of kills and number of resources obtained in the game, as well as highlight video data are output.
[0156] In some embodiments, if the target operation includes the game character used by the game player switching from a first character to a second character, the configuration data and / or associated background data of the first character and the second character are obtained; wherein the first character includes one or more characters; the second character includes one or more characters, and the first character and the second character do not completely overlap, that is, the target trigger event is the game player switching from operating the first character to operating the second character; in this scenario, the target feature data includes the configuration data and / or associated background data of the first character and the second character. The configuration data includes the number of first characters, the number of second characters, the props, level, identity, achievement data, combat power, etc. configured by the game player for each character; the associated background data may include map data of the area where the first character and the second character are located, and / or background data and environmental data of the scene where the characters are located. The target application data is input into the large model to generate corresponding switching prompt audio data, as well as audio data related to the second character.
[0157] In some embodiments, if the target operation includes interactive input with a remote game player in the target game application, at least one of the configuration data, behavior data, and evaluation data of the target game player (including other local players or other remote players) in the target game application is obtained, the target application data is input into the big model, and audio data of the target game player is generated.
[0158] In this way, through the audio data generation method provided by the embodiment of the present disclosure, the real-time background of the game and the behavior of the player character are used to generate corresponding emotional sound effects or prompts to enhance the player's gaming experience.
[0159] Fig.10A second optional flow chart of the audio data generating method provided in an embodiment of the present disclosure is shown, and will be explained according to each step.
[0160] In some embodiments, steps S1101 to S1104 may be implemented based on a large model.
[0161] Step S1101, analyzing the screen to determine the information of the target game application.
[0162] In some embodiments, the user starts the target game application through a startup operation, and the large model determines the name of the target game application and the game running path based on the electronic device display screen.
[0163] Further, in response to the user starting a game, the large model determines information about the first target object selected by the user based on a display screen of the electronic device, where the information about the first target object includes character information and skill information of the first target object.
[0164] Step S1102: Acquire voice data of the first target object.
[0165] In some embodiments, the large model searches for the voice data of the first target object based on the game running path, wherein the voice data may include the voice data of the first target object in the game, and may also include the voice data recorded by the user.
[0166] In some embodiments, if the voice data is not found, the voice data is obtained from the official website corresponding to the target game application; if the voice data does not exist in the official website corresponding to the target game application, it is confirmed whether the user has customized an exclusive sound effect.
[0167] In response to the user's decision to customize the exclusive sound effect, a prompt message is displayed to instruct the user to record voice data based on the prompt message; the characteristics of the voice data are learned, and the voice data recorded by the user is determined in combination with the skill name and effect of the first target object.
[0168] Step S1103: Generate target voice data based on the voice data.
[0169] In some embodiments, the large model generates target voice data with emotional attributes based on the skill emotional characteristics and skill name of the first target object. For example, for highly aggressive skills, aggressive target voice data is output; for resurrection, treatment and other type skills, compassionate target voice data is output.
[0170] In some optional embodiments, the large model can also generate target speech data with different emotional attributes for the same skill name.
[0171] Step S1104: output corresponding target voice data based on the user operation.
[0172] In some embodiments, the large model determines the skills released by the user operating the first target object based on the display screen of the electronic device, and outputs corresponding target voice data based on the skills released by the first target object.
[0173] In some embodiments, the large model can also output target voice data with emotional attributes based on the current game information and the skills released by the user operating the first target object. For example, in a favorable battle, when the attack is started, the generated sound effect is an excited sound effect "Attack, take them down!"; in an unfavorable battle, when the counterattack is started, the generated sound effect is a sad and angry sound effect "Fight back, for the alliance!"
[0174] In this way, through the audio data generation method provided by the embodiment of the present disclosure, sound effects other than in-game voices can be generated, and corresponding sound effects can be generated according to the different skills released by different characters used by players; corresponding emotional sound effects can be generated according to the real-time background of the game and the behavior of the player's character, thereby enhancing the user's gaming experience.
[0175] Next, the multimedia data generation method involved in the embodiment of the present disclosure is illustrated by examples in combination with actual scenarios.
[0176] In response to the target application being a conference type application, if the target trigger event is entering a conference, the target feature data includes conference related information, that is, obtaining subject information, host information, etc., and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes an entry conference identification, a camera status identification, and a microphone status identification. Optionally, the target multimedia data may also include audio / video data of a welcome message for the conference, and the welcome message may include subject information and host information. For example, "Welcome to join the conference with the subject of AAA hosted by Zhang San."
[0177] In response to the target application being a conference-type application, if the target trigger event is a scene switch of the remote host (from camera capture to screen projection), the target feature data includes the scene information after the switch, and the content information displayed after the switch, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes text data, audio data, and video data of scene switch prompts and evaluation prompts of new scenes. For example, the target multimedia data can be audio and video data of "scene switched", and can also be text data of evaluation prompts for new scenes to facilitate users to quickly comment.
[0178] In response to the target application being a conference type application, if the target triggering event is a host switch or a speaker switch, the target feature data includes new host information or new speaker information, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes a host or speaker switch prompt, text data of an evaluation prompt of the new host or new speaker, and audio and video data of the new host or new speaker information. For example, the target multimedia data can be audio and video data of "host / speaker has switched", and can also be text data of an evaluation prompt for the new host or new speaker to facilitate users to quickly comment, such as "Welcome Teacher XX to speak", "Finally, Teacher XX has arrived", etc.
[0179] In response to the target application being a conference-type application, if the target trigger event is a switch in the projection content, the target characteristic data includes new projection content, and the target characteristic data is generated into target multimedia data based on the method described in step S702 or step S902. The target multimedia data includes audio and video data of the switched projection content, and may also include an introduction or summary of the new projection content.
[0180] In response to the target application being a conference type application, if the target triggering event is speech of other remote users and likes of other remote users, the target feature data includes speech information and likes information of the remote users, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes speech prompts of remote users, such as "XX is speaking." Or operation prompts of remote users, such as "XX has liked."
[0181] In response to the target application being a conference-type application, if the target trigger event is to start screen recording or end screen recording, the target feature data includes user operation data, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes an audio and video prompt of "screen recording has started" or an audio and video prompt of "screen recording has ended".
[0182] In response to the target application being a conference type application, if the target triggering event is a near-end user operation, the target feature data includes near-end user operation data, such as speech data, camera on, camera off, like, etc., and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes "microphone is on, you can speak", "camera is on / off", "you like XX".
[0183] In response to the target application being a conference-type application, if the target triggering event is the near-end user turning on the on-hook mode or the hosting mode, the target feature data includes historical user behavior data and user attention data, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes user behavior learned based on historical user behavior data. For example, after the user turns on the on-hook mode or the hosting mode, the content that the user is interested in can be determined based on the historical user behavior data, and related operations can be performed based on the content that the user is interested in, such as recording, marking key points, generating meeting minutes, helping the near-end user record and responding to other people's interactions, etc.
[0184] In response to the target application being an audio-visual application, if the target triggering event is entering a video or audio playback page, the target characteristic data includes content data played by the application, and the target characteristic data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes a welcome message for welcoming the user to watch the video or listen to the audio. For example, "The next video to be played for you is XX", "Welcome to watch XX starring AA".
[0185] In response to the target application being an audio-visual type application, if the target triggering event is playing a video or audio, or pausing the playing of a video or audio, the target feature data includes a play status prompt, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes audio and video of the play status prompt. For example, audio data of "playing XX for you", "pause playing XX", "continue playing XX", and an image or video prompt of the play status.
[0186] In response to the target application being an audio-visual type application, if the target triggering event is opening barrage, closing barrage, opening comments, closing comments, the target feature data includes the barrage playback state or the comment display state, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes audio and video data or graphic data of the barrage playback state or the comment display state. For example, audio and video data or graphic data of "barrage is opened / closed" and "comments are opened or closed".
[0187] In response to the target application being an audio and video type application, if the target trigger event is to like at least one barrage or to dislike at least one barrage, the target feature data includes a like operation or a dislike operation, and the target feature data is generated into target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes a like animation and audio, or a dislike animation and audio.
[0188] In response to the target application being an audio-visual type application, if the target trigger event is to post a bullet screen or a comment, the target feature data includes the content data, actor data, scene data, action data, bullet screen data, and comment data played by the application when the target trigger event is received, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes at least one bullet screen or comment text generated based on the current content data, actor data, scene data, action data, bullet screen data, and comment data, as well as bullet screens or comments posted by users in the same or similar scenes in historical user data. So that users can select the bullet screen or comment that is consistent with what they want to express based on the at least one generated bullet screen or comment text, and quickly post a bullet screen or comment.
[0189] In response to the target application being a live broadcast type application, if the target triggering event is entering a live broadcast room, exiting a live broadcast room, opening a live broadcast room, and closing a live broadcast room, the target feature data includes live broadcast room information and live broadcast room status information, and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes user status information and live broadcast room status information. For example, "Welcome to XX live broadcast room", "You have exited XX live broadcast room", "XX live broadcast room has started broadcasting", "XX live broadcast room has stopped broadcasting".
[0190] In response to the target application being a live broadcast type application, if the target triggering event is the anchor displaying a product, the target feature data includes the product information of the displayed product (including specifications, models, prices, brands, appearance and functions), and the target feature data is used to generate target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes audio and video information or graphic information corresponding to the product information. Optionally, if the user has searched for the displayed product in historical behavior (searched in other shopping software or searched in the live broadcast room), or the user sets the displayed product as a product of interest, the target multimedia data may also include prompt information, such as "the product you are interested in is being explained."
[0191] In response to the target application being a live broadcast type application, if the target trigger event is changing the anchor, the target feature data includes new anchor information and prompt information for the anchor, and the target feature data is generated into target multimedia data based on the method described in step S702 or step S902, and the target multimedia data includes anchor information, anchor evaluation information and evaluation prompts for the anchor.
[0192] Fig.11 An optional structural schematic diagram of a multimedia data generating device provided in an embodiment of the present disclosure is shown, and will be described according to each part.
[0193] In some embodiments, the multimedia data generating device 1200 includes a first acquiring unit 1201 and a first generating unit 1202 .
[0194] The first acquisition unit 1201 is used to obtain target feature data of a target application in response to a target trigger event;
[0195] The first generating unit 1202 is configured to generate target multimedia data matching the target triggering event based on the target feature data;
[0196] The target trigger event is an event that can cause the operating parameters of the target application to change.
[0197] The first acquisition unit 1201 is specifically used for at least one of the following:
[0198] Obtaining category information of the target trigger event, and obtaining target feature data of the target application based on the category information;
[0199] Obtaining attribute information of a target application associated with the target trigger event, and obtaining target feature data corresponding to the target trigger event based on the attribute information;
[0200] Obtaining source information of the target triggering event, and obtaining target feature data of the target application based on the source information;
[0201] A historical log associated with the target triggering event is obtained, and target feature data of the target application is obtained based on the historical log.
[0202] The first acquisition unit 1201 is specifically used for at least one of the following:
[0203] In response to obtaining a first operation that triggers the start of a target application, obtaining application identification data of the target application;
[0204] In response to obtaining a second operation acting on the target application, obtaining behavior data and / or configuration data of an operated object in the target application, wherein the second operation is an operation capable of triggering the target application to provide a target function;
[0205] In response to obtaining a third operation of switching the operated object in the target application from a first object to a second object, obtaining configuration data and / or associated background data of the first object and the second object;
[0206] In response to an operated object in the target application moving from a first position to a second position, obtaining environmental data of the first position and the second position;
[0207] In response to monitoring an increase in the load of a target processor, obtaining application identification data of a first application, the first application being an application causing the increase in the load of the target processor;
[0208] In response to the target application switching from a first running state to a second running state, obtaining evaluation data for a first target object in the target application, wherein resource occupation of the target application in the second running state is less than resource occupation in the first running state;
[0209] In response to the target application switching from the first running state to the third running state, obtaining behavior data and / or configuration data of a first target object in the target application, wherein resource usage of the target application in the third running state is less than or greater than resource usage in the first running state;
[0210] In response to obtaining an interactive input with a target object in the target application, at least one of configuration data, behavior data, and evaluation data for the target object is obtained.
[0211] The first generating unit 1202 is specifically configured to do at least one of the following:
[0212] Calling a target processing model to generate and process the target feature data to obtain at least one of target audio data, target image data, target video data, and target text data that matches the target trigger event;
[0213] The target feature data is provided to a first application capable of calling a target processing model, and the target feature data is generated and processed by calling the target processing model using the first application to obtain at least one of target audio data, target image data, target video data, and target text data that matches the target trigger event;
[0214] Interaction input data between a target user and the target application is obtained, and target multimedia data matching the interaction input data is generated based on the interaction input data and the target feature data.
[0215] The first generating unit 1202 is specifically configured to do at least one of the following:
[0216] Determine the required processing model based on at least one of the category information of the target trigger event, the interactive input data between the target user and the target application, and the user portrait data of the target user, so as to call the determined target processing model to generate and process the target feature data;
[0217] The emotional attribute of the target user represented by the target feature data is determined based on the first processing model, and the emotional attribute and the target feature data are input into the second processing model for generation processing to obtain the target multimedia data.
[0218] The first generating unit 1202 is specifically configured to do at least one of the following:
[0219] Obtaining template data corresponding to the target feature data, and generating corresponding target multimedia data based on the template data and the target feature data;
[0220] The target multimedia data is output to a target area based on the running state of the target application and / or the user portrait data of the target user.
[0221] Fig.12 An optional structural diagram of the audio data generating device provided in an embodiment of the present disclosure is shown, and will be described according to each part.
[0222] In some embodiments, the audio data generating device 1300 includes a second acquiring unit 1301 and a second generating unit 1302 .
[0223] The second acquisition unit 1301 is used to obtain target application data of the target game application in response to a target operation acting on the target game application;
[0224] The second generating unit 1302 is used to generate target audio data matching the target operation based on the target application data;
[0225] The target operation can trigger changes in operating parameters of the target game application.
[0226] The second acquisition unit 1301 is specifically used for at least one of the following:
[0227] In response to obtaining an operation of releasing a game ultimate move in a target game application, obtaining at least one of a skill parameter of the game ultimate move, character information of a game player, and game application screen data;
[0228] In response to obtaining an operation that triggers the start of a target game application, obtaining application identification data of the target application;
[0229] In response to obtaining an operation of switching a game character used by a game player in the target game application from a first character to a second character, obtaining configuration data and / or associated background data of the first character and the second character;
[0230] In response to obtaining interactive input with a remote game player in the target game application, at least one of configuration data, behavior data, and evaluation data of the target game player in the target game application is obtained.
[0231] The second generating unit 1302 is specifically used for at least one of the following:
[0232] Calling the audio generation model to generate and process the target application data to obtain target audio data matching the target operation;
[0233] Obtaining audio template data corresponding to the target application data, and generating the target audio data based on the audio template data and the target application data;
[0234] An emotional attribute of a game player is determined based on the target application data, and the target audio data is generated based on the emotional attribute and the target application data.
[0235] The present disclosure also provides an electronic device, comprising a processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following:
[0236] In response to a target triggering event, obtaining target feature data of a target application;
[0237] Generate target multimedia data matching the target trigger event based on the target feature data;
[0238] The target trigger event is an event that can cause the operating parameters of the target application to change.
[0239] Alternatively, the electronic device includes a processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following:
[0240] In response to a target operation acting on a target game application, obtaining target application data of the target game application;
[0241] generating target audio data matching the target operation based on the target application data;
[0242] The target operation can trigger changes in operating parameters of the target game application.
[0243] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0244] Fig.13A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0245] like Fig.13 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0246] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0247] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a multimedia data generation method or an audio data generation method. For example, in some embodiments, the multimedia data generation method or the audio data generation method may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the multimedia data generation method or the audio data generation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the multimedia data generating method or the audio data generating method in any other appropriate manner (for example, by means of firmware).
[0248] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0249] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0250] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0251] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0252] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0253] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0254] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0255] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0256] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A method for generating multimedia data, comprising: In response to a target triggering event, obtaining target feature data of a target application; Generate target multimedia data matching the target trigger event based on the target feature data; The target trigger event is an event that can cause the operating parameters of the target application to change.
2. The method according to claim 1, wherein: In response to the target trigger event, target feature data of the target application is obtained, including at least one of the following: Obtaining category information of the target trigger event, and obtaining target feature data of the target application based on the category information; Obtaining attribute information of a target application associated with the target trigger event, and obtaining target feature data corresponding to the target trigger event based on the attribute information; Obtaining source information of the target triggering event, and obtaining target feature data of the target application based on the source information; A historical log associated with the target triggering event is obtained, and target feature data of the target application is obtained based on the historical log.
3. The method according to claim 1 or 2, wherein: In response to the target trigger event, target feature data of the target application is obtained, including at least one of the following: In response to obtaining a first operation that triggers the start of a target application, obtaining application identification data of the target application; In response to obtaining a second operation acting on the target application, obtaining behavior data and / or configuration data of an operated object in the target application, wherein the second operation is an operation capable of triggering the target application to provide a target function; In response to obtaining a third operation of switching the operated object in the target application from a first object to a second object, obtaining configuration data and / or associated background data of the first object and the second object; In response to an operated object in the target application moving from a first position to a second position, obtaining environmental data of the first position and the second position; In response to monitoring an increase in the load of a target processor, obtaining application identification data of a first application, the first application being an application causing the increase in the load of the target processor; In response to the target application switching from a first running state to a second running state, obtaining evaluation data for a first target object in the target application, wherein resource occupation of the target application in the second running state is less than resource occupation in the first running state; In response to the target application switching from the first running state to the third running state, obtaining behavior data and / or configuration data of a first target object in the target application, wherein resource usage of the target application in the third running state is less than or greater than resource usage in the first running state; In response to obtaining an interactive input with a target object in the target application, at least one of configuration data, behavior data, and evaluation data for the target object is obtained.
4. The method according to claim 1, wherein: Generating target multimedia data matching the target trigger event based on the target feature data includes at least one of the following: Calling a target processing model to generate and process the target feature data to obtain at least one of target audio data, target image data, target video data, and target text data that matches the target trigger event; The target feature data is provided to a first application capable of calling a target processing model, and the target feature data is generated and processed by calling the target processing model using the first application to obtain at least one of target audio data, target image data, target video data, and target text data that matches the target trigger event; Interaction input data between a target user and the target application is obtained, and target multimedia data matching the interaction input data is generated based on the interaction input data and the target feature data.
5. The method according to claim 4, wherein: Calling the target processing model to generate and process the target feature data includes at least one of the following: Determine the required processing model based on at least one of the category information of the target trigger event, the interactive input data between the target user and the target application, and the user portrait data of the target user, so as to call the determined target processing model to generate and process the target feature data; The emotional attribute of the target user represented by the target feature data is determined based on the first processing model, and the emotional attribute and the target feature data are input into the second processing model for generation processing to obtain the target multimedia data.
6. The method according to claim 1, further comprising at least one of the following: Obtaining template data corresponding to the target feature data, and generating corresponding target multimedia data based on the template data and the target feature data; The target multimedia data is output to a target area based on the running state of the target application and / or the user portrait data of the target user.
7. A method for generating audio data, comprising: In response to a target operation acting on a target game application, obtaining target application data of the target game application; generating target audio data matching the target operation based on the target application data; The target operation can trigger changes in operating parameters of the target game application.
8. The method according to claim 7, wherein: Obtaining target application data of the target game application includes at least one of the following: In response to obtaining an operation of releasing a game ultimate move in a target game application, obtaining at least one of a skill parameter of the game ultimate move, character information of a game player, and game application screen data; In response to obtaining an operation that triggers the start of a target game application, obtaining application identification data of the target application; In response to obtaining an operation of switching a game character used by a game player in the target game application from a first character to a second character, obtaining configuration data and / or associated background data of the first character and the second character; In response to obtaining interactive input with a remote game player in the target game application, at least one of configuration data, behavior data, and evaluation data of the target game player in the target game application is obtained.
9. The method according to claim 7 or 8, wherein: Generating target audio data matching the target operation based on the target application data includes at least one of the following: Calling the audio generation model to generate and process the target application data to obtain target audio data matching the target operation; Obtaining audio template data corresponding to the target application data, and generating the target audio data based on the audio template data and the target application data; An emotional attribute of a game player is determined based on the target application data, and the target audio data is generated based on the emotional attribute and the target application data.
10. An electronic device comprising a processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following: In response to a target triggering event, obtaining target feature data of a target application; Generate target multimedia data matching the target trigger event based on the target feature data; in, The target trigger event is an event that can cause the operating parameters of the target application to change.