Application processing method, related device and medium
By generating emotions curves in the application and recording application review videos, the problem of low effectiveness of information interaction in the prior art is solved, personalized information summary and review are realized, and the effective amount of information in application interaction is improved.
Patent Information
- Application Number
- CN202410100866.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-25
AI Technical Summary
Existing applications lack personalized information interaction during use, resulting in a small amount of effective information and low validity of information interaction, and users are unable to effectively summarize and review previous behaviors.
By receiving the trigger of the recording control in the application, multiple emotional expression elements of the target object during the application process are obtained, the emotion subcurve prediction model is used to generate the emotion curve, and the application video is recorded when the predetermined conditions are met, providing the application review video for users to review.
It improves the effective amount and personalization of application interaction, ensures that users can better summarize and review previous behaviors, and enhances the effectiveness of information interaction.
Smart Images

Figure CN120378683A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of application data processing, and particularly to an application processing method, a related device, and a medium. Background Art
[0002] Currently, after an object logs in to an application (such as a session application, a video browsing application, etc.), during the use of the application, there may be many process segments that the object pays attention to during the use process. These process segments are important information for summarizing experience in the future and overviewing the main actions that have been executed before. After the target object finishes using the application, no information is left for summarizing experience or summarizing the previous behavior later. Therefore, the degree of information interaction is poor and the effectiveness of information interaction is low.
[0003] Although some special applications such as games allow the object to record during use, it is often recorded in response to the trigger of the object, or there is a template stored in the game, and the content prompted in the template is recorded for all objects without distinction according to the template. These contents do not necessarily represent the information that the object pays the most attention to and needs to review later when using the game. Therefore, the effective information content of information interaction is small, the degree of personalization is low, and the effectiveness of information interaction is low. Summary of the Invention
[0004] Embodiments of the present disclosure provide an application processing method, a related device, and a medium, which can improve the effective information amount and the degree of personalization of application interaction, thereby improving the information effectiveness of application interaction.
[0005] According to one aspect of the present disclosure, there is provided an application processing method, including:
[0006] After the target application starts, receiving a trigger on a recording control on the target application page;
[0007] Obtaining a plurality of emotion manifestation elements of the target object during the process of the target application;
[0008] For each of the emotion manifestation elements, inputting the emotion manifestation element into an emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain an emotion sub-curve of the target object corresponding to the emotion manifestation element;
[0009] Generating an emotion curve of the target object based on the emotion sub-curve of the target object corresponding to the emotion manifestation element;
[0010] Recording an application video when the emotion curve of the real-time emotion meets a predetermined condition into an application review video;
[0011] In response to the triggering of the application review video playback control on the target application page after the target application ends, play the application review video.
[0012] According to an aspect of the present disclosure, there is provided an application processing device, including:
[0013] A first receiving unit, configured to receive the triggering of the recording control on the target application page after the target application starts;
[0014] An obtaining unit, configured to obtain a plurality of emotional manifestation elements of the target object during the progress of the target application;
[0015] An input unit, configured to input, for each of the emotional manifestation elements, the emotional manifestation element into the emotional sub - curve prediction model corresponding to the emotional manifestation element, and obtain the emotional sub - curve of the target object corresponding to the emotional manifestation element;
[0016] A generating unit, configured to generate an emotional curve of the target object based on the emotional sub - curves of the target object corresponding to the emotional manifestation elements;
[0017] A recording unit, configured to record the application video when the emotional curve of the real - time emotion meets a predetermined condition as an application review video;
[0018] A playback unit, configured to play the application review video in response to the triggering of the application review video playback control on the target application page after the target application ends.
[0019] Optionally, before playing the application review video in response to the triggering of the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0020] A first display unit, configured to display a duration input area in response to the triggering of the duration setting control on the target application page;
[0021] A second receiving unit, configured to receive the set duration input in the duration input area;
[0022] Wherein, the application review video has the set duration.
[0023] Optionally, before playing the application review video in response to the triggering of the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0024] A second display unit, configured to display a role duration ratio input area in response to the triggering of the role duration ratio control on the target application page;
[0025] A third receiving unit, configured to receive the first set duration ratios of each role in the target application in the application review video, which are input in the role duration ratio input area;
[0026] Wherein, the duration ratio allocated to each role in the target application in the application review video is equal to the first set duration ratio.
[0027] Optionally, before playing the application review video in response to a trigger on the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0028] A third display unit, configured to display an emotion duration ratio input area in response to a trigger on the emotion duration ratio control on the target application page;
[0029] A fourth receiving unit, configured to receive the second set duration ratios of various emotions associated with the target object in the application review video, which are input in the emotion duration ratio input area;
[0030] Wherein, the video duration ratio of various emotions associated with the target object in the application review video is equal to the second set duration ratio.
[0031] Optionally, the multiple emotion manifestation elements include at least one of the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data;
[0032] Obtaining multiple emotion manifestation elements of the target object during the process of the target application includes at least one of the following:
[0033] Obtaining the application win / loss data and the application process data of the target object during the process of the target application;
[0034] Obtaining an image of the target object through a camera, and identifying the target object's expression and the target object's body language from the image;
[0035] Collecting the target object's voice of the target object through a radio;
[0036] Obtaining the application comments from the application screen;
[0037] Obtaining the target object's heart rate of the target object from a heart rate detection device;
[0038] Obtaining the target object's screen pressing force of the target object from the application screen;
[0039] Obtain the historical performance data of the target object from the data source of the target application.
[0040] Optionally, the obtaining of the application process data of the target object during the progress of the target application includes:
[0041] Obtain the first application subprocess data from the application log of the target application;
[0042] Obtain the second application subprocess data from the application code of the target application;
[0043] Obtain the third application subprocess data from the screenshot of the target application;
[0044] Integrate the first application subprocess data, the second application subprocess data, and the third application subprocess data into the application process data.
[0045] Optionally, the obtaining of the pressing force of the target object screen of the target object from the application screen includes:
[0046] Obtain the touch behavior pattern of the target object on the application screen from the application screen;
[0047] Input the touch behavior pattern into the force prediction model to obtain the predicted pressing force of the target object screen.
[0048] Optionally, the obtaining of the historical performance data of the target object from the data source of the target application includes:
[0049] Determine the data source of the target application;
[0050] Obtain a data extraction tool;
[0051] Utilize the data extraction tool to obtain the historical performance data of the target object from the data source through the target data interface of the data source.
[0052] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0053] The generating of the emotion curve of the target object based on the emotion sub-curves of the target object corresponding to the emotion manifestation elements includes:
[0054] Obtain the first weight of each emotion manifestation element;
[0055] Utilize the first weight to perform a weighted sum on the emotion sub-curves of the target object corresponding to each emotion manifestation element to obtain the emotion curve of the target object.
[0056] Optionally, the emotion manifestation elements include the target object's voice;
[0057] For each of the emotion manifestation elements, inputting the emotion manifestation element into the emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain the emotion sub-curve of the target object corresponding to the emotion manifestation element includes:
[0058] Preprocessing the target object's voice;
[0059] Extracting emotion-related features from the preprocessed target object's voice;
[0060] Inputting the emotion-related features into the emotion sub-curve prediction model corresponding to the target object's voice to obtain the emotion sub-curve of the target object corresponding to the target object's voice.
[0061] Optionally, the emotion manifestation elements include application reviews, and the emotion sub-curve prediction model corresponding to the application reviews includes an emotion dictionary, a natural language emotion recognition model, and a speech emotion recognition model;
[0062] For each of the emotion manifestation elements, inputting the emotion manifestation element into the emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain the emotion sub-curve of the target object corresponding to the emotion manifestation element includes:
[0063] For the text reviews in the application reviews, obtaining the emotion type of the text reviews through the emotion dictionary, thereby obtaining the first emotion score curve corresponding to the text reviews;
[0064] Inputting the text reviews in the application reviews into the natural language emotion recognition model to obtain the second emotion score curve corresponding to the text reviews;
[0065] Integrating the first emotion score curve and the second emotion score curve to obtain the first integrated score curve corresponding to the text reviews;
[0066] Inputting the voice reviews in the application reviews into the speech emotion recognition model to obtain the third emotion score curve corresponding to the voice reviews;
[0067] Determining the emotion sub-curve based on the first integrated score curve and the third emotion score curve.
[0068] Optionally, each of the emotion manifestation elements includes the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data. Among them, the first weights of the target object's expression, the target object's body language, the target object's voice, the application comments, the target object's heart rate, the target object's screen pressing force, the application process data, and the target object's historical performance data decrease in sequence.
[0069] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0070] Recording the application video when the emotion curve of the real-time emotion meets a predetermined condition as an application review video includes:
[0071] On the emotion curve, intercept a segmented target emotion curve according to the review video generation rule;
[0072] Determine the object emotion corresponding to each segmented target emotion curve;
[0073] For the object emotion corresponding to the segmented target emotion curve, select a matching base image corresponding to the target object and the object emotion from the material library;
[0074] Input the emotion curve, the matching base image, and the background data of the target application into the generation model to obtain the application review video.
[0075] Optionally, the review video generation rule is: in at least a first ratio of the application review video, the emotion value of the target object meets a first condition;
[0076] The intercepting the segmented target emotion curve on the emotion curve according to the review video generation rule includes:
[0077] On the emotion curve, determine a first candidate part where the emotion value of the target object meets the first condition;
[0078] Based on the video duration of the application review video and the first ratio, determine the duration to be intercepted;
[0079] Based on the duration to be intercepted, intercept the segmented target emotion curve from the first candidate part.
[0080] Optionally, the review video generation rule is: the application review video includes multiple video parts, and in each video part, the emotion value of the target object meets a second condition associated with the video part, and the duration ratios of the multiple video parts are a predetermined duration ratio;
[0081] On the emotion curve, intercepting a target emotion curve segment according to a review video generation rule includes:
[0082] For each of the video parts, on the emotion curve, determining a second candidate part where the emotion value of the target object meets the second condition associated with the video part;
[0083] Based on the video duration of the application review video and the preset duration ratio, determining the video part duration of each of the video parts;
[0084] Based on the video part duration, intercepting an intercepted part from the second candidate parts corresponding to the video parts, and integrating the intercepted parts corresponding to the respective video parts into the target emotion curve segment.
[0085] Optionally, the review video generation rule is: the duration ratio of each character in the application review video is equal to the preset duration ratio of the character, and each of the characters includes the target character of the target object in the target application;
[0086] On the emotion curve, intercepting a target emotion curve segment according to a review video generation rule includes:
[0087] On the emotion curve, determining a third candidate part where the emotion value of the target object meets the third condition, and in the target application, for each other character, determining the time period when the other character appears, and searching for a fourth candidate part corresponding to the time period on the emotion curve;
[0088] Based on the video duration of the application review video and the preset duration ratio of the target character, determining the target character video duration occupied by the target character, and based on the video duration of the application review video and the preset duration ratio of each other character, determining the other character video duration occupied by each other character;
[0089] Based on the target character video duration, intercepting a target character curve part from the third candidate part, and based on the other character video duration, intercepting an other character curve part from the fourth candidate part;
[0090] Integrating the target character curve part and the other character curve parts into the target emotion curve segment.
[0091] Optionally, the review video generation rule is: the duration ratio of various emotions associated with the target object in the application review video is equal to the preset duration ratio corresponding to the emotion;
[0092] On the emotional curve, intercepting the target emotional curve segment according to the review video generation rule includes:
[0093] For each emotion of the target object, on the emotional curve, determine the fifth candidate part corresponding to the emotion;
[0094] Based on the video duration of the application review video and the preset duration ratio corresponding to the emotion, determine the duration corresponding to the emotion;
[0095] For each emotion, select the sub-segment corresponding to the emotion from the fifth candidate part corresponding to the emotion according to the duration corresponding to the emotion;
[0096] Integrate the sub-segments corresponding to various emotions into the target emotional curve segment.
[0097] Optionally, determining the object emotion corresponding to each target emotional curve segment includes:
[0098] Obtain the emotion value corresponding to each time point in the target emotional curve segment;
[0099] Based on the emotion value, with reference to the correspondence between the object emotion and the emotion value, determine the object emotion at each time point in the target emotional curve segment.
[0100] Optionally, the material library includes a first material library and a second material library. Among them, the first material in the first material library has an object expression map, an object label, and a first emotion label; the second material in the second material library has a character expression map, a character label, and a second emotion label;
[0101] Selecting a matching basic map corresponding to the target object and the object emotion from the material library for the object emotion corresponding to the target emotional curve segment includes:
[0102] Query the first material library. If the object label of a first material in the first material library corresponds to the target object and the first emotion label corresponds to the object emotion, use the object expression map of the first material as the matching basic map;
[0103] If the matching basic map is not found by querying the first material library, query the second material library. If the character label of a second material in the second material library corresponds to the target role of the target object in the target application and the second emotion label corresponds to the object emotion, use the character expression map of the second material as the matching basic map.
[0104] Optionally, the first material library is generated in the following manner:
[0105] Obtain materials with the object emoticon map from the Internet as the first material;
[0106] Input the first material into an emotion recognition model to obtain the first emotion label;
[0107] Input the object label map into an object recognition model to obtain the object label;
[0108] Integrate each of the first materials with the first emotion label and the object label into the first material library.
[0109] Optionally, the second material library is generated in the following manner:
[0110] Obtain materials with the character emoticon map from the Internet as the second material;
[0111] Input the second material into an emotion recognition model to obtain the second emotion label;
[0112] Obtain the character corresponding to the character label map as the character label;
[0113] Integrate each of the second materials with the second emotion label and the character label into the second material library.
[0114] Optionally, the generation model includes a copywriting part generation sub-model, a video part generation sub-model, and an audio part generation sub-model;
[0115] Inputting the emotion curve, the matching base map, and the background data of the target application into the generation model to obtain the application review video includes:
[0116] Input the emotion curve and the matching base map into the copywriting part generation sub-model to obtain the copywriting part;
[0117] Input the emotion curve, the matching base map, and the background data of the target application into the video part generation sub-model to obtain the video part;
[0118] Input the emotion curve and the matching base map into the audio part generation sub-model to obtain the audio part;
[0119] Integrate the copywriting part, the video part, and the audio part into the application review video.
[0120] Optionally, the generation model is obtained in the following manner:
[0121] Obtain a training sample set, where the training samples in the training sample set include sample applications, in - progress data of the sample object during the sample application process, and a sample application review video corresponding to the sample application;
[0122] Based on the sample application and the in - progress data, obtain the sample emotion curve of the sample object, the sample matching base map, and the sample background data of the sample application;
[0123] Input the sample emotion curve, the sample matching base map, and the sample background data into the generation model, and train the generation model based on the comparison between the model generation result and the sample application review video;
[0124] Optimize the parameters of the trained generation model;
[0125] Test the generation model after parameter optimization.
[0126] Optionally, the playback unit is specifically configured to:
[0127] When playing the application review video, make the actions of the target character corresponding to the target object scale based on the real - time emotion of the target object.
[0128] Optionally, the real - time emotion includes an emotion curve that changes over time;
[0129] The playback unit is specifically further configured to:
[0130] Input the emotion curve of the target object into the generation model, obtain the action amplitude of the target character that changes over time, and generate the actions of the target character in the application review video through skeletal animation and animation blending.
[0131] Optionally, the playback unit is specifically further configured to:
[0132] When playing the application review video, make the style and rhythm of the application review video adjusted based on the real - time emotion of the target object.
[0133] Optionally, the real - time emotion includes an emotion curve that changes over time;
[0134] The playback unit is specifically further configured to:
[0135] Input the emotion curve of the target object into the generation model, obtain the style parameters and rhythm parameters of the target character that change over time, and generate video frames of the application review video using the style parameters and the rhythm parameters, where the style of the video frame corresponds to the style parameters, and the rhythm of the video frame corresponds to the rhythm parameters.
[0136] Optionally, the playback unit is further specifically configured to:
[0137] When playing the application review video, adjust the face size and facial expression of the target character corresponding to the target object based on the real-time emotion of the target object.
[0138] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0139] The playback unit is further specifically configured to:
[0140] Input the emotion curve of the target object into the generation model to obtain the face size parameters of the target character that change over time, and generate video frames of the application review video using the emotion curve and the face size parameters, where the facial expression of the video frame corresponds to the emotion value at each time point on the emotion curve, and the face size of the video frame corresponds to the face size parameters.
[0141] Optionally, the playback unit is further specifically configured to:
[0142] When playing the application review video, add enhancement elements based on the real-time emotion of the target object.
[0143] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0144] The playback unit is further specifically configured to:
[0145] Input the emotion curve of the target object into the generation model to obtain the enhancement elements of the target character that change over time, and add the enhancement elements to the video frames of the application review video.
[0146] Optionally, the playback unit is further specifically configured to:
[0147] When playing the application review video, associate the facial image of the target character corresponding to the target object in the application review video with the inherent facial image of the target character and the facial image of the target object.
[0148] Optionally, the playback unit is further specifically configured to:
[0149] Acquire an image of the target object through a camera, and extract the facial image of the target object from the image;
[0150] Acquire the inherent facial image of the target character from the target application;
[0151] Extract the first image feature from the inherent facial image of the target character, and extract the second image feature from the facial image of the target object;
[0152] Through a generative adversarial network, fuse the first image feature and the second image feature to obtain a fused feature, and use the fused feature for 3D modeling to obtain the facial image of the target character in the application review video.
[0153] Optionally, after the target application starts and before receiving a trigger on the recording control on the target application page, the application processing device further includes:
[0154] A fourth display unit, configured to display a setting interface before the target application starts, where the setting interface includes a free screen recording start control, a camera start control, a sound collection start control, and a data collection permission control;
[0155] A fifth receiving unit, configured to receive start commands for the free screen recording start control, the camera start control, the sound collection start control, and the data collection permission control on the setting interface.
[0156] Optionally, after responding to a trigger on the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0157] A fifth display unit, configured to display a sharing control;
[0158] A sharing unit, configured to share the generated application review video to other target terminals in response to a trigger on the sharing control.
[0159] According to an aspect of the present disclosure, there is provided an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the application processing method described above is implemented.
[0160] According to an aspect of the present disclosure, there is provided a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the application processing method described above is implemented.
[0161] According to an aspect of the present disclosure, there is provided a computer program product, where the computer program product includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the application processing method described above.
[0162] In the embodiments of the present disclosure, an application review video can be automatically generated for an object after a target application ends according to the real-time emotion of the target object using the target application during the process of the target application. For example, at this time, the application review video contains points where the emotion of the target object is relatively intense during the application process (such as excitement, anger). After the target application starts, if the recording control on the target application page is triggered, then the process of the target object of the current target application during the target application will be recorded, and the application video when the emotion curve of the real-time emotion meets a predetermined condition will be recorded as the application review video. After the target application ends, if the application review video generation control on the target application page is triggered, an application review video associated with the real-time emotion of the target object using the target application during the process of the target application will be generated. Since these points of intense emotion represent the information that the object will be more interested in in the future, as well as the most important information when summarizing previous behaviors in the future, therefore, recording this information and providing it to the object in the future, or using it as the basis for other analyses, improves the effective information volume and personalization degree of application interaction, thereby improving the information effectiveness of application interaction.
[0163] Other features and advantages of the present disclosure will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0164] The drawings are used to provide a further understanding of the technical solutions of the present disclosure, and constitute a part of the specification. They are used together with the embodiments of the present disclosure to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0165] Figure 1 is a block diagram of the system to which the application processing method according to the embodiment of the present disclosure is applied;
[0166] Figures 2A - 2H is a schematic diagram of an interface in the scenario where the embodiment of the present disclosure is applied to generate and play a game review video in a target game application;
[0167] Figure 3 is a first flowchart of the application processing method according to an embodiment of the present disclosure;
[0168] Figure 4 is a second flowchart of the application processing method according to an embodiment of the present disclosure;
[0169] Figures 5A - 5B is a schematic diagram of an interface for enabling required controls in the settings interface of a target game according to an embodiment of the present disclosure;
[0170] Figures 6A - 6H Schematic diagram of an interface in a scenario where a custom method is used to generate a game review video in a target game application according to an embodiment of the present disclosure;
[0171] Figure 7 Schematic diagram of an interface in a scenario where different emotional setting duration ratios are assigned to positive emotions and negative emotions in a target game application according to an embodiment of the present disclosure;
[0172] Figure 8 is Figure 3 Schematic diagram of step 320 in [[ ]] for collecting multiple emotional manifestation elements during the process of the target application;
[0173] Figure 9 is Figure 3 Flowchart of step 320 in [[ ]] for obtaining game win / loss data and application process data in the emotional manifestation elements;
[0174] Figure 10 is Figure 3 Schematic diagram of step 320 in [[ ]] for obtaining application comments in the emotional manifestation elements from the application screen;
[0175] Figure 11 is Figure 3 Flowchart of step 320 in [[ ]] for obtaining the screen pressing intensity of the target object in the emotional manifestation elements;
[0176] Figure 12 is Figure 3 Flowchart of step 320 in [[ ]] for obtaining the historical performance data of the target object in the emotional manifestation elements;
[0177] Figure 13 Schematic diagram of an emotional curve according to an embodiment of the present disclosure;
[0178] Figure 14 is Figure 3 Flowchart of step 330 in [[ ]] for determining the real-time emotion of the target object;
[0179] Figure 15 is Figure 3 Flowchart of step 330 in [[ ]] for determining the emotional sub-curve corresponding to the target object's voice of the target object;
[0180] Figure 16 is Figure 3 Flowchart of step 340 in [[ ]] for determining the emotional sub-curve corresponding to the application comment of the target object;
[0181] Figure 17 is Figure 3 Flowchart of step 350 in [[ ]] for generating an application review video;
[0182] Figure 18 isFigure 17 The flowchart of step 1710 in
[0183] Figure 19 is Figure 17 The flowchart of step 1710 in
[0184] Figure 20 is Figure 17 The flowchart of step 1710 in
[0185] Figure 21 is Figure 17 The flowchart of step 1710 in
[0186] Figure 22 is Figure 17 The flowchart of step 1720 in
[0187] Figure 23A The schematic diagram of the first material library according to an embodiment of the present disclosure;
[0188] Figure 23B The schematic diagram of the second material library according to an embodiment of the present disclosure;
[0189] Figure 24 is Figure 17 The flowchart of step 1730 in
[0190] Figure 25 is Figure 17 The flowchart of step 1740 in
[0191] Figures 26A - 26B The schematic diagram of the interface of sharing the generated application review video to other object terminals according to an embodiment of the present disclosure;
[0192] Figure 27 is Figure 3 The first flowchart of step 360 in
[0193] Figure 28 is Figure 3 The second flowchart of step 360 in
[0194] Figure 29 is Figure 3 The third flowchart of step 360 in
[0195] Figure 30 is Figure 3 the fourth flowchart of step 360 regarding the playback application review video in
[0196] Figure 31 is Figure 3 the fifth flowchart of step 360 regarding the playback application review video in
[0197] Figure 32 is Figure 31 the flowchart of step 3110 regarding the co - association of the inherent facial image of the target character and the facial image of the target object in
[0198] Figure 33 the module diagram of the application processing device provided according to an embodiment of the present disclosure;
[0199] Figure 34 is executed according to an embodiment of the present disclosure Figure 3 the structural diagram of the object terminal of the application processing method shown in
[0200] Figure 35 is executed according to an embodiment of the present disclosure Figure 3 the structural diagram of the application server of the application processing method shown in Detailed implementation manners
[0201] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure, and are not used to limit the present disclosure.
[0202] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0203] Cloud Gaming: Also known as Gaming On Demand, it is an online game technology based on cloud computing technology. Cloud gaming technology enables lightweight devices (Thin Client) with relatively limited graphics processing and data computing capabilities to run high - quality games. In the cloud gaming scenario, the game does not run on the player's game terminal, but on the cloud server. The cloud server renders the game scene into a video - audio stream and transmits it to the player's game terminal through the network. The player's game terminal only needs to have basic streaming media playback capabilities and the ability to obtain the player's input instructions and send them to the cloud server.
[0204] Game character: It refers to an active object in a virtual scene. The active object can be a virtual character, a virtual animal, an anime character, etc. For example, the characters, animals, plants, oil drums, walls, stones, etc. displayed in the virtual scene. The virtual object can be a virtual image in the virtual scene used to represent the user. The virtual scene can include multiple virtual objects, and each virtual object has its own shape and volume in the virtual scene, occupying a part of the space in the virtual scene.
[0205] Virtual scene: It is the virtual scene displayed (or provided) when the application runs on the terminal. The virtual scene can be a simulation environment of the real world, a semi-simulated and semi-fictional virtual environment, or a purely fictional virtual environment. The virtual scene can be any one of a two-dimensional virtual scene, a 2.5D virtual scene, or a three-dimensional virtual scene, etc. The embodiments of the present application do not limit the dimension of the virtual scene. For example, the virtual scene can include the sky, land, ocean, etc. The land can include environmental elements such as deserts and cities, and the user can control the virtual object to move in the virtual scene.
[0206] Emotion: It refers to a mental state generated by humans in a specific situation. Emotions are usually interrelated with aspects such as physiology, cognition, and behavior. Emotions usually include various different states such as happiness, sadness, anger, surprise, fear, etc., and can be short-term or continuous. Emotions have an important impact on people's behavior, decision-making, and social interactions, etc.
[0207] Emotion recognition: It refers to the process of recognizing and understanding the emotions of others. Emotion recognition judges their current emotional state by observing an individual's facial expressions, body language, voice, and speech expressions, etc. Emotion recognition usually classifies and identifies an individual's emotional state by analyzing and identifying specific patterns and features in emotional expressions. For example, in human-computer interaction, emotion recognition can help the computer better understand the user's emotional state, so as to provide more personalized services; in mental health assessment, emotion recognition can help doctors better understand the patient's emotional state for more accurate diagnosis and treatment.
[0208] Content recognition: It refers to the process of recognizing, judging, classifying, etc. the information content obtained on the network. The objects of content recognition mainly include text, images, audio, video, etc. The purpose of recognition is to determine whether it is the target content required. Through content recognition, it can help the computer system better understand and process various types of data.
[0209] Artificial Intelligence Generated Content (AIGC): It refers to the process of using artificial intelligence technology to generate content. AIGC is considered a new content production method following User Generated Content (UGC) and Professional Generated Content (PGC). AI painting, AI writing, etc. all belong to the branches of AIGC. The content generated by AIGC can include various forms such as text, images, audio, and video.
[0210] Natural Language Processing: It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life. So it has a close connection with the research of linguistics, but there are also important differences. Natural language processing is mainly applied to machine translation, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, speech recognition, Chinese OCR, etc.
[0211] Generative Adversarial Network (GAN): It is a network of a deep learning model. The generative adversarial network consists of two neural network models: the generator and the discriminator. The role of the generator is to generate data samples from random noise, such as images, text, or audio, etc. Its goal is to generate fake data similar to real data samples. The role of the discriminator is to distinguish between the fake data samples generated by the generator and the real data samples, that is, to judge whether the input data is real or generated by the generator. The goal of the discriminator is to distinguish real data and fake data as accurately as possible.
[0212] System Architecture and Scenario Description Applied in Embodiments of the Present Disclosure
[0213] Figure 1 It is the system architecture diagram applied by the application processing method according to the embodiments of the present disclosure. It includes: object terminal 110, Internet 120, gateway 130, and application server 140.
[0214] The target terminal 110 is a device for the target to view game review videos. It includes various forms such as desktop computers, laptops, PDAs (Personal Digital Assistants), mobile phones, in-vehicle terminals, home theater terminals, dedicated terminals, etc. Additionally, it can be a single device or a collection of multiple devices. For example, multiple devices are connected through a local area network and share a display device for collaborative work, jointly constituting a terminal. The target terminal 110 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.
[0215] The gateway 130 is also known as an internetwork connector and protocol converter. The gateway 130 realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway 130 is a translator. At the same time, the gateway 130 can also provide filtering and security functions. Messages sent from the target terminal 110 to the application server 140 need to be sent to the corresponding application server 140 through the gateway 130. Messages sent from the application server 140 to the target terminal 110 also need to be sent to the corresponding target terminal 110 through the gateway 130.
[0216] The application server 140 refers to a computer system that can provide application review video generation services to the target terminal 110. Compared with the target terminal 110, the application server 140 has higher requirements in terms of stability, security, performance, etc. The application server 140 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc. The application server 140 can also communicate with the Internet 120 in a wired or wireless manner to exchange data.
[0217] The embodiments of the present disclosure can be applied in various scenarios, such as Figures 2A - 2H the scenario of generating a game review video with one key after the target game ends as shown, etc. The game review video at this time is the application review video generated in the game application scenario.
[0218] Currently, after the target logs in to an application (such as a session application, video browsing application, etc.) and during the use of the application, there may be many process segments that the target pays attention to during the use process. These process segments are important information for summarizing experience in the future and overviewing the main actions that have been executed before. After the target finishes using the application, no information is left for summarizing experience or summarizing the previous behavior later. Therefore, the degree of information interaction is poor and the effectiveness of information interaction is low.
[0219] Although some special applications such as games allow objects to be recorded during use, it is often recorded in response to the trigger of the object, or there are templates stored in the game, and the content prompted in the template is recorded for all objects without distinction. These contents do not necessarily represent the information that the object is most concerned about and most needs to review in the future when using the game. Therefore, the effective information content of information interaction is small, the degree of personalization is low, and the effectiveness of information interaction is low.
[0220] For the application of the game scenario, the object terminal 110 that generates the personalized game review video can be the target object terminal participating in the target game, or the opponent object terminal participating in the target game, or the friend object terminal of the target object participating in the target game. In the subsequent description of the embodiments of the present disclosure, it is described that the target object participating in the target game performs the generation operation of the game review video.
[0221] Such as Figure 2A As shown, on the target game page of the target game application in the object terminal 110, it includes operation controls in the game operation area, and also includes a recording control on the target game page ( Figure 2A the "Record" control in the lower right corner of the target game page in Figure 2A and a control for storing the game review video that the target object has already recorded (
[0222] the "Mine" control in the lower right corner of the target game page in
[0223] ). The recording control is used to start and stop recording the content during the target game process. The control for storing the video that the target object has already recorded can be used to manage and browse the game content of the player object, including the already recorded game review video, game strategy, save file, etc. The target object can start recording the content during the target game process by triggering the "Record" in the lower right corner. The target object can also view the already recorded game review video by triggering "Mine".
[0224] It should be noted that in the "My" control of the target game application, it is possible to store the game review video after editing the recorded video clip, or the original recorded video.
[0225] As Figure 2B shown, after the target game ends, a prompt of "Game over, game review video has been generated" is displayed on the target game page of the object terminal 110.
[0226] As Figure 2C shown, after the target game ends, in the "My" recorded videos, the target object selects the video for which a personalized plot event needs to be generated and makes an AI intelligent generation selection. By triggering the video for which a personalized plot event needs to be generated, the AI personalized editing page for this video is entered. The AI personalized editing page includes a "Cancel" option, a "One-click video generation" option, and a "Manual selection" option. The "One-click video generation" option and the "Manual selection" option both belong to the game review video generation controls. Among them, the "Cancel" option refers to the option to cancel the AI personalized editing of the current game review video. The "One-click video generation" option refers to the option to perform video editing on the game review video according to the default video editing configuration information. The "Manual selection" option refers to the option to perform video editing on the game review video according to the custom video editing configuration information. It should be understood that at this time, both the "One-click video generation" option and the "Manual selection" option are associated with the real-time emotions of the target object participating in the target game during the process of the target game to perform video editing on the original game review video. It can be understood that the video for which a personalized plot event needs to be generated can be a newly recorded video or a video for which a personalized plot event has already been generated to achieve secondary creation.
[0227] As Figure 2D shown, after the target object triggers the "One-click video generation" option on the AI personalized editing page, the target game application performs video editing on the video for which a personalized plot event needs to be generated according to the default video editing configuration information. At this time, a prompt of "One-click video generation in progress..." is displayed on the display screen of the object terminal 110.
[0228] As Figure 2E shown, after the editing of the game review video is completed, a prompt of "Editing completed, saved as a new video" is displayed on the display screen of the object terminal 110. At this time, the new game review video has been stored in the storage area corresponding to the "My" control of the target game application.
[0229] As Figure 2FAs shown, after the target object triggers the "My" control of the target game application, the page of "My Videos" is entered. Videos obtained by editing the game review video and the original videos of the game review video are stored on this page. If the target object wants to play the newly edited game review video, the game review video playback control of the newly edited game review video can be triggered to play the corresponding game review video.
[0230] As Figure 2G shown, after the target object triggers the game review video playback control of the newly edited game review video, as Figure 2H shown, the game review video is played on the display screen of the object terminal 110.
[0231] The embodiments of the present disclosure can be applied not only to the application processing in the above-mentioned game application scenario for generating game reviews, but also to application processing such as video viewing reviews in video application scenarios and music appreciation reviews in music application scenarios.
[0232] General description of the embodiments of the present disclosure
[0233] According to an embodiment of the present disclosure, an application processing method is provided.
[0234] The application processing method of the present application is directed to the process of generating an application review video associated with the real-time emotions of the target object participating in the target application during the progress of the target application. The target application refers to the application participated by the target object. The target object refers to the object that wants to participate in the target application. The real-time emotion refers to the specific emotion of the target object during the progress of the target application, such as excitement, anger, irritability, nervousness, complaint, pleasure, etc. The application review video refers to the video of the generated personalized plot event.
[0235] Currently, the processing method of the application review video often gives the content production tools of website-side creators. This content production tool has clear input and output ends, and is more the screencast effect of the application page, which is relatively boring. It is difficult to feel the emotional state of the application object during the competition process. Although the video production cost can be reduced, there is a lack of a certain sense of immersion and interest.
[0236] In the embodiments of the present disclosure, an application review video can be generated for an object after a target application ends according to the real-time emotions of the target object participating in the target application during the process of the target application. For example, the application review video at this time includes points where the emotions of the target object are relatively intense during the application process (such as excitement, anger). After the target application starts, if the recording control on the target application page is triggered, the process of the target object in the current target application will be recorded. After the target application ends, if the application review video generation control on the target application page is triggered, an application review video associated with the real-time emotions of the target object participating in the target application during the process of the target application will be generated. Subsequently, if the application review video playback control on the target application page is triggered, the generated application review video will be played. The embodiments of the present disclosure can fully consider the information associated with the performance of the user in the application during the process of generating the application review video, improve the effective information volume and personalization degree of the application review video, and thus improve the information effectiveness of effective interaction.
[0237] The application processing method of the embodiments of the present disclosure can be executed on the object terminal 110, can also be executed on the application server 140, or can be partly executed on the object terminal 110 and partly executed on the application server 140.
[0238] In one embodiment, as Figure 3 shown, the application processing method of the embodiments of the present disclosure includes:
[0239] Step 310: After the target application starts, receive the trigger of the recording control on the target application page;
[0240] Step 320: Obtain multiple emotion manifestation elements of the target object during the process of the target application;
[0241] Step 330: For each emotion manifestation element, input the emotion manifestation element into the emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain the emotion sub-curve of the target object corresponding to the emotion manifestation element;
[0242] Step 340: Generate an emotion curve of the target object based on the emotion sub-curve of the target object corresponding to the emotion manifestation element;
[0243] Step 350: Record the application video when the emotion curve of the real-time emotion meets a predetermined condition as an application review video;
[0244] Step 360: In response to the trigger of the application review video playback control on the target application page after the target application ends, play the application review video.
[0245] The above steps 310 - 360 are the application processing procedures of the target application under the default settings. To fully consider the information related to the performance of the user in the application, embodiments of the present disclosure may also customize some options on the settings interface of the target application to increase the effective information content and personalization degree of the application review video.
[0246] Therefore, in another embodiment, as Figure 4 shown, the application processing method of the embodiments of the present disclosure includes:
[0247] Step 3001: Before the target application starts, display a settings interface, where the settings interface includes a free screen recording activation control, a camera activation control, a sound collection activation control, and a data collection permission control;
[0248] Step 3002: On the settings interface, receive activation commands for the free screen recording activation control, the camera activation control, the sound collection activation control, and the data collection permission control;
[0249] Step 310: After the target application starts, receive the trigger of the recording control on the target application page;
[0250] Step 3101: After the target application ends, receive the trigger of the application review video generation control on the target application page;
[0251] Step 320: Obtain multiple emotional manifestation elements of the target object during the process of the target application;
[0252] Step 330: For each emotional manifestation element, input the emotional manifestation element into the emotional sub - curve prediction model corresponding to the emotional manifestation element to obtain the emotional sub - curve of the target object corresponding to the emotional manifestation element;
[0253] Step 340: Generate an emotional curve of the target object based on the emotional sub - curve of the target object corresponding to the emotional manifestation element;
[0254] Step 350: Record the application video when the emotional curve of the real - time emotion meets a predetermined condition as an application review video;
[0255] Step 3201: In response to the trigger of the custom video option control on the target application page, display a custom option bar, receive the custom video setting data configured in the custom option bar, and the application review video is associated with the configured custom video setting data;
[0256] Step 3202: Display a sharing control;
[0257] Step 3203: In response to the trigger of the sharing control, share the generated application review video to other target terminals;
[0258] Step 360: In response to the triggering of the application review video playback control on the target application page after the target application ends, play the application review video.
[0259] Note that although Figure 4 the embodiment of Figure 3 adds Step 3001, Step 3002, Step 3101, Step 3201, and Step 3202 compared to the embodiment of
[0260] The above steps 3001 - 360 will be described in detail below.
[0261] Detailed description of Step 3001 and Step 3002
[0262] In Step 3001, before the target application starts, display a settings interface, which includes a free screen recording enable control, a camera enable control, a sound collection enable control, and a data collection permission control.
[0263] The settings interface refers to an interface in the target application that allows players to make personalized settings according to their own needs before the target application starts. The embodiments of the present disclosure mainly take the application processing process of game applications as an example for illustration, and in other applications, it can be replaced according to the actual situation.
[0264] For game scene applications, as Figure 5A shown, the menu bar of the settings interface ( Figure 5A the left menu bar area) includes multiple categories of custom settings, such as basic settings, image settings, combat settings, operation settings, layout settings, sound settings, video recording settings, function / privacy settings. When the target object triggers the control of the video recording settings, it enters the Figure 5A shown video recording settings interface. At this time, the video recording settings interface includes a free screen recording enable control 510 ( Figure 5A "Free Screen Recording" in Figure 5A ), a camera enable control 520 ( Figure 5A "Automatic Video Recording" in
[0265] It should be noted that in actual applications, the format of the settings interface and the categories of custom settings included can be flexibly adjusted according to actual needs, and specific limitations are not made here.
[0266] The free screen recording activation control is a control that allows the target object to select whether to activate the application screen recording function. When the target object activates free screen recording, the target application will record the application process of the target object on the target application page in the background, including application operations, application screens, etc.
[0267] The camera activation control is a control that allows the target object to select whether to activate the camera function. If the target object activates the camera of the object terminal, the target application may use the camera of the object terminal to capture the facial expressions or body movements of the target object during the target application process.
[0268] The data collection permission control is a control that allows the target object to select whether to allow the target application to collect personal data. If the target object chooses to consent to the target application to collect and use their personal data, the target application will collect the data in the target object's application and past data for integrated collection and analysis.
[0269] The manual recording permission control is a control that allows the target object to select whether to allow the manual activation of the application recording function. If the target object activates the manual recording permission control, the target application may provide a button or shortcut key for the target object to manually activate the recording function during the application process to record their own application segments. For example, the application segments of the first 30 seconds of their own glorious moments can be manually recorded.
[0270] The XX moment permission control is a control that allows the target object to select a specific moment in the target application for recording. Such a function allows the target object to record at a specific moment in the target application, such as when achieving important achievements, completing special tasks, or having a wonderful battle.
[0271] The recording clarity selection control is a control that allows the target object to select the recording clarity of the target application. Since an increase in clarity is likely to cause the object terminal to experience lag, to improve the fluency of application operations, the target object can select different recording clarities according to their device performance and storage space, such as ultra-high definition, high definition, standard definition, or custom clarity.
[0272] The design of these controls allows the target object to make application recording settings according to their own needs and preferences, improving the interactivity and personalization of the application. At the same time, the target application needs to be able to make corresponding recording settings according to the target object's selection to ensure that the target object has more freedom and control over the application recording function.
[0273] If the target object triggers the control for sound settings in the settings page, then it enters Figure 5BThe interface of the sound settings shown. At this time, the interface of the sound settings includes a sound collection enabling control 540. The sound collection enabling control may include a total application sound enabling sub-control, an out-of-application music enabling sub-control, an in-application music enabling sub-control, and an application sound effect enabling sub-control. For applications in a game scenario, the sound collection enabling control may include a total game sound enabling sub-control, an out-of-game music enabling sub-control, an in-game music enabling sub-control, and a game sound effect enabling sub-control. In one embodiment, after enabling the corresponding sub-control, recording can be started according to the volume set by default for each sub-control, or the volume levels corresponding to at least one sub-control can be adjusted according to the needs of the target object itself, so as to start recording according to the customized sound volume.
[0274] The sound collection enabling control refers to a control that allows the target object to select whether to enable the sound collection function. If the target object enables the sound collection enabling control, the target application may record the voice of the target object in the actual scenario, the ambient sound, and the voice in the application for communication or other special functions in the target application.
[0275] The total application sound enabling sub-control refers to a control that allows the target object to select whether to enable the overall sound of the target application. If the target object does not enable the total application sound enabling sub-control, then all sound effects in the target application will be muted.
[0276] The out-of-application music enabling sub-control refers to a control that allows the target object to select whether to enable the background music of the target application when the target object is not in the in-application area. For example, the music in the application menu interface or the application loading interface of the target application.
[0277] The in-application music enabling sub-control refers to a control that allows the target object to select whether to enable the background music in the in-application area of the target application. For example, the music played in the levels or battle scenes of the target application.
[0278] The application sound effect enabling sub-control refers to a control that allows the target object to select whether to enable the sound effects in the target application. For example, the action sounds of characters and the ambient sound effects during the target application process.
[0279] The design of these sub-controls allows the target object to have more precise control over the sound effects in the target application and adjust the audio experience of the application according to their own preferences. Through these sub-controls, players can turn on or off different types of application sounds according to their own preferences, improving the personalization and user experience of the application.
[0280] Note that when obtaining the facial data, voice data, and personal data of the target object, the consent of the target object must be obtained in advance. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. When obtaining the consent of the target object, separate permission or separate consent of the target object can be obtained through pop-up windows or by jumping to the confirmation or setting page. That is to say, the activation commands of these controls are the target object's permission for the application to collect relevant information of the target object during the target application process.
[0281] In step 3002, on the setting interface, receive the activation commands for the free screen recording activation control, camera activation control, voice collection activation control, and data collection permission control.
[0282] On the setting interface, the target object triggers the free screen recording activation control, camera activation control, voice collection activation control, and data collection permission control. The object terminal of the target object receives the activation commands for these controls, enabling the target application to recognize the setting selections of the target object and configure the screen recording, camera, voice, and data collection functions of the application according to these selections. Therefore, the target application also has corresponding permission management and setting storage functions to ensure that the player's selections can be correctly applied after the application starts.
[0283] Detailed description of step 310
[0284] In step 310, after the target application starts, receive the trigger of the recording control on the target application page.
[0285] The recording control is used to start and stop the recording of the content during the target application process. This recording function allows players to record wonderful moments during the application process, or create application guides, video content, etc.
[0286] It should be noted that the target application can register an event catcher for the corresponding recording control after the target application starts. This event catcher is used to capture the trigger events of the recording control. This event catcher can capture various operations of the recording control, such as start recording, end recording, pause recording, etc.
[0287] Once the recording control is triggered, the target application can call the corresponding recording function interface to perform specific recording operations. As Figure 2A shown, if the target object clicks the "Record" button, the target application can call the recording function interface and start recording the application screen, application voice, the facial expressions and body movements of the target object, etc. according to the activation commands received for the free screen recording activation control, camera activation control, voice collection activation control, and data collection permission control.
[0288] Through the above steps, after the target application starts, it is possible to receive the trigger of the recording control on the target application page and perform corresponding recording operations according to the trigger event, thereby realizing the interaction of the application recording function.
[0289] Detailed descriptions of step 3101 and step 3201
[0290] In step 3101, after the target application ends, receive the trigger of the application review video generation control on the target application page.
[0291] The application review video generation control refers to the control used to trigger the generation of an application review video after the target application ends. The application review video generation control can be a button, icon, or other interactive element displayed on the target application page.
[0292] The target object can trigger this control to generate an application review video based on the recorded video. If the target object triggers the application review video generation control, the target application can call the video generation function interface to generate the recorded video into an application review video. This application review video can include the target object's wonderful operations, application screens, and other relevant content in the target application. The generated application review video allows the target object to review the application process or share their application achievements and wonderful moments with other players, improving the user experience. In the process of generating the application review video in the embodiments of the present disclosure, information related to the performance of the user in the application will be fully considered. At this time, the application review video is associated with the real-time emotions of the target object participating in the target application during the process of the target application, thereby increasing the effective information volume and personalization degree of the application review video.
[0293] In one embodiment, for applications in a game scenario, such as Figure 2C As shown, the game review video generation control of the embodiments of the present disclosure includes a "one-click video creation" generation control and a "manual selection" generation control. By triggering the "one-click video creation" generation control to generate a game review video, a game review video can be generated according to the default personalized video configuration information. For example, the default personalized video configuration information includes: the video duration of the default game review video is 30 seconds, the duration ratio of the target character of the target object in the target game in the game review video is 90%, and the remaining 10% of the duration in the game review video is the description of other characters and the environment background.
[0294] In another embodiment, as Figure 6A As shown, by triggering the "manual selection" generation control to generate a game review video, a game review video can be generated according to the personalized video configuration information customized by the target object. Therefore, before step 330, according to an embodiment of the present disclosure, the application processing method further includes step 3201.
[0295] In step 3201, in response to the triggering of a custom video option control on the target application page, a custom option bar is displayed, custom video setting data configured in the custom option bar is received, and the application review video is associated with the configured custom video setting data.
[0296] The custom video option control refers to a control that allows the target object to select whether to perform personalized video configuration for customization. The custom video option control at this time can be Figure 6A the "manually select" generation control in
[0297] The custom option bar refers to a structure that includes options allowing the target object to perform personalized configuration. As Figure 6B shown, the custom option bar includes a time setting control, a role duration ratio control, and an emotion duration ratio control. The controls included in the custom option bar can be flexibly adjusted according to actual needs and are not limited here. The custom video setting data refers to a data set generated after each option in the custom option bar is configured. The target object can separately trigger the time setting control, the role duration ratio control, and the emotion duration ratio control in the custom option bar to perform personalized settings for the data in each option control. In this way, after triggering the "generate with one key" control, the set data in each option control can be packaged to generate custom video setting data associated with the application review video. Therefore, the application review video is associated with the configured custom video setting data, which can effectively improve the effective information volume and personalization degree of the application review video.
[0298] In one embodiment, as Figure 4 shown, step 3201 includes:
[0299] Step 320111: In response to the triggering of a duration setting control on the target application page, a duration input area is displayed;
[0300] Step 320112: Receive the set duration input in the duration input area.
[0301] In step 320111, the duration setting control refers to a control for setting the video duration of the application review video. The duration input area refers to an area for inputting the video duration set for the application review video. This duration input area can be a text input box, a slider, or other interactive elements, etc. As Figure 6CAs shown, the duration input area 610 can be a multi - option box, including options for preset durations and a custom option. The preset durations include 15 seconds (s), 30 seconds, 1 minute (min), 2 minutes, and 5 minutes. These preset durations can be expressed in seconds, minutes, or other time units, depending on the requirements of the application. The custom option allows the target object to set the duration of the generated application review video according to their own needs. After clicking "Customize", a text input box can appear on the target application page, and the target object fills in the specific duration in this text input box. Therefore, the application review video has this set duration.
[0302] In step 320112, the target application receives the set duration input by the target object in the duration input area. After triggering the "Generate in One Click" control, it can package the set duration in the duration setting control with other setting data to generate custom video setting data associated with the application review video.
[0303] Through the above steps, the target application can allow the target object to customize the video by setting the duration of the generated application review video on the target application page. This method can effectively improve the personalization level of the generated application review video to meet the needs and preferences of different target objects.
[0304] In another embodiment, as Figure 4 shown, step 3201 further includes:
[0305] Step 320121, in response to the triggering of the role duration ratio control on the target application page, display the role duration ratio input area;
[0306] Step 320122, receive the first set duration ratio of each role in the target application in the application review video input in the role duration ratio input area.
[0307] In step 320121, the role duration ratio control refers to a control for setting the appearance duration of each role in the video for the application review video. This control usually allows users to adjust the appearance duration ratio of different roles in the video according to their own needs and preferences to achieve a more personalized viewing experience. The role duration ratio input area refers to the area for setting the appearance duration ratio of each role in the application review video in the application review video. This role duration ratio input area can be a text input box, a slider, or other interactive elements, etc. As Figure 6D shown, the role duration ratio input area 620 can be a multi - slider area, including role sliders corresponding to all roles included in the target application process.
[0308] In step 320122, the first set duration ratio refers to the ratio of the appearance duration of each role in the target application in the application review video to the total set duration. The first set duration ratio can be in the form of a percentage, a decimal, or other forms, and can be flexibly set according to the needs of the target object. That is to say, the duration ratio of each role in the target application allocated in the application review video is equal to the first set duration ratio. As Figure 6D shown, the object terminal corresponding to the target application can receive the first set duration ratio of each role in the application review video input by the target object in the role duration ratio input area. After triggering the "One - click Generation" control, multiple first set duration ratios in the role duration ratio control can be packaged with other setting data to generate custom video setting data associated with the application review video.
[0309] For an application in a game scenario, as Figure 6D shown, a target game of the target application includes 10 roles, namely role 1 to role 10. Role 1 is the target role of the target object during the target game process. At this time, the first set duration ratio of each role in this target game can be input respectively in the role duration ratio input area. For example, the first set duration ratio of role 1 is 80%, the first set duration ratio of role 2 is 5%, the first set duration ratio of role 3 is 5%, the first set duration ratio of role 4 is 5%, the first set duration ratio of role 7 is 3%, the first set duration ratio of role 8 is 2%, and the first set duration ratio of other roles is 0%. That is to say, in the generated application review video, 80% of the video frames contain role 1. It can be understood that in the role duration ratio input area, the sum of the first set duration ratios corresponding to each role is 1 or 100%.
[0310] Through the above steps, the target application can allow the target object to customize the appearance duration ratio of each role in the application review video, so that the application review video better meets the expectations of the target object and better displays the wonderful moments in the application.
[0311] In another embodiment, as Figure 4 shown, step 3201 further includes:
[0312] Step 320131, in response to the triggering of the emotion duration ratio control on the target application page, display the emotion duration ratio input area;
[0313] Step 320132, receive the second set duration ratio of various emotions associated with the target object in the application review video input in the emotion duration ratio input area.
[0314] In step 320131, the emotion duration ratio control refers to a control for setting the occupied durations of different emotions in the application review video. This control generally allows users to adjust the appearance duration ratios of different emotions in the video according to their own needs and preferences to achieve a more personalized viewing experience. Emotions include excitement, anger, irritability, nervousness, complaint, pleasure, etc. The emotion duration ratio input area refers to an area for setting the appearance duration ratios of different emotions in the application review video. This emotion duration ratio input area can be a text input box, a slider, or other interactive elements, etc. As Figure 6E shown, the emotion duration ratio input area 630 can be a multi-slider area, including emotion sliders corresponding to different emotions.
[0315] In step 320132, the second set duration ratio refers to the ratio of the appearance durations of different emotions in the target application in the application review video to the total set duration. The second set duration ratio can be in the form of a percentage, a decimal, or other forms, and can be flexibly set according to the needs of the target object. That is to say, the video duration ratio of various emotions associated with the target object in the application review video is equal to the second set duration ratio. As Figure 6E shown, the object terminal corresponding to the target application can receive the second set duration ratios of each emotion in the application review video input by the target object in the emotion duration ratio input area. After triggering the "one-key generation" control, multiple second set duration ratios in the emotion duration ratio control can be packaged with other setting data to generate custom video setting data associated with the application review video.
[0316] As Figure 6E shown, the target application has preset 10 emotions in the emotion duration ratio input area, namely emotion 1 to emotion 10, and emotion 1 to emotion 10 are used to represent different emotions. At this time, the second set duration ratios of different emotions in the target application can be input respectively in the emotion duration ratio input area. For example, the second set duration ratio of emotion 1 is 50%, the second set duration ratio of emotion 2 is 10%, the second set duration ratio of emotion 3 is 10%, the second set duration ratio of emotion 4 is 5%, the second set duration ratio of emotion 5 is 5%, the second set duration ratio of emotion 6 is 5%, the second set duration ratio of emotion 7 is 5%, the second set duration ratio of emotion 8 is 5%, the second set duration ratio of emotion 9 is 3%, and the second set duration ratio of emotion 10 is 2%. That is to say, if emotion 1 is excitement, at this time, in the generated application review video, 50% of the video frames are the operation frames of the target object in the excitement emotion during the target application process. It can be understood that in the emotion duration ratio input area, the sum of the second set duration ratios corresponding to each emotion is 1 or 100%.
[0317] Through the above steps, the target application can allow the target object to customize the proportion of the appearance duration of different emotions in the application review video, so that the application review video better meets the expectations of players and showcases different emotional experiences in the application.
[0318] For applications in game scenarios, such as Figure 6F shown, after the target object completes the corresponding settings in the duration input area, the character duration proportion input area, and the emotion duration proportion input area, the target object can click the "Generate in One Click" control on the target application page to generate a personalized application review video according to the set custom video setting data.
[0319] As Figure 6G shown, after triggering the "Generate in One Click" control on the target application page, the target application performs video editing on the video that needs to generate a personalized plot event according to the custom video setting data. At this time, a prompt of "Generating video in one click..." is displayed on the target application page.
[0320] As Figure 6H shown, after completing the editing of the application review video, a prompt of "Editing completed and saved as a new video" is displayed on the display screen of the object terminal 110. At this time, the new application review video has been stored in the storage area corresponding to the "My" control of the target application.
[0321] In another embodiment, compared with Figure 6E , the embodiment of the present disclosure can also pre-divide multiple emotions into two categories: positive emotions and negative emotions. Positive emotions are used to represent emotions in a positive emotional state. Positive emotions include joy, happiness, satisfaction, excitement, etc. Negative emotions are used to represent emotions that make the target object feel stressed. Negative emotions include anxiety, anger, sadness, fear, frustration, etc. At this time, the emotion duration proportion input area refers to the area for setting the proportion of the appearance duration of positive emotions and negative emotions in the application review video in the application review video. This emotion duration proportion input area can be a text input box, a slider, or other interactive elements, etc. As Figure 7 shown, the emotion duration proportion input area at this time can be a slider area. By adjusting the position of the sliding control in the slider area, different emotion setting duration proportions can be assigned to positive emotions and negative emotions. The emotion setting duration proportion can be in the form of a percentage, a decimal, or other forms, and can be flexibly set according to the needs of the target object. Moreover, the sum of the emotion setting duration proportions of positive emotions and negative emotions is 1 or 100%. For example, the emotion setting duration proportion of positive emotions is 80%, and the emotion setting duration proportion of negative emotions is 100% - 80% = 20%.
[0322] It should be noted that by usingFigure 7 After setting the proportion of the emotional duration shown, it can be combined with, for example, Figures 6F - 6H the following process to complete the custom video setting of the application review video.
[0323] When generating an application review video in the embodiments of the present disclosure, the target object can independently select a certain segment or set the default to automatically select the segment with the largest and most intense emotional fluctuations for priority display. The target object can also customize the duration of the generated application review video, the proportion of the appearance duration of each character, and the proportion of the appearance duration of different emotions, so that the application review video better meets the expectations of players to display different emotional experiences in the application. That is to say, in the embodiments of the present disclosure, the target object can effectively improve the effective information content and personalization degree of the application review video through custom duration control, selection of character proportions, and selection of emotion proportions.
[0324] Detailed description of steps 320 to 350
[0325] In step 320, the emotional manifestation element refers to the characteristics that can show the real-time emotion of the target object during the process of the target application. Such as Figure 8 shown, multiple emotional manifestation elements 810 include at least one of the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data. In the target application, obtaining various emotional elements shown by the target object during the process of the target application can better understand the emotional changes of the target object during the process of the target application, and can fully consider the information related to the performance of the user in the application during the process of generating the application review video, improve the effective information content and personalization degree of the application review video, and thus improve the information effectiveness of effective interaction.
[0326] In one embodiment, step 320 includes at least one of the following:
[0327] Obtain the application process data of the target object during the process of the target application;
[0328] Obtain the image of the target object through a camera, and identify the target object's expression and the target object's body language from the image;
[0329] Collect the target object's voice of the target object through a radio;
[0330] Obtain application comments from the application screen;
[0331] Obtain the target object's heart rate of the target object from a heart rate detection device;
[0332] Obtain the target object's screen pressing force of the target object from the application screen;
[0333] Obtain the historical performance data of the target object from the data source of the target application.
[0334] It should be noted that for the application in the game scenario, the emotional manifestation elements also include game win-loss data. Game win-loss data refers to the records of the target object's victory or defeat during the progress of the target game. Game win-loss data includes the win-loss data of the target game at the end of a match, and also includes the win-loss data of the battles with enemy characters during the match. Therefore, by analyzing the game performance of the target object, the real-time emotion of the target object can be analyzed.
[0335] Application process data refers to various data generated by the target object during the progress of the target application. Application process data includes application levels, time consumption, object scores, application status, character information, the technical performance of the target object during the progress of the target application (such as kill count, economy, damage amount, etc.). By analyzing the application process data, the emotional performance of the target object in the current application status can be predicted. For example, for the application in the game scenario, when the target object continuously fails or adopts an aggressive game strategy, it may indicate that the emotion of the target object is frustration or anger.
[0336] In one embodiment, as Figure 9 shown, obtain the application process data of the target object during the progress of the target application, including:
[0337] Step 910: Obtain the first application subprocess data from the application log of the target application;
[0338] Step 920: Obtain the second application subprocess data from the application code of the target application;
[0339] Step 930: Obtain the third application subprocess data from the screenshots of the target application;
[0340] Step 940: Integrate the first application subprocess data, the second application subprocess data, and the third application subprocess data into application process data.
[0341] In step 910, the target application generates an application log on the local terminal or server to record the detailed information of the target application. The first application subprocess data refers to the specific detail data of the target object during the progress of the target application obtained in real time from the application log of the target application. Therefore, by parsing the application log, the first application subprocess data recorded in real time can be extracted.
[0342] In step 920, the application code refers to the source code of the target application. The application code may include the application's logic, algorithms, interface design, etc. The second application subprocess data refers to the specific detail data of the target object during the process of the target application obtained in real time from the application code of the target application. The method of obtaining the second application subprocess data from the application code of the target application may be to obtain it through reverse engineering of the application code or by using the built-in debugging tools of the application, which is not specifically limited herein.
[0343] It should be noted that, as Figure 5A shown, if the target object triggers the opening of the data collection permission control 530 ( Figure 5A "AI data collection" in it) of the setting interface, that is, the target application is enabled based on the AI data collection capability. At this time, the second application subprocess data can be directly obtained from the source code of the target application.
[0344] In step 930, the screenshot of the target application refers to the process of capturing the screen of the target application page during the process of the target application by the target object. The third application subprocess data refers to the specific detail data of the target object during the process of the target application obtained in real time from the screenshot of the target application. Since the target object has enabled the recording function on the target application page, the target application can use the screen capture (or screen scraping) technology to obtain the screenshot image of the target application. And use image recognition and machine learning technologies to extract the third application subprocess data from the screenshot image. The process of extracting the third application subprocess data from the screenshot image in the embodiments of the present disclosure is not specifically limited.
[0345] In step 940, data integration refers to the process of merging data from different sources or formats into unified data. Among them, since the application subprocess data from different sources may use different data structures, naming specifications, or units, at this time, these application subprocess data need to be subjected to data conversion and standardization (including but not limited to renaming fields, converting data types, unit conversion, etc.) to ensure the consistency and comparability of the data. After obtaining the first application subprocess data, the second application subprocess data, and the third application subprocess data, these application subprocess data are integrated to obtain the application process data. The application process data has been described in detail in the above embodiments and will not be elaborated herein.
[0346] It should be noted that for the application in the game scenario, the game win / loss data and application process data of the target object during the process of the target application are obtained, including: obtaining the game win / loss data from the application programming interface of the target application and obtaining the first game subprocess data; obtaining the second game subprocess data from the application log of the target application; obtaining the third game subprocess data from the application code of the target application; obtaining the fourth game subprocess data from the screenshots of the target application; integrating the first game subprocess data, the second game subprocess data, the third game subprocess data, and the fourth game subprocess data into the application process data.
[0347] For the application in the game scenario, the application programming interface (API) of the target application refers to a set of programming interfaces provided by game developers or game platforms. The application programming interface of the target application is used to enable developers to interact and communicate with the target application in a programming way for developers to obtain application data. Through the application programming interface of the target application, the record data of the target object winning or losing in the target game process can be obtained in real time. Specifically, it may include the number of wins, the number of losses, the win rate, the number of kills, the economy, the amount of damage, etc. of the target object. The first game subprocess data refers to the specific detail data of the target object during the process of the target game obtained in real time from the application programming interface of the target game. This specific detail data can be the game strategy, game behavior, equipment selection, time consumption, score, character interaction, tasks or challenges, etc. of the target object during the target game process.
[0348] The application programming interface of the target application is usually based on specific programming languages and technologies, such as Java, C++, Python, etc. At the same time, the application programming interface of the target application also needs to follow specific security and stability requirements to ensure the security and confidentiality of the data in the target application and the personal information of the target object.
[0349] The second game subprocess data is the same as the above first application subprocess data, the third game subprocess data is the same as the above second application subprocess data, and the fourth game subprocess data is the same as the above third application subprocess data. It is just for the application in the game scenario here. To save space, it will not be elaborated here.
[0350] The advantage of the above embodiment is that by analyzing the technical performance of the target object during the process of the target application, such as the application process data, the accuracy of the real-time emotion judgment of the target object can be improved in the embodiments of the present disclosure, so as to increase the effective information volume of the generated application review video.
[0351] In one embodiment, the target object's expression refers to the real-time expression of the target object during the process of the target application. The expressions of the target object include smiling, angry, sad, surprised, etc. It is recorded through a camera and the expression of the target object is analyzed based on facial expression recognition technology to infer the emotional state of the target object. For example, if the target object has a smiling expression, it means the target object is in a happy mood; if the target object has a frowning expression, it means the target object is in a frustrated mood.
[0352] The target object's body language refers to the real-time body movement performance of the target object during the process of the target application. The target object's body language may include the target object's body tilt, body swaying, waving, nodding, etc. For example, if the target object suddenly shows a nodding body movement, it may indicate that the current target object is in a happy mood. It should be noted that the meanings of the target object's expression and the target object's body language may vary from person to person and may also be affected by the specific situation. Therefore, making a comprehensive judgment with other data of the target object can improve the accuracy of judging the real-time emotion of the target object, so as to increase the effective information content of the generated application review video.
[0353] As Figure 5A shown, if the target object triggers the opening of the camera activation control 520 ( Figure 5A "Automatic recording" in) of the settings interface, at this time, the target application can obtain the image of the target object during the process of the target application through the camera of the object terminal to recognize the expression of the target object and the body language of the target object from the image.
[0354] In one embodiment, the target object's voice refers to the real-time voice expression of the target object during the process of the target application. The target object's voice includes the speech content and intonation. It can be understood that if the target application includes a voice chat function, the voice and language analysis technology can be used to analyze the voice of the target object to infer the real-time emotional state of the target object. As Figure 5B shown, if the target object triggers the opening of the sound collection activation control 540 of the settings interface, at this time, the target application can collect the target object's voice during the process of the target application through the microphone of the object terminal. For example, if the pitch of the target object's voice increases or the speech becomes excited, it indicates that the current target object is in a tense or angry mood.
[0355] In one embodiment, an application comment refers to a comment during the process of a target application. Application comments include comments sent by the target object itself and comments sent by other objects during the process of the target application. Comments sent by the target object itself may indicate the current emotional state of the target object. Comments sent by other objects may also affect the emotional change of the target object. For example, if the target object makes a mistake during the target application process and other objects send accusatory comments to the target object, the target object may feel sad after seeing them. As Figure 10 shown, application comments will appear in real time in the comment area 1010 on the target application page of each object during the target application process. Therefore, take a screenshot of the target application page to obtain a screenshot image of the current target application page. And use image text recognition technology to extract image text from the comment area 1010 to extract application comments from the screenshot image. The embodiments of the present disclosure do not specifically limit the process of extracting application comments from the screenshot image of the current target application page.
[0356] In one embodiment, the heart rate of the target object refers to the heart rate of the target object during the process of the target application collected in real time by a heart rate detection device. The target application can obtain the heart rate of the target object from the heart rate detection device, which has also been pre-approved by the target object. The heart rate detection device refers to a device that is communicatively connected to the object terminal and is used to detect the heart rate of the object. The heart rate detection device can be a heart rate bracelet, a smart watch, a heart rate earphone, etc., and these heart rate detection devices are equipped with optical heart rate sensors. Therefore, the change in light absorption when blood flows through the skin through the pulse can be measured by an LED light source and a photoelectric sensor, so as to record the heart rate data of the target object in real time. For example, if the target object is nervous or excited, the heart rate of the target object will increase.
[0357] In one embodiment, the screen pressing force of the target object refers to the numerical value of the force used by the target object to press the screen of the object terminal during the process of the target application. The screen pressing force of the target object can reflect the current emotional state of the target object. It can be understood that some object terminals are configured with pressure-sensitive screens. A pressure sensor is installed on this pressure-sensitive screen, which can be used to detect the touch pressure of the target object on the application screen. For example, 3D Touch or Force Touch, etc. At this time, a device including such a screen can directly detect the force with which the target object presses the screen and transmit the obtained screen pressing force of the target object to the target application. Therefore, if the target application of the target object runs on these devices, the screen pressing force of the target object can be directly obtained by programming.
[0358] In another embodiment, the screen pressing force of the target object can be obtained through an external device. For example, an application controller or a dedicated pressure sensing detection device, etc. These devices can be connected to the object terminal used by the target object through Bluetooth or other wireless technologies. And, when actually collecting the screen pressing force of the target object, additional programming and configuration may be required, which are not specifically defined here.
[0359] In another embodiment, as Figure 11 shown, obtaining the screen pressing force of the target object from the application screen includes:
[0360] Step 1110: Obtain the touch behavior pattern of the target object on the application screen from the application screen;
[0361] Step 1120: Input the touch behavior pattern into the force prediction model to obtain the predicted screen pressing force of the target object.
[0362] In step 1110, if the object terminal is not equipped with a pressure sensing device and without the aid of an external device, the target application can also predict the screen pressing force of the target object by analyzing the touch behavior pattern of the target object on the application screen. The touch behavior pattern refers to the operation method of the target object on the application screen during the process of the target application. These operation methods and patterns may include the touch position, the movement trajectory of the fingers of the target object, the touch frequency of the fingers, etc. By analyzing these touch behavior patterns, the operation behavior of the target object on the touch screen can be understood, and thus the real-time emotion of the target object can be predicted according to this operation behavior. For example, when the target object is in a tense mood, the touch frequency of the finger touching the application screen will continuously increase.
[0363] In step 1120, the force prediction model is a machine learning model used to predict the screen pressing force of the target object according to the input data of the touch behavior pattern. The force prediction model can be constructed based on regression analysis, support vector machine, neural network, etc., which are not specifically defined here. During the model training process, the force prediction model is trained and optimized according to the constructed force sample set to enable it to accurately predict the screen pressing force of the target object.
[0364] In one embodiment, as Figure 12 shown, obtaining the historical performance data of the target object from the data source of the target application includes:
[0365] Step 1210: Determine the data source of the target application;
[0366] Step 1220: Obtain a data extraction tool;
[0367] Step 1230: Use the data extraction tool to obtain the historical performance data of the target object from the data source through the target data interface of the data source.
[0368] In step 1210, the data source of the target application refers to the structure that stores the historical data of the target application. The data sources of the target application include application servers, local storage, etc. As Figure 5A shown, if the target object triggers the enabling of the data collection permission control 530 ( Figure 5A "AI data collection" in it), that is, the target application obtains the data access permission.
[0369] In step 1220, the data extraction tool refers to a tool that extracts the historical performance data of the target object from the data source of the target application. The data extraction tool can be in the form of writing code or using existing data scraping tools, which is not specifically limited here.
[0370] In step 1230, the target data interface of the data source refers to the interface that helps the data extraction tool call the data in the data source. The target data interface of the data source can be a RESTful API, a SOAP API, etc. The historical performance data of the target object refers to the performance data of the target object during the historical process of the target application. The historical performance data of the target object includes historical data such as the emotional state, voice, application ranking, favorite heroes, proficiency level, etc. of the target object in historical application behaviors, so as to analyze the application habits and emotional experiences of the target object.
[0371] It is understandable that when calling different interfaces or libraries, corresponding authentication and authorization information may be required, and the corresponding authentication and authorization steps can be executed according to actual needs, which is not specifically limited here.
[0372] Using the data extraction tool, through the target data interface of the data source, obtain the historical performance data of the target object from the data source, including: the target application uses the data extraction tool to send an HTTP request, responds to the HTTP request, and reads the file of the data source from the target data interface of the data source by calling the API function. It is understandable that during the process of obtaining data, it is necessary to ensure the solution of possible errors and exceptions, for example, network failures, data format errors, etc. For network failures, the historical performance data of the target object can be repeatedly obtained from the data source according to a preset time interval. For data format errors, when the data is obtained, the data can be first verified for format to obtain the historical performance data of the target object with the correct format.
[0373] It should be noted that when using a data extraction tool to extract data through the target data interface of a data source, the extracted data needs to be preprocessed to obtain more effective historical performance data of the target object and improve the efficiency of generating application review videos. The preprocessing operations on the extracted data may include removing duplicate data, filling in missing values, converting data formats, etc. To ensure the secure storage and backup of the historical performance data of the target object, the preprocessed data is stored in a pre-set database or file system for subsequent analysis and use, thereby preventing data loss and leakage.
[0374] It should be noted that after obtaining the historical performance data of the target object, data analysis and visualization tools can also be used to analyze the historical performance data of the target object to calculate statistical indicators, draw charts, establish prediction models, etc., so as to accurately predict the real-time emotion of the target object during the target application process, improve the effective information content and personalization degree of the application review video, and thus improve the information effectiveness of effective interaction.
[0375] In step 330, based on the emotion manifestation elements, the real-time emotion of the target object can be determined. The emotion manifestation elements have been described in detail in the above embodiments and will not be elaborated here. The real-time emotion refers to the emotion of the target object during the progress of the target application.
[0376] In one embodiment, the real-time emotion includes an emotion curve that changes over time. The emotion curve that changes over time refers to a curve graph presented by the change of the real-time emotion state of the target object during the process of the target application over time. As Figure 13 shown, the abscissa of the emotion curve refers to the time from the start of recording to the end of recording of the target object during the target application process, and the ordinate of the emotion curve refers to the represented emotion value. The length of an emotion curve is the length of the video corresponding to the target object from the start of recording to the end of recording during the target application process. In the embodiments of the present disclosure, the emotion value corresponding to a positive emotion is marked as a value greater than 0, and the emotion value corresponding to a negative emotion is marked as a value less than 0. Different emotion values correspond to different emotion states, and the greater the absolute value of the emotion value, the stronger the emotion. For example, an emotion value of 2.6 represents an excited emotion; an emotion value of 3.6 represents a pleasant emotion; an emotion value of -1.3 represents an angry emotion; an emotion value of -3.7 represents an irritable emotion.
[0377] It should be noted that in the emotion curve, positive emotions and negative emotions can also be replaced by positive and negative. At this time, the emotion curve can be divided into 11 levels, including an emotion value of 5 (indicating very positive), an emotion value of 4 (indicating very positive), an emotion value of 3 (indicating relatively positive), an emotion value of 2 (indicating generally positive), an emotion value of 1 (indicating slightly positive), an emotion value of 0 (indicating calm), an emotion value of -1 (indicating slightly negative), an emotion value of -2 (indicating generally negative), an emotion value of -3 (indicating relatively negative), an emotion value of -4 (indicating very negative), and an emotion value of -5 (indicating extremely negative). Therefore, for the actually determined emotion value, the positive emotion value can be rounded up, and the negative emotion value can be rounded down. For example, if the emotion value is 2.6, it can correspond to the emotion of the level with an emotion value of 3; if the emotion value is 0, it can correspond to the emotion of the level with an emotion value of 0; if the emotion value is -4.8, it can correspond to the emotion of the level with an emotion value of -5.
[0378] The emotion sub-curve prediction model refers to a model that predicts the emotion sub-curve corresponding to the emotion manifestation elements of the target object during the target application process based on the input. An emotion sub-curve refers to a curve predicted corresponding to different emotion manifestation elements. In one embodiment, the emotion sub-curve prediction models corresponding to different emotion manifestation elements adopt different model structures. In this way, the prediction accuracy of the emotion sub-curves corresponding to different emotion manifestation elements can be effectively improved. However, this will also increase the model complexity of generating the emotion curve at the same time, thereby affecting the generation efficiency of the application review video.
[0379] In another embodiment, the emotion sub-curve prediction model corresponding to the emotion manifestation elements in the embodiment of the present invention can be of the same model structure. In this way, the model complexity of generating the emotion curve can be effectively simplified, thereby improving the generation efficiency of the application review video.
[0380] In one embodiment, the emotion manifestation elements include the target object's voice. As Figure 14 shown, step 330 includes:
[0381] Step 1410: Preprocess the target object's voice;
[0382] Step 1420: Extract the emotion-related features from the preprocessed target object's voice;
[0383] Step 1430: Input the emotion-related features into the emotion sub-curve prediction model corresponding to the target object's voice to obtain the emotion sub-curve of the target object corresponding to the target object's voice.
[0384] In step 1410, the target object's speech has been described in detail in the above embodiments and will not be elaborated here. When predicting the emotional sub-curve of the target object's speech, preprocessing the audio data corresponding to the target object's speech can improve the accuracy of subsequent emotion recognition. Preprocessing the target object's speech includes, but is not limited to: sampling rate conversion, noise reduction, filtering, speech segmentation, pre-emphasis, endpoint detection, speech enhancement, etc. Among them, pre-emphasis refers to highlighting the high-frequency part in the target object's speech to improve the clarity of the speech signal. Endpoint detection refers to detecting the start point and end point of the target object's speech to determine the valid part of the speech signal. Speech enhancement refers to using digital signal processing technology to enhance the target object's speech to improve the quality and intelligibility of the speech signal. Therefore, by preprocessing the target object's speech, speech data can be better prepared, and the accuracy of the subsequent speech emotion prediction task can be improved.
[0385] In step 1420, the emotion-related features refer to the features related to the emotional representation of the target object. The emotion-related features include fundamental frequency, energy, spectrum, etc.
[0386] In step 1430, input the emotion-related features into the emotion sub-curve prediction model corresponding to the target object's speech to obtain the emotion sub-curve corresponding to the target object's speech. The emotion sub-curve prediction model corresponding to the target object's speech is used to perform emotion classification on the emotion-related features to identify the emotion categories corresponding to the speech segments in the target object's speech. For example, happy, sad, angry, etc. In this way, determine the emotion values according to the emotion categories corresponding to the speech segments, and obtain the emotion sub-curve corresponding to the target object's speech based on these emotion values and the positions of the speech segments in the target object's speech.
[0387] It should be noted that after obtaining the emotion sub-curve corresponding to the target object's speech, this emotion sub-curve can be output to the application developer. In this way, the application developer can make corresponding responses according to the different emotional state changes of the target object during the target application process. For example, provide appropriate application feedback or adjust the application difficulty.
[0388] It should be noted that there are still certain errors in the emotion recognition technology for emotion-related features. Therefore, in actual applications, other emotion manifestation elements need to be combined, such as application status, target object behavior, etc., to comprehensively judge the emotion state of the target object.
[0389] In one embodiment, the emotion manifestation elements include application reviews. The emotion sub-curve prediction model corresponding to the application reviews includes an emotion dictionary, a natural language emotion recognition model, and a speech emotion recognition model. As Figure 15 shown, step 330 includes:
[0390] Step 1510: For the text comments in the application reviews, obtain the sentiment type of the text comments through a sentiment dictionary, so as to obtain the first emotion score curve corresponding to the text comments.
[0391] Step 1520: Input the text comments in the application reviews into a natural language emotion recognition model to obtain the second emotion score curve corresponding to the text comments.
[0392] Step 1530: Integrate the first emotion score curve and the second emotion score curve to obtain the first integrated score curve corresponding to the text comments.
[0393] Step 1540: Input the voice comments in the application reviews into a voice emotion recognition model to obtain the third emotion score curve corresponding to the voice comments.
[0394] Step 1550: Determine the emotion sub-curve based on the first integrated score curve and the third emotion score curve.
[0395] In step 1510, the application reviews include text comments and voice comments. The text comments can be the text comments sent by the target object or the text comments sent by other objects jointly participating in the target application. The sentiment dictionary refers to a set containing words and corresponding sentiment categories. Here, emotion and sentiment have the same meaning. The first emotion score curve is a curve that changes over time constructed according to the emotion values corresponding to the sentiment types of the text comments through the matching method of the sentiment dictionary. The sentiment dictionary can be used in sentiment analysis tasks to classify sentiment by matching the words appearing in the text comments with the words in the sentiment dictionary. Therefore, by classifying the sentiment of the text comments in the application reviews through the sentiment dictionary, the sentiment types corresponding to multiple words in the text comments can be obtained. At this time, according to the emotion values corresponding to the sentiment types of multiple words, the first emotion score curve corresponding to the text comments is obtained.
[0396] It should be noted that the application reviews can be data obtained from the competition platform or social media of the target application. When obtaining the application reviews, preprocessing such as noise removal, word segmentation, and stop word removal can be performed on the text comments in the application reviews to obtain the final text comments.
[0397] It should be noted that when performing sentiment analysis on the preprocessed text comments using the sentiment dictionary, the text can be divided into different sentiment categories such as positive (equivalent to the above-mentioned positive emotion), neutral (equivalent to the above-mentioned calm emotion), and negative (equivalent to the above-mentioned negative emotion).
[0398] In step 1520, the natural language emotion recognition model: The natural language emotion recognition model refers to a model used to analyze the emotion categories in review texts. The natural speech emotion recognition model is a model trained using natural language processing techniques. By learning a large amount of labeled emotion data, the natural speech emotion recognition model can automatically classify review texts into different emotion categories, such as happy, sad, angry, etc. The second emotion score curve refers to a curve that changes over time constructed based on the emotion values corresponding to the emotion types of text reviews through the recognition method of the natural language emotion recognition model. Natural language processing techniques are used to understand and analyze human language, including spoken and written language. Such techniques can be used to identify emotion information in speech and text. For example, by analyzing the semantics and emotional color of review texts, the emotion state of the object corresponding to the review text can be judged.
[0399] In step 1530, the first integrated score curve refers to the curve obtained by integrating the first emotion score curve and the second emotion score curve. In one embodiment, the way to integrate the first emotion score curve and the second emotion score curve can be: calculate the mean value of the emotion values at the corresponding time positions of the first emotion score curve and the second emotion score curve, and construct the first integrated score curve based on the calculated emotion values.
[0400] In another embodiment, the way to integrate the first emotion score curve and the second emotion score curve can also be: set corresponding second weights for the first emotion score curve and the second emotion score curve; use the second weights to perform weighted summation on the first emotion score curve and the second emotion score curve to obtain the first integrated score curve.
[0401] In step 1540, the voice review can be a review of the voice emitted by the target object or a review of the voice emitted by other objects jointly participating in the target application. The voice emotion recognition model is a model that uses speech recognition technology to analyze the emotion categories expressed in voice reviews. The third emotion score curve refers to a curve that changes over time constructed based on the emotion values corresponding to the emotion types of voice reviews through the recognition method of the voice emotion recognition model. The voice emotion recognition model extracts the acoustic features of voice reviews and applies machine learning or deep learning algorithms for emotion classification. The voice emotion recognition model can analyze features such as the pitch, volume, and speech rate of voice reviews to identify the emotion state of the speaker.
[0402] It should be noted that if the voice emotion recognition model is a model constructed based on deep learning algorithms, after training the deep learning model, the accuracy of emotion recognition for voice reviews can be effectively improved.
[0403] In step 1550, the emotion sub-curve corresponding to the application comment refers to the curve integrated based on the first integrated sub-curve and the third emotion sub-curve. In one embodiment, the way to integrate the first integrated sub-curve and the third emotion sub-curve can be: calculating the average value of the emotion values at the corresponding time positions of the first integrated sub-curve and the third emotion sub-curve, and constructing the emotion sub-curve corresponding to the application comment according to the calculated emotion values.
[0404] In another embodiment, the way to integrate the first integrated sub-curve and the third emotion sub-curve can also be: setting corresponding third weights for the first integrated sub-curve and the third emotion sub-curve respectively; using the third weights to perform weighted sum on the first integrated sub-curve and the third emotion sub-curve to obtain the emotion sub-curve corresponding to the application comment.
[0405] The advantages of the above steps 1510 - 1550 are that by performing emotion classification on text comments through an emotion dictionary and a natural language emotion recognition model, and performing emotion recognition on voice comments through a voice emotion recognition model, the accuracy of emotion recognition corresponding to application comments can be achieved, so as to accurately predict the real-time emotion of the target object during the target application process, improve the effective information volume and personalization degree of the application review video, and thus improve the information effectiveness of effective interaction.
[0406] In one embodiment, as Figure 16 shown, step 340 includes:
[0407] Step 1610, obtaining the first weight of each emotion manifestation element;
[0408] Step 1620, using the first weight to perform weighted sum on the emotion sub-curves corresponding to each emotion manifestation element of the target object to obtain the emotion curve of the target object.
[0409] In step 1610, the first weight of each emotion manifestation element is obtained. The first weight refers to the numerical value of the influence degree of each emotion manifestation element on the generated emotion curve. The larger the first weight, the higher the influence degree of the emotion sub-curve corresponding to the corresponding emotion manifestation element on the final emotion curve.
[0410] In one embodiment, each emotion manifestation element includes the target object's expression, the target object's body language, the target object's voice, the application comment, the target object's heart rate, the target object's screen pressing force, the application process data, and the target object's historical performance data. Among them, the first weights of the target object's expression, the target object's body language, the target object's voice, the application comment, the target object's heart rate, the target object's screen pressing force, the application process data, and the target object's historical performance data decrease in turn. And, the sum of the first weights of the included emotion manifestation elements is 1.
[0411] In another embodiment, for the target application in the game scenario, each emotion manifestation element includes game win / loss data, the target object's expression, the target object's body language, the target object's voice, game comments, the target object's heart rate, the target object's screen pressing force, game process data, and the target object's historical performance data. Among them, the first weights of the game win / loss data, the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data decrease in sequence. And the sum of the first weights of each included emotion manifestation element is 1. For example, the first weight of the game win / loss data is 0.4, the first weight of the target object's expression is 0.18, the first weight of the target object's body language is 0.12, the first weight of the target object's voice is 0.12, the first weight of the application comments is 0.08, the first weight of the target object's heart rate is 0.04, the first weight of the target object's screen pressing force is 0.03, the first weight of the application process data is 0.02, and the first weight of the target object's historical performance data is 0.01.
[0412] It should be noted that the above emotion manifestation elements have been described in detail in the above embodiments and will not be elaborated here. Since the game win / loss data may have a more direct and significant impact on the result, it is given a higher weight. Some emotion manifestation elements may better represent the true state and situation of the target object, so they can also be given a higher weight to better reflect the true situation of the target object. While the target object's historical performance data has a smaller impact on the target object during the current target application process, it is given a smaller weight.
[0413] In the multi-element fusion of the embodiments of the present disclosure, more attention is paid to direct and representative elements. Therefore, the decreasing first weights can reflect the influence degree of different elements on the final result, which is more in line with the actual situation and effectively improves the prediction accuracy of the emotion curve.
[0414] In step 1620, using the first weights, perform weighted summation on the emotion sub-curves of the target object corresponding to each emotion manifestation element to obtain the emotion curve of the target object, including: performing a multiplication operation on the emotion sub-curves of the target object corresponding to each emotion manifestation element and the corresponding first weights to obtain the weighted emotion sub-curves of each emotion manifestation element; performing a summation calculation on the emotion values at the corresponding time in the weighted emotion sub-curves of each emotion manifestation element to obtain the emotion curve of the target object.
[0415] The advantages of the above embodiments are that the use of the first weight can highlight the influence degree of certain emotion manifestation elements on the emotion of the target object. By adjusting the first weights of different emotion manifestation elements, the emotion curve can be flexibly adjusted to better adapt to specific requirements and scenarios. Through the weighted sum method, the emotion sub-curves of different emotion manifestation elements can be synthesized, so as to more comprehensively reflect the emotion state of the target object, making the final emotion curve more representative, and further better reflecting the change of the emotion state of the target object.
[0416] In another embodiment, the real-time emotion may further include emotion text that changes over time. The emotion text that changes over time refers to the combination of emotion words corresponding to the real-time emotion state of the target object during the process of the target application over time. Moreover, preset delimiters may be included between different emotion words in the emotion text to accurately analyze the change of the emotion state of the target object during the target application process. For example, an emotion text may be "calm|calm|slightly positive|comparatively positive|very negative|slightly negative|calm", so the change of the emotion state of the target object can be analyzed according to this emotion text.
[0417] It should be noted that the emotion text at this time can also be obtained by classifying emotions in the same way as steps 1410 - step 1430, steps 1510 - step 1550, and steps 1610 - step 1620 above, except that the generated emotion sub-curves and emotion curves are both replaced in the form of text, which will not be elaborated here.
[0418] In step 350, based on the real-time emotion, the application video when the emotion curve meets a predetermined condition is recorded as an application review video. The predetermined condition refers to a pre-set emotion condition for generating the application review video. The present application can generate the application review video through a generation model. The generation model refers to a model that can be associated with the real-time emotion of the target object participating in the target application and generate a video of a personalized plot event.
[0419] In one embodiment, the real-time emotion includes an emotion curve that changes over time, as Figure 17 shown. At this time, step 350 includes:
[0420] Step 1710: On the emotion curve, intercept the target emotion curve segments according to the review video generation rule;
[0421] Step 1720: Determine the object emotion corresponding to each target emotion curve segment;
[0422] Step 1730: For the object emotion corresponding to the target emotion curve segment, select a matching base image corresponding to the target object and the object emotion from the material library;
[0423] Step 1740: Input the emotion curve, the matching base map, and the background data of the target application into the generation model to obtain the application review video.
[0424] In step 1710, the specific meaning and acquisition method of the emotion curve have been described in detail in the above embodiments and will not be elaborated here. The review video generation rule refers to the rule for segmenting the emotion curve specified in advance according to the emotional changes of the target object during the process of the target application. The review video generation rule can be a rule determined based on factors such as the emotion value of the target object, the duration ratio of different emotions, and the duration ratio of different roles. The target emotion curve segmentation refers to the segment selected for generating the application review video after the emotion curve of the target object is segmented according to the review video generation rule. Multiple target emotion curve segments can be intercepted according to the review video generation rule.
[0425] In one embodiment, the review video generation rule is that in at least the first ratio in the application review video, the emotion value of the target object meets the first condition. The first ratio can be the proportion of the duration occupied by the target object in each emotional state set in advance. The first condition refers to the emotion value condition that the emotion value of the target object needs to meet. For example, the duration of the positive emotion peak 5 > the duration of the negative emotion peak -5 > the duration of the positive small wave peak 3 > the duration of the negative small wave peak -3 > the duration of calm 0, and the first ratio represents: the duration of the positive emotion peak 5: the duration of the negative emotion peak -5: the duration of the positive small wave peak 3: the duration of the negative small wave peak -3: the duration of calm 0 = 4:2:1.8:1.2:1. At this time, the first condition means that the emotion value of the target object meets any one of the positive emotion peak 5, negative emotion peak -5, positive small wave peak 3, negative small wave peak -3, and calm 0.
[0426] At this time, as Figure 18 shown, step 1710 includes:
[0427] Step 1810: On the emotion curve, determine the first candidate part where the emotion value of the target object meets the first condition;
[0428] Step 1820: Based on the video duration of the application review video and the first ratio, determine the duration to be intercepted;
[0429] Step 1830: Based on the duration to be intercepted, intercept the target emotion curve segment from the first candidate part.
[0430] In step 1810, on the emotion curve, the emotion values at each time point are judged to determine the first candidate part where the emotion value of the target object meets the first condition. The emotion value of the target object meeting the first condition means that the emotion value of the target object is equal to an emotion value included in the first condition within the error value. The first candidate part refers to the curve segment determined by the emotion value of the target object that meets the first condition. For example, if the first condition indicates that the emotion value of the target object meets any one of the positive emotion peak value 5, negative emotion peak value -5, positive small wave peak value 3, negative small wave peak value -3, and calm value 0, then 5 first candidate parts that meet the first condition will be obtained.
[0431] It should be noted that to avoid the situation where the target object lacks the corresponding emotion value in the first condition, the embodiments of the present disclosure will recognize the emotion values of the target object that meet the first condition within the error value as the first candidate parts that meet the first condition. The error value here can be 0.2, 0.3, etc., and can be flexibly set according to actual requirements.
[0432] In step 1820, the video duration of the application review video refers to the total duration of the generated video. When the application review video is generated by default, the video duration is the default set duration. When the application review video is generated customarily, the video duration is the duration customarily set by the target object. Since the duration of the determined first candidate part may be a segment longer than the video duration, at this time, it is necessary to determine the duration to be intercepted based on the video duration of the application review video and the first ratio. For example, the video duration is 100 seconds, and the first ratio indicates that the duration of the positive emotion peak value 5: the duration of the negative emotion peak value -5: the duration of the positive small wave peak value 3: the duration of the negative small wave peak value -3: the duration of the calm value 0 = 4:2:1.8:1.2:1. In this way, the duration to be intercepted for the positive emotion peak value 5 can be determined to be 40 seconds, the duration to be intercepted for the negative emotion peak value -5 is 20 seconds, the duration to be intercepted for the positive small wave peak value 3 is 18 seconds, the duration to be intercepted for the negative small wave peak value -3 is 12 seconds, and the duration to be intercepted for the calm value 0 is 10 seconds.
[0433] In step 1830, after determining the duration to be intercepted and the first candidate part, the target emotion curve segment is intercepted from the first candidate part according to the duration to be intercepted. The segment length of the target emotion curve segment is the length of the duration to be intercepted.
[0434] The advantages of the above steps 1810 - 1830 are that according to the first condition, the first candidate part that meets the first condition in the emotion value of the target object can be selected. According to the total duration of the application review video and the first ratio, the duration of the segment to be intercepted is determined for cutting the emotion curve. In this way, the information related to the user's performance in the application can be fully considered, improving the effective information amount and personalization degree of the application review video.
[0435] In another embodiment, the rule for generating the review video is as follows: The application review video includes multiple video segments, and the emotional value of the target object in each video segment meets a second condition associated with the video segment. The duration ratio of the multiple video segments is a predetermined duration ratio. To reflect different emotional combinations in each video segment, embodiments of the present disclosure may set corresponding second conditions for each video segment. The second condition refers to the condition of the emotional value that the target object needs to meet. The predetermined duration ratio refers to the ratio of the duration of each video segment to the video duration of the application review video. For example, the application review video includes two video segments. In the first video segment, the emotional value of the target object meets the second condition associated with the video segment: the emotional value of the target object meets any one of the positive emotional peak value of 5 and the positive small wave peak value of 3. In the second video segment, the emotional value of the target object meets the second condition associated with the video segment: the emotional value of the target object meets any one of the negative emotional peak value of -5, the negative small wave peak value of -3, and the calm value of 0. The predetermined duration ratio of the first video segment is 20%, and the predetermined duration ratio of the second video segment is 80%.
[0436] At this time, as Figure 19 shown, step 1710 includes:
[0437] Step 1910: For each video segment, on the emotional curve, determine a second candidate segment where the emotional value of the target object meets the second condition associated with the video segment;
[0438] Step 1920: Based on the video duration of the application review video and the predetermined duration ratio, determine the video segment duration of each video segment;
[0439] Step 1930: Based on the video segment duration, intercept an intercepted part from the second candidate segment corresponding to the video segment, and integrate the intercepted parts corresponding to each video segment into a target emotional curve segment.
[0440] In step 1910, on the emotional curve, for each video segment, judge the emotional value at each time point of each video segment to determine a second candidate segment where the emotional value of the target object meets the second condition associated with the video segment. The emotional value of the target object meeting the second condition associated with the video segment means that the emotional value of the target object is equal to one of the emotional values included in the second condition associated with the video segment within the error value. The second candidate segment refers to the curve segment determined by the emotional value of the target object that meets the second condition associated with the video segment. For example, the application review video includes two video segments, namely the first video segment and the second video segment. In this way, a second candidate segment where the emotional value of the target object meets the second condition associated with the first video segment, and a second candidate segment where the emotional value of the target object meets the second condition associated with the second video segment can be determined.
[0441] In step 1920, since the duration of the determined second candidate part may be a segment longer than the video duration, at this time, it is necessary to determine the video part duration of each video part based on the video duration of the application review video and the preset duration ratio. For example, the application review video includes two video parts, namely the first video part and the second video part. The video duration is 90 seconds, the preset duration ratio of the first video part is 20%, and the preset duration ratio of the second video part is 80%. In this way, the video part duration of the first video part can be determined as 90×20% = 18 seconds, and the video part duration of the second video part is 90×80% = 72 seconds.
[0442] In step 1930, after determining the video part duration of each video part and the second candidate part corresponding to the video part, the intercepted part is obtained by intercepting from the second candidate part corresponding to the video part according to the video part duration of each video part. The duration of the intercepted part corresponding to each video part is the video part duration of each video part. Then, according to the time of each video part on the emotion curve, the intercepted parts corresponding to each video part are integrated into the target emotion curve segment.
[0443] The advantages of the above steps 1910 - 1930 are that according to the second condition associated with the video part, the second candidate part that satisfies the second condition of each video part can be selected from the emotion values of the target object. According to the video duration of the application review video and the preset duration ratio, the video part duration of each video part to be intercepted is determined for the cutting of the emotion curve. In this way, the information associated with the user's performance in the application can be fully considered, improving the effective information amount and personalization degree of the application review video.
[0444] In another embodiment, for the application of the game scenario, the rule for generating the review video is: the duration ratio of each character in the application review video is equal to the preset duration ratio of the character, and each character includes the target character of the target object in the target application. The preset duration ratio of the character refers to the ratio of the duration of each character in the application review video to the video duration of the application review video. As Figure 6D shown, there are 10 characters in a target application, namely character 1 to character 10. Character 1 is the target character of the target object during the target application process. For example, the preset duration ratio of character 1 is 80%, the preset duration ratio of character 2 is 2%, the preset duration ratio of character 3 is 5%, the preset duration ratio of character 4 is 5%, the preset duration ratio of character 5 is 0%, the preset duration ratio of character 6 is 1%, the preset duration ratio of character 7 is 5%, the preset duration ratio of character 8 is 2%, the preset duration ratio of character 9 is 0%, and the preset duration ratio of character 10 is 0%.
[0445] At this time, asFigure 20 As shown, step 1710 includes:
[0446] Step 2010: On the emotion curve, determine the third candidate part where the emotion value of the target object meets the third condition, and in the target application, for each other role, determine the time period when the other role appears, and search for the fourth candidate part corresponding to the time period on the emotion curve;
[0447] Step 2020: Based on the video duration of the application review video and the preset duration ratio of the target role, determine the target role video duration occupied by the target role, and based on the video duration of the application review video and the preset duration ratio of each other role, determine the other role video duration occupied by each other role;
[0448] Step 2030: Based on the target role video duration, intercept the target role curve part from the third candidate part, and based on the other role video duration, intercept the other role curve part from the fourth candidate part;
[0449] Step 2040: Integrate the target role curve part and each other role curve part into the target emotion curve segment.
[0450] In step 2010, the third condition refers to the condition of the emotion value that the emotion value of the target object needs to meet. Similar to the meaning of the above second condition, it will not be elaborated here. On the emotion curve, judge the emotion value of the target object at each time point to determine the third candidate part where the emotion value of the target object meets the third condition. That the emotion value of the target object meets the third condition means that the emotion value of the target object is equal to an emotion value included in the third condition within the error value. The third candidate part refers to the curve segment determined by the emotion value of the target object that meets the third condition. The fourth candidate part refers to the curve segment determined by searching for the emotion value corresponding to the time period on the emotion curve according to the time period when each other role appears in the recorded video of the target application.
[0451] In step 2020, after determining the third candidate part and the fourth candidate part, based on the video duration of the application review video and the preset duration ratio of the target role, the target role video duration occupied by the target role is determined. At the same time, based on the video duration of the application review video and the preset duration ratio of each other role, the other role video duration occupied by each other role is determined. In the application of the game scenario, for example, a target application includes 5 roles, namely role 1 to role 5. Role 1 is the target role of the target object during the target application process. At this time, if the preset duration ratio of role 1 is 80% and the video duration is 90 seconds, then the target role video duration occupied by the target role is 90×80% = 72 seconds. If the preset duration ratio of role 2 is 10%, the preset duration ratio of role 3 is 2%, the preset duration ratio of role 4 is 4%, and the preset duration ratio of role 5 is 4%. At this time, the role video duration occupied by role 2 is 90×10% = 9 seconds, the role video duration occupied by role 3 is 90×2% = 1.8 seconds, the role video duration occupied by role 2 is 90×4% = 3.6 seconds, and the role video duration occupied by role 2 is 90×4% = 3.6 seconds.
[0452] In step 2030, after determining the target role video duration and the third candidate part corresponding to the target role, the target role curve part is intercepted from the third candidate part according to the target role video duration. After determining the other role video duration occupied by the other role and the fourth candidate part of the other role, the other role curve part is intercepted from the corresponding fourth candidate part of the other role according to the other role video duration occupied by each other role.
[0453] In step 2040, according to the time of the emotion curve, the target role curve part and each other role curve part are integrated into the target emotion curve segment.
[0454] The advantages of the above steps 2010 - 2040 are that in the application of the game scenario, according to the third condition corresponding to the target role, the third candidate part with the emotion value of the target object meeting the third condition can be selected. In this target application, for each other role, determine the time period when the other role appears, and find the fourth candidate part corresponding to the time period on the emotion curve. At this time, based on the determined target role video duration of the target role, the target role curve part can be intercepted from the third candidate part, and based on the other role video duration of each other role, the other role curve part can be intercepted from the corresponding fourth candidate part. In this way, the information related to the performance of the user in the application can be fully considered, and the effective information volume and personalization degree of the application review video can be improved.
[0455] It should be noted that the embodiments of the present disclosure can adjust the duration ratios of different roles to select and switch perspectives of different roles, so as to produce multiple application review videos with different roles as the protagonists. In this way, the emotional states of multiple roles can be felt, more, better, and more interesting content can be produced, and the willingness of users to share, spread, and even consume can be enhanced. At the same time, the way of generating application review videos in this way can greatly reduce the cost of second-creation videos for application reviews, not only improving the efficiency of user creation and editing, but also increasing users' interest in content production to a certain extent and improving the effectiveness of information in effective interactions.
[0456] In another embodiment, the rule for generating a review video is: the duration ratio of various emotions associated with the target object in the application review video is equal to the preset duration ratio corresponding to the emotion. The preset duration ratio corresponding to an emotion refers to the ratio of the duration of each emotion in the application review video to the video duration of the application review video. As Figure 6E shown, the target application can preset 10 emotions in the emotion duration ratio input area, namely Emotion 1 to Emotion 10, and Emotion 1 to Emotion 10 are used to represent different emotions. At this time, the preset duration ratios corresponding to different emotions can be respectively input in the emotion duration ratio input area. For example, the preset duration ratio corresponding to Emotion 1 is 50%, the preset duration ratio corresponding to Emotion 2 is 10%, the preset duration ratio corresponding to Emotion 3 is 10%, the preset duration ratio corresponding to Emotion 4 is 5%, the preset duration ratio corresponding to Emotion 5 is 5%, the preset duration ratio corresponding to Emotion 6 is 5%, the preset duration ratio corresponding to Emotion 7 is 5%, the preset duration ratio corresponding to Emotion 8 is 5%, the preset duration ratio corresponding to Emotion 9 is 3%, and the preset duration ratio corresponding to Emotion 10 is 2%. That is to say, if Emotion 1 is excitement, then in the generated application review video, 50% of the video frames are the operation frames of the target object in the excitement emotion during the target application process.
[0457] At this time, as Figure 21 shown, step 1710 includes:
[0458] Step 2110: For each emotion of the target object, determine the fifth candidate part corresponding to the emotion on the emotion curve;
[0459] Step 2120: Based on the video duration of the application review video and the preset duration ratio corresponding to the emotion, determine the duration corresponding to the emotion;
[0460] Step 2130: For each emotion, select the sub-segment corresponding to the emotion from the fifth candidate part corresponding to the emotion according to the duration corresponding to the emotion;
[0461] Step 2140: Integrate the sub-segments corresponding to various emotions into the target emotion curve segments.
[0462] In step 2110, the fifth candidate part refers to a curve segment determined according to the emotion values of each emotion on the emotion curve.
[0463] In step 2120, based on the video duration of the application review video and the preset duration ratio corresponding to the emotion, the duration corresponding to the emotion is determined. For example, if the preset duration ratio corresponding to emotion 1 is 50% and the video duration is 90 seconds, then the duration corresponding to emotion 1 is 90×50% = 45 seconds.
[0464] In step 2130, after determining each emotion and the duration corresponding to each emotion, a sub-segment corresponding to the emotion is selected from the fifth candidate part corresponding to each emotion.
[0465] In step 2140, the target emotion curve segmentation refers to a curve obtained by integrating the segments corresponding to various emotions.
[0466] The advantages of the above steps 2110 - 2140 are that, according to each emotion of the target object and the preset duration ratio corresponding to the emotion, a sub-segment corresponding to the emotion is selected from the fifth candidate part corresponding to the emotion, and the sub-segments corresponding to various emotions are integrated into the target emotion curve segmentation. In this way, information related to the performance of users in the application can be fully considered, improving the effective information content and personalization degree of the application review video.
[0467] In step 1720, the object emotions corresponding to each target emotion curve segmentation are determined. A target emotion curve segmentation may include multiple object emotions, and one time point in the target emotion curve segmentation corresponds to one object emotion. The object emotion can be a specific emotion on the emotion curve as Figure 13 shown, such as excited, angry, irritable, nervous, happy, etc.
[0468] In one embodiment, as Figure 22 shown, step 1720 includes:
[0469] Step 2210, obtaining the emotion value corresponding to each time point in the target emotion curve segmentation;
[0470] Step 2220, based on the emotion value, referring to the correspondence between the object emotion and the emotion value, determining the object emotion corresponding to each time point in the target emotion curve segmentation.
[0471] In step 2210, the visual structure of the target emotion curve segmentation is the same as Figure 13The structures of the shown emotion curves are the same. According to the numerical values of the horizontal and vertical coordinates where the target emotion curve segment is located, the emotion value corresponding to each time point in the target emotion curve segment can be determined. For example, if the target emotion curve segment is the time interval from the start time to the first calm value, the emotion values corresponding to each time point in the target emotion curve segment can be obtained according to the preset time point interval as: 0, 2.7, 0.
[0472] In step 2220, since there is a corresponding relationship between each object emotion and the emotion value, in order to ensure that the corresponding relationship between the intermediate emotion value and the object emotion is unique, each object emotion can correspond to an emotion value interval. For example, the emotion value of the calm emotion is 0, the emotion value interval of the slightly nervous emotion is (0, 0.4], the emotion value interval of the excited emotion is (1.8, 2.8], and the emotion value interval of the pleasant emotion is (2.8, 3.8]. Therefore, based on the emotion value and referring to the corresponding relationship between the object emotion and the emotion value, the object emotion corresponding to each time point in the target emotion curve segment can be determined. For example, if the target emotion curve segment is the time interval from the start time to the first calm value, the emotion values corresponding to each time point in the target emotion curve segment can be obtained according to the preset time point interval as: 0, 2.7, 0. At this time, the object emotions corresponding to each time point in the target emotion curve segment are "calm, excited, calm".
[0473] The advantages of the above embodiments are that by determining the emotion value corresponding to each time point in the target emotion curve segment, detailed emotion value information can be provided, so as to better understand the emotional changes at each time point in the video, thereby providing richer information for subsequent emotional analysis and understanding.
[0474] In step 1730, for the object emotion corresponding to the target emotion curve segment, a matching base image corresponding to the target object and the object emotion is selected from the material library.
[0475] The material library refers to a database or resource library containing various multimedia materials. The materials contained in the material library include pictures, videos, audios, etc. The content in the material library is sorted and classified so that the target application can conveniently search and obtain the required materials. The matching base image refers to the material image selected from the material library that matches the target object and the object emotion. These base images can be used as the visual presentation of the emotional analysis results and are used in video production to show the emotional changes in the video. For example, if the object emotion corresponding to the target emotion curve segment is "happy", then an image that can express a happy emotion can be selected from the material library, which may be a smiling face, a happy scene or other pictures that can convey a happy emotion. In this way, in the subsequent application review video, these images can be combined with the video content to more vividly display the emotional changes in the video.
[0476] The material library includes a first material library and a second material library. Among them, the first materials in the first material library have object expression maps, object tags, and first emotion tags; the second materials in the second material library have character expression maps, character tags, and second emotion tags.
[0477] In other words, the first material library is a material library related to real objects. As Figure 23A shown, one object expression map in the first materials corresponds to one object tag and one first emotion tag. However, one object tag can correspond to multiple object expression maps, and one first emotion tag can correspond to the object expression maps corresponding to multiple object tags. For example, one object tag represents object A, and the first emotion tags include "happy", "angry", and "calm". At this time, the object expression maps of object A when "happy" can include expression map P1 and expression map P2, the object expression maps of object A when "angry" can include expression map P3, and the object expression maps of object A when "calm" can include expression map P4 and expression map P5. It can be understood that the object expression maps may contain some texts expressing emotions.
[0478] The second material library is a material library related to the characters in the target application. As Figure 23B shown, one character expression map in the second materials corresponds to one character tag and one second emotion tag. However, one character tag can correspond to multiple character expression maps, and one second emotion tag can correspond to the character expression maps corresponding to multiple character tags. For example, one object tag represents character B, and the second emotion tags include "happy", "angry", and "irritable". At this time, the character expression maps of character B when "happy" can include expression map P6 and expression map P7, the character expression maps of character B when "angry" can include expression map P8, and the character expression maps of character B when "irritable" can include expression map P9 and expression map P10. It can be understood that the character expression maps may contain some texts expressing emotions.
[0479] In one embodiment, the first material library is generated in the following manner:
[0480] Obtain materials with object expression maps from the Internet as the first materials;
[0481] Input the first materials into an emotion recognition model to obtain the first emotion tags;
[0482] Input the object label maps into an object recognition model to obtain the object tags;
[0483] Integrate each of the first materials with the first emotion tags and object tags into the first material library.
[0484] The first material refers to the material obtained from the Internet with the target emoticon pictures. The process of obtaining the material with the target emoticon pictures from the Internet can be achieved by obtaining videos, audios, texts, images, etc. from the Internet. Moreover, the material with the target emoticon pictures obtained from the Internet may come from various websites, social media platforms, forums, blogs, etc. It should be noted that these materials with the target emoticon pictures are all data obtained after the permission of websites, platforms, forums, etc.
[0485] The emotion recognition model refers to a model that can recognize the emotion label corresponding to the first material. The emotion recognition model can be a model constructed based on machine learning, natural language processing technology, etc. The first emotion label refers to the label determined after classifying the emotion of the first material. For example, the first emotion label can include "happy", "sad", "angry", "surprised", etc., to assign corresponding labels to each first material. If the first material is a video or an image, an emotion recognition model based on computer vision technology can be used to recognize the facial expressions in the video or image to determine the first emotion label corresponding to the first material. If the first material contains text, the emotional color of the text content can be analyzed through an emotion recognition model based on natural language processing technology to determine the first emotion label corresponding to the first material. If the first material contains material audio, an emotion recognition model based on speech recognition technology can be used to analyze the speech emotion in the audio to determine the first emotion label corresponding to the first material.
[0486] The object recognition model refers to a model that can recognize the object label corresponding to the object label picture. The object recognition model can be a model such as a convolutional neural network (CNN), a recurrent neural network (RNN), etc. By analyzing the object label picture, its corresponding object label can be determined. The object label refers to the object to which the corresponding object label picture belongs. The object label can include object A, object B, object C, etc. At this time, after determining the first emotion label and the object label corresponding to the first material, each first material with the first emotion label and the object label can be integrated into the first material library.
[0487] The advantages of the above embodiments are that the first materials obtained after labeling are integrated into a material library related to the correspondence between emotions and objects, so that materials can be searched according to the emotion category and the corresponding object category, improving the interest and personalization of reviewing application videos, and helping users better feel the dramatic performance in the application process. At the same time, it also reserves a database for the subsequent generation of AIGC personalized plot competitions.
[0488] In one embodiment, the second material library is generated in the following manner:
[0489] Obtain the material with the character emoticon pictures from the Internet as the second material;
[0490] Input the second material into the emotion recognition model to obtain the second emotion label;
[0491] Obtain the character corresponding to the character label map as the character label;
[0492] Integrate each second material with the second emotion label and the character label into a second material library.
[0493] The second material refers to the material with character expression maps obtained from the Internet. The method of obtaining the first material from the Internet is the same as that of obtaining the second material, which will not be elaborated here. At this time, the second material can be videos, audios, text contents, etc. related to each application character. At this time, the second material can also be the content videos recreated by netizens.
[0494] The emotion recognition model at this time is the same as the emotion recognition model used for the first material, which will not be elaborated here. The second emotion label refers to the label determined after classifying the emotions of the second material. For example, the second material label can be "happy", "sad", "angry", "surprised", etc. to assign corresponding labels to each second material.
[0495] Different methods can be used to obtain the character corresponding to the character label map. If the second material is a video or an image, computer vision technology can be used to identify the character image in the video or image to determine the character label corresponding to the character label map. If the second material contains text, the character name mentioned in the text content can be analyzed through natural language processing technology to determine the character label corresponding to the character label map. If the second material contains material audio, speech recognition technology can be used to analyze the character name in the audio to determine the character label corresponding to the character label map. For example, the character label can be set to categories such as "tank", "warrior", "mage", etc., or specific characters such as character A, character B, character C, etc., which are not specifically limited here.
[0496] At this time, after determining each second material with the second emotion label and the character label, the second material can be integrated into a second material library.
[0497] The advantages of the above embodiments are that the second materials obtained after labeling are integrated into a material library corresponding to emotions and characters, so that materials can be searched according to emotion categories and corresponding character categories, improving the interest and personalization of the application review video, and helping users better feel the dramatization performance in the application process. At the same time, it also reserves a database for the subsequent generation of AIGC personalized plot events.
[0498] In one embodiment, as Figure 24 shown, step 1730 includes:
[0499] Step 2410: Query the first material library. If the object label of a first material in the first material library corresponds to the target object and the first emotion label corresponds to the object emotion, use the object expression map of the first material as the matching base map.
[0500] Step 2420: If no matching base map is found by querying the first material library, query the second material library. If the role label of a second material in the second material library corresponds to the target role of the target object in the target application and the second emotion label corresponds to the object emotion, use the role expression map of the second material as the matching base map.
[0501] In step 2410 of some embodiments, for the object emotion corresponding to each time point in the segmented target emotion curve, first query the first material library. At this time, the object label corresponding to the obtained matching base map can correspond to the target object, and the first emotion label corresponding to the matching base map can correspond to the object emotion corresponding to this time point in the target emotion curve segment.
[0502] In step 2420 of some embodiments, if no matching base map is found by querying the first material library, query the second material. At this time, the role label corresponding to the obtained matching base map can correspond to the target role of the target object in the target application, and the second emotion label corresponding to the matching base map can correspond to the object emotion corresponding to this time point in the target emotion curve segment.
[0503] It should be noted that after finding a matching base map by querying the first material library, the second material library can also be continued to be queried to determine the matching base map related to the role. In this way, according to multiple matching base maps, the generated application review video can be enriched, and the interest and personalization degree of the application review video can be improved.
[0504] The advantages of the above embodiments are that by searching for the matching base map that matches the object emotion corresponding to the target emotion curve segment in the first material library and the second material library, the interest and personalization degree of the application review video can be improved, and it helps users better feel the dramatic performance in the application process.
[0505] In step 1740, input the emotion curve, the matching base map, and the background data of the target application into the generation model to obtain the application review video. Specifically, each target emotion curve segment, the matching base map, and the background data of the target application can be input into the generation model to obtain the application review video. The application curve and the matching base map have been described in detail in the above embodiments. To save space, they will not be elaborated here.
[0506] The background data of the target application refers to the background materials and data used in the development of the target application. The background data of the target application includes application scenarios and maps, characters, props, application audio, etc. Application scenario and map data refer to various scenarios and maps in the target application, including terrain, buildings, roads, water bodies, etc. These data usually exist in the form of images or 3D models and are used to construct the environment of the application world. Character data refers to various character models, actions, appearances, etc. in the target application. Character data may include character models, animations, texture maps, etc., and is used to present the character images in the application. Prop data refers to various prop, equipment, item, etc. data in the target application. Prop data may include prop models, textures, special effects, etc., and is used to present various interactive items in the application. Application audio data refers to audio materials such as background music and sound effects used in the target application. Generally speaking, the background data of the target application refers to various materials and data used in the development of the target application for constructing the application environment. These data are an important part of the application development process and can directly affect the visual effects and user experience of the application.
[0507] As Figure 25 shown, the generation model 2510 includes a copywriting part generation sub-model, a video part generation sub-model, and an audio part generation sub-model.
[0508] In one embodiment, as Figure 25 shown, step 1740 includes:
[0509] Input the emotion curve and the matching base map into the copywriting part generation sub-model to obtain the copywriting part;
[0510] Input the emotion curve, the matching base map, and the background data of the target application into the video part generation sub-model to obtain the video part;
[0511] Input the emotion curve and the matching base map into the audio part generation sub-model to obtain the audio part;
[0512] Integrate the copywriting part, the video part, and the audio part into an application review video.
[0513] It can be understood that the copywriting part generation sub-model refers to a model used to generate the text part in the application review video. The copywriting part generation sub-model can be a model constructed based on natural language processing (NLP) technology. Therefore, input the emotion curve and the matching base map into the copywriting part generation sub-model to obtain the copywriting part. At this time, the copywriting part is an application review copy that is generated according to the emotions of the target object and the application situation and that conforms to both the object's personality and the application style.
[0514] The video part generation sub-model refers to a model used to generate the video part in the application review video. The video part generation sub-model can be a model constructed based on computer vision and video processing technologies. The video part generation sub-model can be used to generate virtual reality scenes, video special effects, video editing, etc. Therefore, it is necessary to input the emotion curve, the matching base map, and the background data of the target application into the video part generation sub-model to obtain the video part. At this time, the video part can display the virtual application background generated according to the background data of the target application, and the corresponding matching base map can be displayed in the video according to the changes in emotions in the emotion curve, thereby improving the interest and personalization of the application review video and helping users better feel the dramatic performance during the application process.
[0515] The audio part generation sub-model refers to a model used to generate the audio part in the application review video. The audio part generation sub-model can be a model constructed based on audio processing technologies. The audio part generation sub-model can be used for audio synthesis, audio enhancement, speech recognition, etc. Since some of the matching base maps may contain corresponding audio, the emotion curve and the matching base map can be input into the audio part generation sub-model to obtain the audio part, so as to integrate the audio in the matching base map into the application review video.
[0516] After determining the text part, video part, and audio part corresponding to the application review video, the corresponding text part, video part, and audio part are integrated according to the time points of the emotion curve, so that a complete application review video can be obtained. That is to say, according to the video style, the position, font, color, and size of the text part can be automatically matched and adjusted to ensure that the text part can be clearly displayed on the video. At the same time, the volume and mixing of the audio part are intelligently adjusted to ensure that the audio can perfectly match the video. Finally, the video is exported to obtain the final application review video. Therefore, the application review video obtained in this way can make the video content more vivid, rich, and attractive, improving the user experience.
[0517] It should be noted that in the embodiments of the present disclosure, when generating the audio part, the corresponding timbre expression can be matched in combination with the user's emotional state to enhance the expressiveness and appeal of the audio. And the text-to-speech (TTS) technology, such as deep learning or speech models, is used to generate natural and fluent speech.
[0518] In one embodiment, the generation model is obtained in the following manner:
[0519] Obtain a training sample set, where the training samples in the training sample set include sample applications, in-progress data of the sample object during the sample application process, and sample application review videos corresponding to the sample applications;
[0520] Based on the sample application and the ongoing data, obtain the sample emotion curve of the sample object, the sample matching base map, and the sample background data of the sample application;
[0521] Input the sample emotion curve, the sample matching base map, and the sample background data into the generation model, and train the generation model based on the comparison between the model generation result and the sample application review video;
[0522] Perform parameter tuning on the trained generation model;
[0523] Test the generation model after parameter tuning.
[0524] The training sample set refers to a set used to store multiple training samples. The sample application is the same as the target application in the above embodiment, the sample object is the same as the target object in the above embodiment, and the ongoing data of the sample object during the sample application process is the same as the data corresponding to the above emotion manifestation elements. Here, they are all used as training samples. For the sake of brevity, they will not be elaborated here. The sample application review video serves as a sample label for the training samples and is used to test the generation performance of the generation model.
[0525] Based on the sample application and the ongoing data, the sample emotion curve of the sample object is the same as the emotion curve in the above embodiment, the sample matching base map is the same as the matching base map in the above embodiment, and the sample background data of the sample application is the same as the background data of the target application in the above embodiment. The corresponding acquisition steps can be referred to. Here, they are all used as training samples. For the sake of brevity, they will not be elaborated here.
[0526] The process of inputting the sample emotion curve, the sample matching base map, and the sample background data into the generation model to obtain the model generation result is the same as step 1740 above. Here, they are used as training samples. For the sake of brevity, they will not be elaborated here. Based on the comparison between the model generation result and the sample application review video, a similarity calculation can be performed between the model generation result and the sample application review video to obtain a video similarity value. And the generation model is trained according to the video similarity value, that is, the parameters of the generation model are adjusted. Specifically, a first threshold can be preset. The first threshold refers to the minimum value that the video similarity needs to meet. When the video similarity value is greater than or equal to the first threshold, the training process ends. When the video similarity value is less than the first threshold, the parameters of the fusion model are adjusted until the video similarity value is greater than or equal to the first threshold. Therefore, by performing a similarity calculation between the model generation result and the sample application review video to train the generation model, the generation accuracy of the trained model is improved.
[0527] After completing the training of the generation model, some optimizations need to be carried out on the trained generation model to improve the performance of the model. The process of parameter tuning includes adjusting the learning rate, regularization coefficient, network structure, optimizer type, etc., to improve the generalization ability and convergence speed of the model. Therefore, in the process of parameter tuning, more complex models can be used, or some techniques such as data augmentation and transfer learning can be used, which are not specifically limited here. By performing parameter tuning on the trained generation model, the generation model can be made more accurate, fluent and meet expectations when generating content such as text, images, and audio.
[0528] After performing parameter tuning on the trained generation model, some test data need to be used to test the parameter-tuned generation model to test the performance of the model. The test data should be different from the training data but have similar characteristics. Therefore, through testing, it can be evaluated whether the generation model can achieve the expected goal. Testing the parameter-tuned generation model can include two aspects: quantitative evaluation and qualitative evaluation. Quantitative evaluation can quantify the performance of the model by calculating the metrics of the model, such as accuracy, recall, BLEU score, the quality score of the generated images, etc. Qualitative evaluation is to judge the generation effect of the model by manually evaluating the content generated by the model, such as the smoothness of the text, the realism of the image, the quality of the audio, etc.
[0529] Performing parameter tuning and testing on the trained generation model is a process of continuously improving and optimizing the generation model. The generation model obtained in this way can effectively improve the quality and realism of the generated application review videos, thereby increasing the effective information content and personalization degree of the application review videos, and improving the information effectiveness of effective interactions.
[0530] Detailed description of step 3202 and step 3203
[0531] In practical applications, the generated application review videos can be selected to be shared to other social media, or can also be selected to be saved to the local device.
[0532] In step 3202, sharing controls are displayed. The sharing controls refer to the controls that the target object can choose whether to share the generated application review video to other object terminals. The sharing controls can be a sharing button or option on the target application page, allowing the target object to click on it to share the application review video.
[0533] As Figure 26A shown, after the video editing is completed and a new application review video is generated, the current target application page contains a "Cancel" option and an "Upload and Share" option (equivalent to the sharing control).
[0534] In step 3203, in response to the triggering of the sharing control, the generated application review video is shared to other target terminals. The specific generation process of the application review video has been described in detail in the above embodiments and will not be elaborated here.
[0535] It can be understood that if the target object triggers the sharing control, the target application can call the sharing function of the target or call a specific sharing SDK to share the generated application review video to other target terminals, such as social media platforms, emails, messages, etc. As Figure 26B shown, at this time, the target application can share the generated application review video to other target terminals through application platforms such as A number, A application, A circle, A letter, XX camp, A video, etc.
[0536] It should be noted that, therefore, the embodiments of the present disclosure can share and spread the application review video through major platforms by triggering the sharing control, which can effectively improve the user activity of the target application, form a virtuous relay and self-created content production, and thus improve the information effectiveness of effective interaction.
[0537] Detailed description of step 360
[0538] In step 360, in response to the triggering of the application review video playback control on the target application page after the target application ends, the application review video is played.
[0539] The application review video playback control refers to the control on the target application page that executes the playback function for the generated application review video. As Figure 2F shown, after the target object triggers the "My" control of the target application, the "My Video" interface is entered. The videos after editing the application review video and the original videos of the application review video are stored on this interface. As Figure 2G shown, after the target object triggers the application review video playback control for the newly edited application review video, as Figure 2H shown, the application review video is played on the display screen of the target terminal 110.
[0540] In one embodiment, as Figure 27 shown, playing the application review video in step 360 includes:
[0541] Step 2710, when playing the application review video, making the actions of the target role corresponding to the target object scale based on the real-time emotion of the target object.
[0542] The scaling of the actions of the target character based on the real-time emotions of the target object means that when implementing a personalized event plot, the player's emotional performance will be integrated, and the amplitude of the hitting actions will be changed in terms of strength. At this time, the action amplitude of the target character includes the amplitude of limb dynamics (the height of the hand lift or the height of the jump) and speed (the punching speed, the moving speed, etc.).
[0543] Specifically, the real-time emotions include an emotion curve that changes over time. The emotion curve has been described in detail in the above embodiments and will not be elaborated here. At this time, step 2710 includes: inputting the emotion curve of the target object into the generation model to obtain the action amplitude of the target character that changes over time, and generating the actions of the target character in the application review video through skeletal animation and animation blending.
[0544] The action amplitude refers to the amplitude or range when the target character performs actions in the target application. In the application of the game scenario, the action amplitude includes the moving speed of the character, the attack range, the skill release range, the jumping height, etc. Skeletal animation is a common animation technique. It creates animations by moving and rotating a series of "bones" of a character connected together. A character can be decomposed into two parts: bones and skin. The skeletal animation technique can edit and control the actions of the bones and then map the skin texture to the bones to achieve the animation effect of the character. The skeletal animation technique is usually used to create the basic actions of the target character, such as walking, running, jumping, attacking, etc. Then, the amplitude of these actions is adjusted according to the output of the generation model, such as changing the walking speed, the jumping height, the attacking strength, etc.
[0545] Animation blending is a technique that blends multiple animation effects together. Animation blending is used to achieve smooth transitions and complex action performances of characters, and can smoothly transition between different animations or play multiple animations simultaneously. Through the animation blending technique, application developers can blend different action effects (such as walking, running, jumping) with a certain weight to achieve a more natural and smooth transition effect. For example, when the target character transitions from walking to running, the walking and running animations can be smoothly blended. Or, when the target character jumps and attacks, the jumping and attacking animations can be played simultaneously.
[0546] Animation blending technology allows characters to switch and combine different actions more flexibly in an application, enhancing the expressiveness of the characters and the application experience. In the application of game scenarios, for example, the special effects presented by the hitting actions of the target character during the game, such as the fixed special effects of consecutive powerful moves. During the playback of the application review video, when the target character in the application hits during the application process and the emotion of the target object is recognized as extremely angry, it can be analyzed that the target object may have been killed by the opponent in the previous round. At this time, combined with the emotional performance of the user, the hitting special effects of the target character can be strengthened, and the player's body movements will be enhanced and amplified, such as becoming an additional bonus for the consecutive powerful move special effects, so as to enhance the player's feeling of emotional expression and improve the personalization degree of the application review video.
[0547] It should be noted that the embodiments of the present disclosure also introduce reinforcement learning to let the generation model try mistakes in a set environment. That is to say, the behavior of the model can be adjusted according to the results (rewards or punishments) of reinforcement learning. Specifically, the set environment includes the emotion value of the player and the movement amplitude of the target character. Then, the embodiments of the present disclosure can let the generation model try mistakes in this set environment and maximize a certain reward, such as the satisfaction of the player, by adjusting the movement amplitude of the target character.
[0548] It should be noted that the embodiments of the present disclosure can input the emotion value of the target object in the target application into the application AI data to add and rewrite the fixed and patterned postures and dynamics of the original target character. At this time, the higher the absolute value of the emotion value, the more obvious the change in the movement amplitude of the target character over time. That is to say, the amplitude of body dynamics (the height of the hand raised or the height of the jump), speed (the speed of punching, the moving speed, etc.) will increase by a corresponding percentage compared with before. In the application of game scenarios, for example, during the killing process, the current emotion value of the target object is 5 (indicating very positive), that is, dealing with a very excited emotion. Then, at this time, the punching height during the killing can be increased to 200% compared with the original application effect, and the punching speed is also increased to twice the speed. Therefore, the embodiments of the present disclosure can adjust different effects on the actions of the target character according to different emotion values, so as to more accurately express the real-time emotion value of the target object according to the action special effects of the target character in the application review video.
[0549] In one embodiment, as Figure 28 shown, playing the application review video in step 360 includes:
[0550] Step 2810: When playing the application review video, adjust the style and rhythm of the application review video based on the target object's real-time emotion. That is to say, aspects such as the content, music, and shot transitions of the application review video can be dynamically adjusted according to the emotional changes of the target object to better match and reflect the emotional state of the target object.
[0551] Before adjusting the style and rhythm of the application review video according to the target object's real-time emotion, it is necessary to identify the emotion of the target object during the target application process. In the application of a game scenario, for example, if the target object fails frequently or encounters difficulties during the target application process, it may show a sense of frustration or depression. If the target object achieves a major victory during the target application process, it may show joy or excitement. At this time, if the corresponding style and rhythm can be flexibly switched according to the object's emotion, it can better increase the emotional resonance of the viewing object and improve the viewing experience.
[0552] In one embodiment, specifically, Step 2810 includes: inputting the emotion curve of the target object into a generation model to obtain the style parameters and rhythm parameters of the target character that change over time, and using the style parameters and rhythm parameters to generate video frames of the application review video, where the style of the video frame corresponds to the style parameters and the rhythm of the video frame corresponds to the rhythm parameters.
[0553] Since the emotion of the target object changes over time, the style parameters and rhythm parameters in the video frames of the generated application review video will also change accordingly. The style parameters and rhythm parameters are usually used to describe the characteristics and presentation methods of music, animation, or other audio-visual elements in the target application. The style parameters refer to the style characteristics and attributes of music, visual effects, or other elements in the target application. For example, for music, the style parameters can include characteristics such as melody, rhythm, harmony, and instrument selection. For the application screen, the style parameters can include characteristics such as color tone, light and shadow effects, and special effects. Application developers can create different audio-visual styles by adjusting and controlling these style parameters to meet the needs of application design and the expectations of players. The rhythm parameters refer to the rhythm characteristics and presentation methods of music, animation, or other elements in the target application. For music, the rhythm parameters can include characteristics such as beats, tempo, drum beats, and note lengths. For application animation, the rhythm parameters can include characteristics such as action speed, shot transition frequency, and picture change rhythm. By adjusting and controlling the rhythm parameters, application developers can make the music and animation in the application more in line with the application rhythm and enhance the rhythm and atmosphere of the application.
[0554] Input the emotional curve of the target object into the generation model to obtain the style parameters and rhythm parameters of the target character that change over time, and use the style parameters and rhythm parameters to generate video frames for the application review video. For example, if the target object shows a happy or excited emotion, a fast-paced and bright-color style and rhythm can be selected as the style parameters and rhythm parameters of the current video frame. If the target object shows a sense of frustration or depression, a slow-paced and dull-color style and rhythm can be selected as the style parameters and rhythm parameters of the current video frame.
[0555] Therefore, in the application review video, the embodiments of the present disclosure can adjust the style and rhythm of the video frames according to the real-time emotion of the target object, thereby affecting the emotion of the viewing object, better increasing the emotional resonance of the viewing object, and enhancing the viewing experience.
[0556] In one embodiment, as Figure 29 shown, playing the application review video in step 360 includes:
[0557] Step 2910, when playing the application review video, make the face size and facial expression of the target character corresponding to the target object be adjusted based on the real-time emotion of the target object.
[0558] The face size refers to the size and proportion of the face of the target character in the application review video. By adjusting the face size, the image expressiveness of the character can be enhanced, thereby highlighting the real-time emotion of the target object.
[0559] The facial expression refers to the emotion and expression shown by the target character in the application review video through facial muscle movement and expression changes. In the target application, the facial expression can be realized through facial animation technology. The performance of the facial expression can make the target character more vivid and emotional, and can better express the inner world and emotional state of the target character, thereby directly expressing the real-time emotion of the target object.
[0560] In one embodiment, specifically, step 2910 includes: inputting the emotional curve of the target object into the generation model to obtain the face size parameters of the target character that change over time, and using the emotional curve and the face size parameters to generate video frames for the application review video, where the facial expression of the video frame corresponds to the emotion value at each time point on the emotional curve, and the face size of the video frame corresponds to the face size parameters. In this way, by integrating the generated video frames, the application review video can be obtained.
[0561] In the emotional curve of the target object, each object emotion of the target object can be converted into an emotion value. Combining Figure 13As shown, the emotion curve can be divided into 11 levels, including an emotion value of 5 (indicating very positive), an emotion value of 4 (indicating very positive), an emotion value of 3 (indicating relatively positive), an emotion value of 2 (indicating generally positive), an emotion value of 1 (indicating slightly positive), an emotion value of 0 (indicating calm), an emotion value of -1 (indicating slightly negative), an emotion value of -2 (indicating generally negative), an emotion value of -3 (indicating relatively negative), an emotion value of -4 (indicating very negative), and an emotion value of -5 (indicating extremely negative). Moreover, for the actually determined emotion value, the positive emotion value can be rounded up, and the negative emotion value can be rounded down.
[0562] In the embodiments of the present disclosure, based on the absolute value of different emotion values, the facial expression of the target role corresponding to the target object can be presented and relatively magnified by a corresponding multiple. For example, when the absolute value is |5| and the facial size parameter is magnified by 2 times, a very positive and happy or extremely negative state is presented to give a close-up of the facial expression; when the absolute value is |4| and the facial size parameter is magnified by 1.8 times, a very positive and happy or very negative state is presented to give a close-up of the facial expression; when the absolute value is |3| and the facial size parameter is magnified by 1.6 times, a relatively positive and happy or relatively negative state is presented to give a close-up of the facial expression; when the absolute value is |2| and the facial size parameter is magnified by 1.4 times, a generally positive and happy or generally negative state is presented to give a close-up of the facial expression; when the absolute value is |1| and the facial size parameter is magnified by 1.2 times, a slightly positive and happy or slightly negative state is presented to give a close-up of the facial expression.
[0563] For example, if the emotion value of the target object is very high, such as an emotion value of 5 indicating very positive, at this time, the target role can be made to show a happy facial expression. At the same time, the facial expression close-up of the target role corresponding to the target object is magnified by 2 times. That is to say, in the video frame at this time, the facial expression of the target object is happy, and the facial size parameter is magnified by 2 times. If the emotion value of the target object is very low, such as presenting -5 indicating extremely negative, the target role can be made to show a sad facial expression. At the same time, the facial expression of the target role corresponding to the target object is also magnified by 2 times. That is to say, in the video frame at this time, the facial expression of the target object is sad, and the facial size parameter is magnified by 2 times. It can be seen from this that although the two facial size parameters are both magnified by 2 times, the emotions shown are completely opposite.
[0564] In the embodiments of the present disclosure, by adjusting the facial size and facial expression of the target role to adjust the video frames for generating the application review video, the role can be made more vivid and expressive in the application or animation, and can better express the real-time emotion of the target object. Therefore, during the process of generating the application review video, information related to the performance of the user in the application can be fully considered, the effective information volume and personalization degree of the application review video can be improved, and thus the information effectiveness of effective interaction can be improved.
[0565] In one embodiment, as Figure 30 shown, playing the application review video in step 330 includes:
[0566] Step 3010, when playing the application review video, adding enhancement elements based on the real-time emotion of the target object.
[0567] The enhancement elements of real-time emotion refer to the elements that enhance the user's emotional expression in the target application. The enhancement elements may include, but are not limited to: sound effects and music, visual effects, character action special effects, etc. Sound effects and music refer to enhancing the emotional expression in the application or virtual experience through pre-designed sound effects and music. For example, adding suspense music in a tense plot, or adding sad music in a sad scene to enhance the emotional experience of the target object. Visual effects include dynamic light and shadow effects, special effects, picture tones, etc. That is, guiding the emotional experience of the object through visual changes. For example, adding rapid shot transitions and dynamic picture effects in a tense plot to increase the sense of tension. Character action special effects include slow motion, magnification, flashing, etc. to highlight the user's emotional expression. In addition, comments or voices of objects in the target application can also be added to enhance the expressiveness and appeal of the video.
[0568] In one embodiment, specifically, step 3010 includes:
[0569] Inputting the emotion curve of the target object into a generation model to obtain the enhancement elements of the target character that change over time, and adding the enhancement elements to the video frames of the application review video.
[0570] Since the emotion of the target character changes over time, at this time, the enhancement elements corresponding to each video frame in the application review video may be different. Therefore, it is necessary to add the enhancement elements to the video frames of the application review video. In this way, the generated video frames are integrated to obtain the application review video.
[0571] By adding the enhancement elements to the video frames of the application review video in the embodiments of the present disclosure, the user's emotional expression can be enhanced, so as to better express the real-time emotion of the target object in the application review video. Therefore, during the process of generating the application review video, information related to the user's performance in the application can be fully considered, improving the effective information amount and personalization degree of the application review video, thereby improving the information effectiveness of effective interaction.
[0572] In one embodiment, as Figure 31 shown, playing the application review video in step 360 includes:
[0573] Step 3110: When playing the application review video, associate the facial image of the target character corresponding to the target object in the application review video with the inherent facial image of the target character and the facial image of the target object.
[0574] At this time, the facial image of the target character corresponding to the target object is the facial image obtained by fusing the inherent facial image of the target character and the facial image of the target object.
[0575] The inherent image of the target character refers to the inherent facial image set for the character by the character designer of the target application when designing the character. This inherent facial image is usually designed according to factors such as the character's personality, characteristics, and character settings, and is the basic appearance of the character. The facial image of the target object refers to the facial image of the real player.
[0576] The facial image of the target character corresponding to the target object is associated with the inherent facial image of the target character and the facial image of the target object, which means that the inherent facial image of the target character and the facial image of the target object are fused to obtain the facial image of the target character corresponding to the target object. At this time, the facial image of the target character corresponding to the target object contains both the real image of the target object and the image of the target character of the target object during the target application process.
[0577] In one embodiment, specifically, as Figure 32 shown, step 3110 includes:
[0578] Step 3210: Obtain the image of the target object through the camera, and extract the facial image of the target object from the image;
[0579] Step 3220: Obtain the inherent facial image of the target character from the target application;
[0580] Step 3230: Extract the first image feature from the inherent facial image of the target character, and extract the second image feature from the facial image of the target object;
[0581] Step 3240: Through the generative adversarial network, fuse the first image feature and the second image feature to obtain a fusion feature, and use the fusion feature for 3D modeling to obtain the facial image of the target character in the application review video.
[0582] In step 3210, after starting video recording, an image of the target object is obtained through the camera. The image at this time can include an image of the player's personal image. The method for extracting the facial image of the target object from the image includes face detection algorithms (such as Haar cascade detectors, deep learning-based face detectors, etc.), key point detection algorithms (such as the Dlib library, face key point detectors in OpenCV, etc.), face recognition algorithms (such as deep learning-based face recognition models, face recognizers in OpenCV, etc.), and image segmentation algorithms (such as the GrabCut algorithm, deep learning-based image segmentation models, etc.). These techniques can be applied alone or in combination to obtain more accurate and comprehensive facial image extraction results. In this way, the face part in the image can be separated from the background, and the facial image of the target object can be accurately extracted.
[0583] In step 3220, AI data collection is started. At this time, the model, texture, and animation of the target character can be obtained from the target application page, so as to obtain the inherent facial image of the target character. At this time, when obtaining the inherent facial image of the target character, the permission of the application developer needs to be obtained, and the issue of copyright needs to be considered.
[0584] In step 3230, the first image feature refers to the feature extracted from the inherent facial image of the target character. The second image feature refers to the feature extracted from the facial image of the target object. The first image feature and the second image feature can include features such as facial features, body features, and action features. The method for feature extraction can adopt deep learning, convolutional neural network (CNN), etc.
[0585] In step 3240, feature fusion refers to the process of fusing different features or feature representations together to obtain more comprehensive and accurate information. In the embodiment of the present disclosure, the first image feature and the second image feature are fused through a generative adversarial network to obtain a fused feature. The fused feature at this time contains both the features of the inherent facial image of the target character and the features of the facial image of the target object. The process of using the fused feature for 3D modeling is the process of image fusion. That is, the personal image of the target object is fused with the image of the target character to generate a new facial image of the target character, and the facial image of the target character contains both the features of the target object and the features of the target character.
[0586] It should be noted that the method for feature fusion can also be implemented by simple weighted averaging, splicing, cascading, etc.
[0587] It should be noted that the facial image of the target character can also be adjusted according to the real-time emotion of the target object through the above steps 2910 and 3010, which will not be elaborated here.
[0588] In practical applications, when playing an application review video, it can be specifically played in one or more ways among steps 2710, step 2810, step 2910, step 3010, and step 3110, which are not specifically limited herein.
[0589] Compared with the related art that adopts a fixed way of generating an application review video, the embodiments of the present disclosure can generate an application review video for an object after the target application ends according to the real-time emotion of the target object participating in the target application during the process of the target application. This automatically generated application review video is more interesting, can better bring the intensity of the situation at a certain period of the target object's game, and can better strengthen the emotional expression during the game and the catharsis experience during the application process. During the process of generating the application review video, data such as the target object's own expressions, body movements, and voice comments during the game are personalized integrated. At the same time, the performance of the target object in past games can also be obtained, the emotional portrait of the target object is analyzed, and the facial image of a new target character that better conforms to the personal characteristics of the target object is generated. Moreover, according to the real-time emotion of the target object, the style and rhythm of the video, the actions of the target character, and the size and facial expressions of the target character's face are adjusted, so as to improve the effective information volume and personalization degree of the application review video, thereby improving the information effectiveness of effective interaction.
[0590] Description of the devices and equipment of the embodiments of the present disclosure
[0591] It can be understood that although each step in the above-mentioned various flowcharts is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0592] It should be noted that in each specific embodiment of the present application, when it comes to performing relevant processing based on data related to the target object, such as facial image information, voice information, application screen recording information, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain data related to the target object, a separate permission or separate consent of the target object will be obtained through methods such as popping up a window or jumping to a confirmation page (equivalent to the target object enabling the corresponding control in the interface). After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiments of the present application will be obtained.
[0593] Figure 33 FIG. 3300 is a schematic structural diagram of an application processing apparatus 3300 provided by an embodiment of the present disclosure. The application processing apparatus 3300 includes:
[0594] A first receiving unit 3310, configured to receive a trigger on a recording control on a target application page after the target application starts;
[0595] An obtaining unit 3320, configured to obtain multiple emotion manifestation elements of the target object during the progress of the target application;
[0596] An input unit 3330, configured to input, for each emotion manifestation element, the emotion manifestation element into an emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain an emotion sub-curve of the target object corresponding to the emotion manifestation element;
[0597] A generating unit 3340, configured to generate an emotion curve of the target object based on the emotion sub-curve of the target object corresponding to the emotion manifestation element;
[0598] A recording unit 3350, configured to record an application video when the emotion curve of the real-time emotion meets a predetermined condition as an application review video;
[0599] A playing unit 3360, configured to play the application review video in response to a trigger on an application review video playing control on the target application page after the target application ends.
[0600] Optionally, before playing the application review video in response to a trigger on an application review video playing control on the target application page after the target application ends, the application processing apparatus further includes:
[0601] A first display unit (not shown), configured to display a duration input area in response to a trigger on a duration setting control on the target application page;
[0602] A second receiving unit (not shown) that receives the set duration input in the duration input area;
[0603] Among them, the application review video has a set duration.
[0604] Optionally, before playing the application review video in response to the triggering of the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0605] A second display unit (not shown) that, in response to the triggering of the role duration ratio control on the target application page, displays a role duration ratio input area;
[0606] A third receiving unit (not shown) that receives the first set duration ratio of each role in the target application in the application review video input in the role duration ratio input area;
[0607] Among them, the duration ratio allocated to each role in the target application in the application review video is equal to the first set duration ratio.
[0608] Optionally, before playing the application review video in response to the triggering of the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0609] A third display unit (not shown) that, in response to the triggering of the emotion duration ratio control on the target application page, displays an emotion duration ratio input area;
[0610] A fourth receiving unit (not shown) that receives the second set duration ratio of various emotions associated with the target object in the application review video input in the emotion duration ratio input area;
[0611] Among them, the video duration ratio of various emotions associated with the target object in the application review video is equal to the second set duration ratio.
[0612] Optionally, the multiple emotion manifestation elements include at least one of the target object's expression, the target object's body language, the target object's voice, the application comment, the target object's heart rate, the target object's screen pressing force, the application process data, and the target object's historical performance data;
[0613] Obtaining multiple emotion manifestation elements of the target object during the process of the target application includes at least one of the following:
[0614] Obtaining the application process data of the target object during the process of the target application;
[0615] Obtaining an image of the target object through a camera, and identifying the target object's expression and the target object's body language from the image;
[0616] Collect the target object voice of the target object through a radio receiver;
[0617] Obtain application reviews from the application screen;
[0618] Obtain the target object heart rate of the target object from a heart rate detection device;
[0619] Obtain the target object screen pressing force of the target object from the application screen;
[0620] Obtain the target object historical performance data from the data source of the target application.
[0621] Optionally, obtain the application process data of the target object during the process of the target application, including:
[0622] Obtain the first application subprocess data from the application log of the target application;
[0623] Obtain the second application subprocess data from the application code of the target application;
[0624] Obtain the third application subprocess data from the screenshot of the target application;
[0625] Integrate the first application subprocess data, the second application subprocess data, and the third application subprocess data into application process data.
[0626] Optionally, obtain the target object screen pressing force of the target object from the application screen, including:
[0627] Obtain the touch behavior pattern of the target object on the application screen from the application screen;
[0628] Input the touch behavior pattern into the force prediction model to obtain the predicted target object screen pressing force.
[0629] Optionally, obtain the target object historical performance data from the data source of the target application, including:
[0630] Determine the data source of the target application;
[0631] Obtain a data extraction tool;
[0632] Use the data extraction tool to obtain the target object historical performance data from the data source through the target data interface of the data source.
[0633] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0634] Generate the emotion curve of the target object based on the emotion sub-curve of the target object corresponding to the emotion manifestation elements, including:
[0635] Obtain the first weight of each emotion manifestation element;
[0636] Using a first weight, perform a weighted sum on the emotion sub-curves corresponding to each emotion manifestation element of the target object to obtain the emotion curve of the target object.
[0637] Optionally, the emotion manifestation elements include the voice of the target object;
[0638] For each emotion manifestation element, input the emotion manifestation element into the emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain the emotion sub-curve of the target object corresponding to the emotion manifestation element, including:
[0639] Preprocess the voice of the target object;
[0640] Extract the emotion-related features from the preprocessed voice of the target object;
[0641] Input the emotion-related features into the emotion sub-curve prediction model corresponding to the voice of the target object to obtain the emotion sub-curve of the target object corresponding to the voice of the target object.
[0642] Optionally, the emotion manifestation elements include application reviews, and the emotion sub-curve prediction model corresponding to the application reviews includes an emotion dictionary, a natural language emotion recognition model, and a voice emotion recognition model;
[0643] For each emotion manifestation element, input the emotion manifestation element into the emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain the emotion sub-curve of the target object corresponding to the emotion manifestation element, including:
[0644] For the text reviews in the application reviews, obtain the emotion type of the text reviews through the emotion dictionary, so as to obtain the first emotion score curve corresponding to the text reviews;
[0645] Input the text reviews in the application reviews into the natural language emotion recognition model to obtain the second emotion score curve corresponding to the text reviews;
[0646] Integrate the first emotion score curve and the second emotion score curve to obtain the first integrated score curve corresponding to the text reviews;
[0647] Input the voice reviews in the application reviews into the voice emotion recognition model to obtain the third emotion score curve corresponding to the voice reviews;
[0648] Based on the first integrated score curve and the third emotion score curve, determine the emotion sub-curve.
[0649] Optionally, each emotion manifestation element includes the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data. Among them, the first weights of the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data decrease in turn.
[0650] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0651] Recording the application video when the emotion curve of the real-time emotion meets a predetermined condition as an application review video, including:
[0652] On the emotion curve, intercept the target emotion curve segment according to the review video generation rule;
[0653] Determine the object emotion corresponding to each target emotion curve segment;
[0654] For the object emotion corresponding to the target emotion curve segment, select a matching base map corresponding to the target object and the object emotion from the material library;
[0655] Input the emotion curve, the matching base map, and the background data of the target application into the generation model to obtain the application review video.
[0656] Optionally, the review video generation rule is: in at least the first ratio of the application review video, the emotion value of the target object meets the first condition;
[0657] On the emotion curve, intercept the target emotion curve segment according to the review video generation rule, including:
[0658] On the emotion curve, determine the first candidate part where the emotion value of the target object meets the first condition;
[0659] Based on the video duration of the application review video and the first ratio, determine the duration to be intercepted;
[0660] Based on the duration to be intercepted, intercept the target emotion curve segment from the first candidate part.
[0661] Optionally, the review video generation rule is: the application review video includes multiple video parts, and the emotion value of the target object in each video part meets the second condition associated with the video part, and the duration ratio of the multiple video parts is the predetermined duration ratio;
[0662] On the emotion curve, intercept the target emotion curve segment according to the review video generation rule, including:
[0663] For each video segment, on the emotion curve, determine a second candidate segment where the emotion value of the target object meets the second condition associated with the video segment;
[0664] Based on the video duration of the application review video and the preset duration ratio, determine the video segment duration of each video segment;
[0665] Based on the video segment duration, intercept the intercepted part from the second candidate segments corresponding to the video segments, and integrate the intercepted parts corresponding to each video segment into a target emotion curve segment.
[0666] Optionally, the review video generation rule is: the duration ratio of each character in the application review video is equal to the preset duration ratio of the character, and each character includes the target character of the target object in the target application;
[0667] Intercept the target emotion curve segment on the emotion curve according to the review video generation rule, including:
[0668] On the emotion curve, determine a third candidate segment where the emotion value of the target object meets the third condition, and in the target application, for each other character, determine the time period when the other character appears, and find the fourth candidate segment corresponding to the time period on the emotion curve;
[0669] Based on the video duration of the application review video and the preset duration ratio of the target character, determine the target character video duration occupied by the target character, and based on the video duration of the application review video and the preset duration ratio of each other character, determine the other character video duration occupied by each other character;
[0670] Based on the target character video duration, intercept the target character curve part from the third candidate segment, and based on the other character video duration, intercept the other character curve part from the fourth candidate segment;
[0671] Integrate the target character curve part and each other character curve part into a target emotion curve segment.
[0672] Optionally, the review video generation rule is: the duration ratio of various emotions associated with the target object in the application review video is equal to the preset duration ratio corresponding to the emotion;
[0673] Intercept the target emotion curve segment on the emotion curve according to the review video generation rule, including:
[0674] For each emotion of the target object, on the emotion curve, determine the fifth candidate segment corresponding to the emotion;
[0675] Based on the video duration of the application review video and the preset duration ratio corresponding to the emotion, determine the duration corresponding to the emotion;
[0676] For each emotion, select a sub - segment corresponding to the emotion from the fifth candidate part corresponding to the emotion according to the duration corresponding to the emotion;
[0677] Integrate the sub - segments corresponding to various emotions into the target emotion curve segments.
[0678] Optionally, determine the object emotions corresponding to each target emotion curve segment, including:
[0679] Obtain the emotion value corresponding to each time point in the target emotion curve segment;
[0680] Based on the emotion values, with reference to the corresponding relationship between the object emotion and the emotion value, determine the object emotion at each time point in the target emotion curve segment.
[0681] Optionally, the material library includes a first material library and a second material library. Among them, the first material in the first material library has an object expression map, an object label, and a first emotion label; the second material in the second material library has a character expression map, a character label, and a second emotion label;
[0682] For the object emotion corresponding to the target emotion curve segment, select a matching basic map corresponding to the target object and the object emotion from the material library, including:
[0683] Query the first material library. If the object label of a first material in the first material library corresponds to the target object and the first emotion label corresponds to the object emotion, use the object expression map of the first material as the matching basic map;
[0684] If no matching basic map is found by querying the first material library, query the second material library. If the character label of a second material in the second material library corresponds to the target role of the target object in the target application and the second emotion label corresponds to the object emotion, use the character expression map of the second material as the matching basic map.
[0685] Optionally, the first material library is generated in the following way:
[0686] Obtain materials with object expression maps from the Internet as the first materials;
[0687] Input the first materials into an emotion recognition model to obtain the first emotion labels;
[0688] Input the object label map into an object recognition model to obtain the object labels;
[0689] Integrate each first material with the first emotion label and the object label into the first material library.
[0690] Optionally, the second material library is generated in the following way:
[0691] Obtain materials with character expression diagrams from the Internet as the second material;
[0692] Input the second material into the emotion recognition model to obtain the second emotion label;
[0693] Obtain the character corresponding to the character label diagram as the character label;
[0694] Integrate each second material with the second emotion label and the character label into a second material library.
[0695] Optionally, the generation model includes a copywriting part generation sub-model, a video part generation sub-model, and an audio part generation sub-model;
[0696] Input the emotion curve, the matching base map, and the background data of the target application into the generation model to obtain an application review video, including:
[0697] Input the emotion curve and the matching base map into the copywriting part generation sub-model to obtain the copywriting part;
[0698] Input the emotion curve, the matching base map, and the background data of the target application into the video part generation sub-model to obtain the video part;
[0699] Input the emotion curve and the matching base map into the audio part generation sub-model to obtain the audio part;
[0700] Integrate the copywriting part, the video part, and the audio part into an application review video.
[0701] Optionally, the generation model is obtained in the following way:
[0702] Obtain a training sample set, where the training samples in the training sample set include sample applications, in-progress data of the sample object during the sample application process, and sample application review videos corresponding to the sample applications;
[0703] Based on the sample application and the in-progress data, obtain the sample emotion curve of the sample object, the sample matching base map, and the sample background data of the sample application;
[0704] Input the sample emotion curve, the sample matching base map, and the sample background data into the generation model, and train the generation model based on the comparison between the model generation result and the sample application review video;
[0705] Optimize the parameters of the trained generation model;
[0706] Test the generation model after parameter optimization.
[0707] Optionally, the playback unit 3360 is specifically used for:
[0708] When playing the application review video, the actions of the target character corresponding to the target object are scaled based on the real-time emotion of the target object.
[0709] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0710] The playback unit 3360 is specifically further configured to:
[0711] Input the emotion curve of the target object into a generation model to obtain the action amplitude of the target character that changes over time, and generate the actions of the target character in the application review video through skeletal animation and animation blending.
[0712] Optionally, the playback unit 3360 is specifically further configured to:
[0713] When playing the application review video, adjust the style and rhythm of the application review video based on the real-time emotion of the target object.
[0714] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0715] The playback unit 3360 is specifically further configured to:
[0716] Input the emotion curve of the target object into a generation model to obtain the style parameters and rhythm parameters of the target character that change over time, and generate video frames of the application review video using the style parameters and rhythm parameters, where the style of the video frame corresponds to the style parameters and the rhythm of the video frame corresponds to the rhythm parameters.
[0717] Optionally, the playback unit 3360 is specifically further configured to:
[0718] When playing the application review video, adjust the face size and facial expression of the target character corresponding to the target object based on the real-time emotion of the target object.
[0719] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0720] The playback unit 3360 is specifically further configured to:
[0721] Input the emotion curve of the target object into a generation model to obtain the face size parameters of the target character that change over time, and generate video frames of the application review video using the emotion curve and the face size parameters, where the facial expression of the video frame corresponds to the emotion value at each time point on the emotion curve, and the face size of the video frame corresponds to the face size parameters.
[0722] Optionally, the playback unit 3360 is specifically further configured to:
[0723] When playing the application review video, add enhancement elements based on the real-time emotion of the target object.
[0724] Optionally, the real-time emotion includes an emotion curve that changes over time;
[0725] The playback unit 3360 is further specifically configured to:
[0726] Input the emotion curve of the target object into a generation model to obtain the reinforcement factor that changes over time of the target character, and add the reinforcement factor to the video frames of the application review video.
[0727] Optionally, the playback unit 3360 is further specifically configured to:
[0728] When playing the application review video, make the facial image of the target character corresponding to the target object in the application review video be jointly associated with the inherent facial image of the target character and the facial image of the target object.
[0729] Optionally, the playback unit 3360 is further specifically configured to:
[0730] Obtain an image of the target object through a camera, and extract the facial image of the target object from the image;
[0731] Obtain the inherent facial image of the target character from the target application;
[0732] Extract the first image feature from the inherent facial image of the target character, and extract the second image feature from the facial image of the target object;
[0733] Through a generative adversarial network, fuse the first image feature and the second image feature to obtain a fused feature, and use the fused feature for three-dimensional modeling to obtain the facial image of the target character in the application review video.
[0734] Optionally, after the target application starts and before receiving the trigger of the recording control on the target application page, the application processing device further includes:
[0735] A fourth display unit (not shown), configured to display a setting interface before the target application starts, where the setting interface includes a free screen recording start control, a camera start control, a sound collection start control, and a data collection permission control;
[0736] A fifth receiving unit (not shown), configured to receive start commands for the free screen recording start control, the camera start control, the sound collection start control, and the data collection permission control on the setting interface.
[0737] Optionally, after responding to the trigger of the application review video playback control on the target application page after the target application ends, the application processing device further includes:
[0738] A fifth display unit (not shown), configured to display a sharing control;
[0739] A sharing unit (not shown) for sharing the generated application review video to other target terminals in response to the triggering of a sharing control.
[0740] Refer to Figure 34 , Figure 34 FIG. 0001635 is a block diagram of a part of the target terminal 110 for implementing the application processing method according to an embodiment of the present disclosure. The terminal includes components such as a Radio Frequency (RF) circuit 3410, a memory 3415, an input unit 3430, a display unit 3440, a sensor 3450, an audio circuit 3460, a wireless fidelity (WiFi) module 3470, a processor 3480, and a power supply 3490. Those skilled in the art can understand that Figure 34 the structure of the target terminal 110 shown does not limit a mobile phone or a computer, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0741] The RF circuit 3410 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from a base station, it is given to the processor 3480 for processing; in addition, the designed uplink data is sent to the base station.
[0742] The memory 3415 can be used for storing software programs and modules. The processor 3480 executes various functional applications and data processing of the content terminal by running the software programs and modules stored in the memory 3415.
[0743] The input unit 3430 can be used for receiving input digital or character information, and generating key signal inputs related to the settings and function controls of the content terminal. Specifically, the input unit 3430 may include a touch panel 3431 and other input devices 3432.
[0744] The display unit 3440 can be used for displaying input information or provided information and various menus of the content terminal. The display unit 3440 may include a display panel 3441.
[0745] The audio circuit 3460, a speaker 3461, and a microphone 3462 can provide an audio interface.
[0746] In this embodiment, the processor 3480 included in the target terminal 110 can execute the application processing method of the previous embodiment.
[0747] The target terminal 110 according to an embodiment of the present disclosure includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc. Embodiments of the present invention can be applied to various scenarios, including but not limited to human-computer interaction, data processing, etc.
[0748] Figure 35 A partial structural block diagram of the application server 140 for implementing the application processing method of the embodiments of the present disclosure. The application server 140 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 3522 (for example, one or more processors) and a memory 3532, and one or more storage media 3530 (for example, one or more mass storage devices) for storing application programs 3542 or data 3544. Among them, the memory 3532 and the storage media 3530 may be transient storage or persistent storage. The program stored in the storage media 3530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 3522 may be configured to communicate with the storage media 3530 and execute a series of instruction operations in the storage media 3530 on the server.
[0749] The application server 140 may further include one or more power supplies 3526, one or more wired or wireless network interfaces 3550, one or more input / output interfaces 3558, and / or one or more operating systems 3541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0750] The central processing unit 3522 in the application server 140 may be used to execute the application processing method of the embodiments of the present disclosure.
[0751] The embodiments of the present disclosure further provide a computer-readable storage medium for storing program codes for executing the application processing methods of the foregoing various embodiments.
[0752] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above application processing method.
[0753] In the description of the present disclosure and the above - mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar content and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "including" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0754] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated content and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates that the content before and after is an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single - item (one) or plural - item (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0755] It should be understood that in the description of the embodiments of the present disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0756] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.
[0757] The unit described as a separate component may or may not be physically separated, and the component presented as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0758] In addition, in the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the functions of that module or unit.
[0759] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0760] It should also be understood that the various embodiments provided in this disclosure can be combined arbitrarily to achieve different technical effects.
[0761] The above is a specific description of the embodiments of this disclosure, but this disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of this disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this disclosure.
Claims
1. An application processing method, characterized in that, Including: After the target application starts, receiving a trigger of a recording control on the target application page; Obtaining multiple emotional manifestation elements of the target object during the progress of the target application; For each of the emotional manifestation elements, inputting the emotional manifestation element into an emotional sub-curve prediction model corresponding to the emotional manifestation element to obtain an emotional sub-curve of the target object corresponding to the emotional manifestation element; Generating an emotional curve of the target object based on the emotional sub-curve of the target object corresponding to the emotional manifestation element; Recording an application video when the emotional curve of the real-time emotion meets a predetermined condition as an application review video; Responding to a trigger of an application review video playback control on the target application page after the target application ends, and playing the application review video.
2. The application processing method according to claim 1, wherein Before responding to a trigger of an application review video playback control on the target application page after the target application ends and playing the application review video, the application processing method further includes: Responding to a trigger of a duration setting control on the target application page, and displaying a duration input area; Receiving a set duration input in the duration input area; Wherein, the application review video has the set duration.
3. The application processing method according to claim 1, wherein Before responding to a trigger of an application review video playback control on the target application page after the target application ends and playing the application review video, the application processing method further includes: Responding to a trigger of a role duration ratio control on the target application page, and displaying a role duration ratio input area; Receiving a first set duration ratio of each role in the target application in the application review video input in the role duration ratio input area; Wherein, the duration ratio of the application review video allocated to each role in the target application is equal to the first set duration ratio.
4. The application processing method according to claim 1, characterized in that, Before responding to a trigger of an application review video playback control on the target application page after the target application ends and playing the application review video, the application processing method further includes: Responding to a trigger of an emotion duration ratio control on the target application page, and displaying an emotion duration ratio input area; Receiving a second set duration ratio of various emotions associated with the target object in the application review video input in the emotion duration ratio input area; Wherein, the video duration ratio of various emotions associated with the target object in the application review video is equal to the second set duration ratio.
5. The application processing method according to claim 1, wherein, The multiple emotional manifestation elements include at least one of the target object's expression, the target object's body language, the target object's voice, application comments, the target object's heart rate, the target object's screen pressing force, application process data, and the target object's historical performance data; The obtaining of the multiple emotional manifestation elements of the target object during the progress of the target application includes at least one of the following: Obtaining the application process data of the target object during the progress of the target application; Obtaining an image of the target object through a camera, and identifying the target object's expression and the target object's body language from the image; Collect the target object voice of the target object through a radio; Obtain the application review from the application screen; Obtain the target object heart rate of the target object from a heart rate detection device; Obtain the target object screen pressing force of the target object from the application screen; Obtain the target object historical performance data of the target object from the data source of the target application.
6. The application processing method according to claim 1, wherein, The real-time emotion includes an emotion curve that changes over time; Generating the emotion curve of the target object based on the emotion sub-curves of the target object corresponding to the emotion manifestation elements includes: Obtain the first weight of each of the emotion manifestation elements; Using the first weight, perform a weighted sum on the emotion sub-curves of the target object corresponding to each of the emotion manifestation elements to obtain the emotion curve of the target object.
7. The application processing method according to claim 1, wherein The real-time emotion includes an emotion curve that changes over time; Recording the application video when the emotion curve of the real-time emotion meets a predetermined condition as an application review video includes: On the emotion curve, intercept a target emotion curve segment according to the review video generation rule; Determine the object emotion corresponding to each of the target emotion curve segments; For the object emotion corresponding to the target emotion curve segment, select a matching base map corresponding to the target object and the object emotion from the material library; Input the emotion curve, the matching base map, and the background data of the target application into a generation model to obtain the application review video.
8. The application processing method according to claim 7, wherein The material library includes a first material library and a second material library, wherein the first material in the first material library has an object expression map, an object label, and a first emotion label; the second material in the second material library has a character expression map, a character label, and a second emotion label; The selecting a matching base map corresponding to the target object and the object emotion from the material library for the object emotion corresponding to the target emotion curve segment includes: Query the first material library. If the object label of a first material in the first material library corresponds to the target object and the first emotion label corresponds to the object emotion, use the object expression map of the first material as the matching base map; If the matching base map is not found by querying the first material library, query the second material library. If the character label of a second material in the second material library corresponds to the target role of the target object in the target application and the second emotion label corresponds to the object emotion, use the character expression map of the second material as the matching base map.
9. The application processing method according to claim 7, wherein, The generation model includes a copywriting part generation sub-model, a video part generation sub-model, and an audio part generation sub-model; Inputting the emotion curve, the matching base map, and the background data of the target application into the generation model to obtain the application review video includes: Input the emotion curve and the matching base map into the copywriting part generation sub-model to obtain the copywriting part; Input the emotion curve, the matching base map, and the background data of the target application into the video part generation sub-model to obtain the video part; Input the emotional curve and the matching base map into the sub-model for generating the audio part to obtain the audio part; Integrate the copywriting part, the video part and the audio part into the application review video.
10. The application processing method according to claim 1, wherein The playing of the application review video includes: when playing the application review video, scaling the actions of the target character corresponding to the target object based on the real-time emotion of the target object.
11. The application processing method according to claim 1, wherein The playing of the application review video includes: when playing the application review video, adjusting the style and rhythm of the application review video based on the real-time emotion of the target object.
12. The application processing method according to claim 1, wherein The playing of the application review video includes: when playing the application review video, adjusting the facial size and facial expression of the target character corresponding to the target object based on the real-time emotion of the target object.
13. The application processing method according to claim 1, wherein The playing of the application review video includes: when playing the application review video, adding enhancement elements based on the real-time emotion of the target object.
14. The application processing method according to claim 1, wherein The playing of the application review video includes: when playing the application review video, making the facial image of the target character corresponding to the target object in the application review video jointly associated with the inherent facial image of the target character and the facial image of the target object.
15. The application processing method according to claim 14, characterized in that The making the facial image of the target character corresponding to the target object in the application review video jointly associated with the inherent facial image of the target character and the facial image of the target object when playing the application review video includes: Obtain the image of the target object through a camera, and extract the facial image of the target object from the image; Obtain the inherent facial image of the target character from the target application; Extract the first image feature from the inherent facial image of the target character, and extract the second image feature from the facial image of the target object; Through a generative adversarial network, fuse the first image feature and the second image feature to obtain a fused feature, and use the fused feature for three-dimensional modeling to obtain the facial image of the target character in the application review video.
16. An application processing device, characterized in that, It includes: A first receiving unit, configured to receive a trigger on a recording control on a target application page after the target application starts; An obtaining unit, configured to obtain multiple emotion manifestation elements of the target object during the process of the target application; An input unit, configured to input, for each emotion manifestation element, the emotion manifestation element into an emotion sub-curve prediction model corresponding to the emotion manifestation element to obtain an emotion sub-curve of the target object corresponding to the emotion manifestation element; A generating unit, configured to generate an emotion curve of the target object based on the emotion sub-curve of the target object corresponding to the emotion manifestation element; A recording unit, configured to record the application video when the emotion curve of the real-time emotion meets a predetermined condition as an application review video; A playing unit, configured to play the application review video in response to a trigger on an application review video playing control on the target application page after the target application ends.
17. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the application processing method according to any one of claims 1 to 15.
18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the application processing method according to any one of claims 1 to 15.
19. A computer program product, which includes a computer program that is read and executed by a processor of a computer device, so that the computer device executes the application processing method according to any one of claims 1 to 15.