Virtual character generation method, device and equipment and computer readable storage medium
By analyzing program screenshots and audio using a recognition model, virtual character images that conform to user habits are generated, solving the problem that virtual character images cannot be changed and improving user interaction willingness and experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2026-03-03
AI Technical Summary
The existing virtual character designs cannot be changed, leading to reduced user engagement and a poor user experience.
By acquiring program screenshots and audio within a preset time period, a recognition model is used to identify user habits, and virtual character images that meet user needs are generated based on these habits.
It increased users' willingness to interact with virtual characters and improved the user experience.
Smart Images

Figure CN115423909B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of television technology, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating virtual characters. Background Technology
[0002] With the development of television technology, a virtual character is now often featured on televisions. This virtual character may provide information via voice or other means at certain times. However, the appearance of this virtual character is currently fixed and cannot be changed. Often, this appearance does not meet the needs of users. When users interact with the default virtual character, they may become bored, reducing their willingness to interact and resulting in a poor user experience.
[0003] Therefore, how to generate virtual character images and enhance users' willingness to interact is an urgent problem to be solved. Summary of the Invention
[0004] The main objective of this invention is to provide a method, apparatus, device, and computer-readable storage medium for generating virtual characters, aiming to solve the problem of how to generate virtual character images and enhance users' willingness to interact.
[0005] To achieve the above objectives, the present invention provides a method for generating virtual characters, the method comprising the following steps:
[0006] Get screenshots and audio corresponding to programs played within a preset time period;
[0007] The screenshot and audio are identified using a pre-created recognition model to obtain the current recognition result, and the user's usage habits are determined based on the current recognition result and historical recognition results.
[0008] The virtual character image is determined based on the user's usage habits, and the virtual character is generated based on the virtual character image.
[0009] Optionally, the steps for obtaining screenshots and audio corresponding to programs played within a preset time period include:
[0010] Get the current time period and compare it with the preset time period;
[0011] If the current time period is within the preset time period, then according to the preset acquisition frequency, screenshot and audio recording operations are performed on the program played in the current time period to obtain the screenshot and audio corresponding to the program played within the preset time period.
[0012] Optionally, the step of identifying the screenshot and the audio using a pre-created recognition model to obtain the current recognition result includes:
[0013] Key pixels in the screenshot are extracted using a pre-created recognition model, and the first scene information corresponding to the screenshot is determined based on the key pixels.
[0014] The voiceprint features corresponding to the audio are extracted by a pre-created recognition model, and the second scene information corresponding to the audio is determined based on the voiceprint features.
[0015] The current recognition result is determined based on the first scene information and the second scene information.
[0016] Optionally, before the step of determining user habits based on the current recognition results and historical recognition results, the following steps are included:
[0017] Obtain the scene proportion set of the current recognition results, and compare the scene proportion set with the preset confidence level;
[0018] If there is a scene proportion in the scene proportion set that is greater than the preset confidence level, then the current recognition result is stored and the following steps are executed: determine the user's usage habits based on the current recognition result and the historical recognition result;
[0019] If there is no scene proportion in the scene proportion set that is greater than the preset confidence level, then delete the current recognition result and re-execute the step: obtain the screenshots and audio corresponding to the programs played within the preset time period.
[0020] Optionally, the step of determining user habits based on the current recognition results and historical recognition results includes:
[0021] Obtain the previous preset number of historical recognition results corresponding to the current recognition result, and calculate the similarity between the current recognition result and the previous preset number of historical recognition results;
[0022] If the similarity between the current recognition result and the previous preset number of historical recognition results is greater than the preset similarity threshold, then the user's usage habits are determined based on the current recognition result and the previous preset number of historical recognition results.
[0023] Optionally, the steps of determining the virtual character image based on the user's usage habits and generating a virtual character based on the virtual character image include:
[0024] The user's usage habits are input into a pre-created virtual character management module, which then determines the virtual character based on those habits.
[0025] Obtain the user image corresponding to the user's usage habits;
[0026] The virtual character image and the user image are sent to the cloud-based virtual character creation terminal, which generates a virtual character based on the virtual character image and the user image, and then sends the virtual character back to the virtual character image management module.
[0027] Optionally, after the steps of determining the virtual character image based on the user's usage habits and generating a virtual character based on the virtual character image, the process includes:
[0028] Monitor push notifications and display them to users through the virtual character.
[0029] Furthermore, to achieve the above objectives, the present invention also provides a virtual character generation device, the virtual character generation device comprising:
[0030] The acquisition module is used to acquire screenshots and audio corresponding to programs played within a preset time period;
[0031] The recognition module is used to recognize the screenshot and the audio through a pre-created recognition model, obtain the current recognition result, and determine the user's usage habits based on the current recognition result and the historical recognition result;
[0032] The generation module is used to determine the virtual character image based on the user's usage habits, and generate a virtual character based on the virtual character image.
[0033] Furthermore, the acquisition module is also used for:
[0034] Get the current time period and compare it with the preset time period;
[0035] If the current time period is within the preset time period, then according to the preset acquisition frequency, screenshot and audio recording operations are performed on the program played in the current time period to obtain the screenshot and audio corresponding to the program played within the preset time period.
[0036] Furthermore, the identification module is also used for:
[0037] Key pixels in the screenshot are extracted using a pre-created recognition model, and the first scene information corresponding to the screenshot is determined based on the key pixels.
[0038] The voiceprint features corresponding to the audio are extracted by a pre-created recognition model, and the second scene information corresponding to the audio is determined based on the voiceprint features.
[0039] The current recognition result is determined based on the first scene information and the second scene information.
[0040] Furthermore, the identification module is also used for:
[0041] Obtain the scene proportion set of the current recognition results, and compare the scene proportion set with the preset confidence level;
[0042] If there is a scene proportion in the scene proportion set that is greater than the preset confidence level, then the current recognition result is stored and the following steps are executed: determine the user's usage habits based on the current recognition result and the historical recognition result;
[0043] If there is no scene proportion in the scene proportion set that is greater than the preset confidence level, then delete the current recognition result and re-execute the step: obtain the screenshots and audio corresponding to the programs played within the preset time period.
[0044] Preferably, the identification module further includes a determination module, the determination module being used to:
[0045] Obtain the previous preset number of historical recognition results corresponding to the current recognition result, and calculate the similarity between the current recognition result and the previous preset number of historical recognition results;
[0046] If the similarity between the current recognition result and the previous preset number of historical recognition results is greater than a preset similarity threshold, then the user's usage habits are determined based on the current recognition result and the previous preset number of historical recognition results. Preferably, the generation module is further used for:
[0047] The user's usage habits are input into a pre-created virtual character management module, which then determines the virtual character based on those habits.
[0048] Obtain the user image corresponding to the user's usage habits;
[0049] The virtual character image and the user image are sent to the cloud-based virtual character creation terminal, which generates a virtual character based on the virtual character image and the user image, and then sends the virtual character back to the virtual character image management module.
[0050] Preferably, the generation module further includes a display module, the display module being used for:
[0051] Monitor push notifications and display them to users through the virtual character.
[0052] In addition, to achieve the above objectives, the present invention also provides a virtual character generation device, the virtual character generation device comprising: a memory, a processor, and a virtual character generation program stored in the memory and executable on the processor, wherein when the virtual character generation program is executed by the processor, it implements the steps of the virtual character generation method as described above.
[0053] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a virtual character generation program, which, when executed by a processor, implements the steps of the virtual character generation method described above.
[0054] This invention proposes a virtual character generation method that involves acquiring screenshots and audio corresponding to programs played within a preset time period; identifying the screenshots and audio using a pre-created recognition model to obtain the current recognition result; determining user habits based on the current and historical recognition results; determining the virtual character image based on the user habits; and generating a virtual character based on the virtual character image. This invention uses a recognition model to identify screenshots and audio corresponding to programs played within a preset time period to determine user habits, then determines the virtual character image based on these habits, and finally generates a virtual character. This ensures that the virtual character meets the user's needs, thereby increasing the user's willingness to interact with the virtual character and improving the user experience. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention;
[0056] Figure 2 This is a flowchart illustrating the first embodiment of the virtual character generation method of the present invention;
[0057] Figure 3 This is a flowchart illustrating the second embodiment of the virtual character generation method of the present invention;
[0058] Figure 4 This is a schematic diagram of the virtual character generation device of the present invention.
[0059] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0060] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0061] like Figure 1 As shown, Figure 1 This is a schematic diagram of the device structure of the hardware operating environment involved in the embodiments of the present invention.
[0062] The device in this embodiment of the invention can be a PC or a server.
[0063] like Figure 1As shown, the device may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0064] Those skilled in the art will understand that Figure 1 The device structure shown does not constitute a limitation on the device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0065] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a virtual character generation program.
[0066] The operating system is a program that manages and controls the portable virtual character generation device and software resources, and supports the operation of the network communication module, user interface module, virtual character generation program and other programs or software; the network communication module is used to manage and control the network interface 1002; the user interface module is used to manage and control the user interface 1003.
[0067] exist Figure 1 In the virtual character generation device shown, the virtual character generation device calls the virtual character generation program stored in the memory 1005 through the processor 1001 and executes the operations in the various embodiments of the virtual character generation method described below.
[0068] Based on the above hardware structure, an embodiment of the virtual character generation method of the present invention is proposed.
[0069] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the virtual character generation method of the present invention, the method comprising:
[0070] Step S10: Obtain screenshots and audio corresponding to the programs played within the preset time period;
[0071] Step S20: Recognize the screenshot and the audio using a pre-created recognition model to obtain the current recognition result, and determine the user's usage habits based on the current recognition result and historical recognition results;
[0072] Step S30: Determine the virtual character image based on the user's usage habits, and generate a virtual character based on the virtual character image.
[0073] This embodiment of the virtual character generation method is applied to smart devices, including smart TVs, smart terminals, and PC terminals. For ease of description, a smart TV is used as an example. When a user watches a program, the smart TV obtains the current time period and compares it with a preset time period. If the current time period is within the preset time period, it performs screenshot and audio recording operations on the program played in the current time period according to a preset acquisition frequency, obtaining the screenshot and audio corresponding to the program played within the preset time period. The smart TV extracts key pixels from the screenshot using a pre-created recognition model and determines the first scene information corresponding to the screenshot based on the key pixels. The smart TV extracts the voiceprint features corresponding to the audio using a pre-created recognition model and determines the second scene information corresponding to the audio based on the voiceprint features. The smart TV determines the current recognition result based on the first scene information and the second scene information. The smart TV obtains the previous preset number of historical recognition results corresponding to the current recognition result and calculates the similarity between the current recognition result and the previous preset number of historical recognition results. If the similarity between the current recognition result and the previous preset number of historical recognition results is greater than a preset similarity threshold, the user's usage habits are determined based on the current recognition result and the previous preset number of historical recognition results. The smart TV inputs user habits into a pre-created virtual avatar management module, which then determines the virtual avatar based on these habits. The smart TV then retrieves the user image corresponding to these habits and generates the virtual avatar based on the user image and the virtual avatar designation. It's important to note that preset time slots are pre-set in the smart TV, and users can adjust them according to their needs. User habits refer to the types of programs the user watches for extended periods.
[0074] This embodiment of the virtual character generation method involves acquiring screenshots and audio corresponding to programs played within a preset time period; recognizing the screenshots and audio using a pre-created recognition model to obtain the current recognition result; determining user habits based on the current and historical recognition results; determining the virtual character image based on the user habits; and generating the virtual character based on the virtual character image. This invention uses a recognition model to identify screenshots and audio corresponding to programs played within a preset time period to determine user habits, determines the virtual character image based on the user habits, and then generates a virtual character. This ensures that the virtual character meets the user's needs, thereby increasing the user's willingness to interact with the virtual character and improving the user experience.
[0075] The following will provide a detailed explanation of each step:
[0076] Step S10: Obtain screenshots and audio corresponding to the programs played within the preset time period;
[0077] In this embodiment, while the user is watching a TV program, the smart TV acquires screenshots and audio corresponding to the programs played within a preset time period. Optionally, the preset time period is pre-set in the smart TV; generally, the preset time period is 24 hours, but the user can also set the preset time period according to their own needs. Optionally, the smart TV counts the time periods when the user uses the TV to determine the preset time period. For example, if 8 PM to 10 PM is the time period when users use the TV the most, the smart TV can use the 8 PM to 10 PM time period as the preset time period. Or, for example, if users use the TV more frequently throughout the day in July to August and January to February, the smart TV can use the entire day (24 hours) as the preset time period during July to August and January to February. This ensures that the smart TV acquires a large number of valuable screenshots and audio while not wasting its computing resources.
[0078] Understandably, smart TVs are equipped with screenshot and recording modules. When a smart TV detects that a user is watching a program on the TV within a preset time period, it activates the screenshot module to take a screenshot of the program being played and activates the recording module to record the program being played, thereby obtaining the corresponding audio of the program.
[0079] Further, step S10 includes:
[0080] Step a: Obtain the current time period and compare it with the preset time period;
[0081] In this step, after the smart TV starts up, it obtains the current time period and compares it with the preset time period to determine whether the current time period is within the preset time period, so as to determine whether to obtain the screenshot and audio corresponding to the program being played; preferably, when the preset time period is 24 hours a day, the smart TV does not need to obtain the current time period after starting up, and directly obtains the screenshot and audio corresponding to the program being played.
[0082] Step b: If the current time period is within a preset time period, then according to the preset acquisition frequency, perform screenshot and audio recording operations on the program played in the current time period to obtain the screenshot and audio corresponding to the program played within the preset time period.
[0083] In this step, if the smart TV determines that the current time period is within a preset time period, it activates the screenshot module and the recording module. The screenshot module takes a screenshot of the program playing in the current time period according to a preset acquisition frequency, obtaining a screenshot of the program playing within the preset time period. The screenshot module also records the audio of the program playing in the current time period according to the preset acquisition frequency, obtaining the audio of the program playing within the preset time period. Preferably, the preset acquisition frequency is pre-set in the smart TV, typically set to 330ms / time, meaning that a screenshot and recording operation is performed on the currently playing program every 330ms to obtain the corresponding screenshot and audio. Optionally, the preset acquisition frequency can be adjusted according to user needs.
[0084] Step S20: Recognize the screenshot and the audio using a pre-created recognition model to obtain the current recognition result, and determine the user's usage habits based on the current recognition result and historical recognition results;
[0085] In this embodiment, after acquiring screenshots and audio of programs played within a preset time period, the smart TV inputs the screenshots and audio into a pre-created recognition model. The recognition model then identifies the screenshots and audio to obtain the current recognition result. Based on the current recognition result, historical recognition results are obtained, and user habits are determined based on the current and historical recognition results. Preferably, during the process of acquiring screenshots and audio of programs played within the preset time period, the smart TV inputs the screenshots and audio into the recognition model. The recognition model identifies the input screenshots and audio. After the smart TV acquires the screenshots and audio of programs played within the preset time period, the recognition module can output the current recognition result. By recognizing the screenshots and audio during the acquisition process, recognition efficiency is improved. Optionally, after acquiring all the screenshots and audio of programs played within the preset time period, the smart TV inputs all the screenshots and audio into the recognition model to obtain the current recognition result.
[0086] It should be noted that each time a smart TV receives a recognition result, it determines whether the recognition result is valid and stores the valid recognition results in chronological order based on the time the recognition result was received, thus obtaining historical recognition results.
[0087] Specifically, the steps of identifying the screenshot and the audio using a pre-created recognition model to obtain the current recognition result include:
[0088] Step c: Extract key pixels from the screenshot using a pre-created recognition model, and determine the first scene information corresponding to the screenshot based on the key pixels;
[0089] In this step, the smart TV extracts key pixels from the screenshot using a pre-created recognition model and determines the first scene information corresponding to the screenshot based on these key pixels. It's understood that some pixels in the screenshot contain important information. For example, in a screenshot of an NBA live broadcast, there's usually the iconic NBA logo in the upper left corner, along with objects like basketballs and basketball hoops. The pixels used to represent these icons or objects are the key pixels. The recognition module determines the first scene information corresponding to the screenshot based on these key pixels. The first scene information indicates the type of program the screenshot belongs to, such as a basketball program, a news program, or a movie program.
[0090] Step d: Extract the voiceprint features corresponding to the audio using a pre-created recognition model, and determine the second scene information corresponding to the audio based on the voiceprint features;
[0091] In this step, the smart TV extracts the voiceprint features corresponding to the audio through a pre-created recognition model. Based on the voiceprint features, it determines the second scene information corresponding to the audio. It can be understood that the voiceprint features of the audio record the key information of the audio. The recognition model extracts the voiceprint features of the audio and analyzes the second scene information corresponding to the audio based on the voiceprint features. For example, if the voiceprint features are identified as matching those of a piano, it can be determined that the second scene information may be a piano performance. If the voiceprint features are identified as matching those of a basketball hitting the floor, it can be determined that the second scene information may be a basketball program.
[0092] Furthermore, when the recognition model detects human voices in the audio, it can convert the voices into text, and then determine the second scene information by combining the recognition of the text with the recognition of voiceprint features.
[0093] Step e: Determine the current recognition result based on the first scene information and the second scene information.
[0094] In this step, the smart TV uses a recognition model to identify screenshots and audio corresponding to programs played within a preset time period to obtain first and second scene information. Based on this information, the current recognition result is determined. It's understood that the first and second scene information obtained by the smart TV each include multiple different scenes. The smart TV statistically analyzes these multiple scenes to determine the current recognition result, which includes the percentages of each different scene. For example, the current recognition result might be 80% for basketball programs, 10% for news programs, 5% for music programs, and 5% for science and education programs.
[0095] Specifically, the steps for determining user habits based on the current identification results and historical identification results include:
[0096] Step f: Obtain the previous preset number of historical recognition results corresponding to the current recognition result, and calculate the similarity between the current recognition result and the previous preset number of historical recognition results;
[0097] In this step, the smart TV calculates the similarity between the current recognition result and the previous preset number of historical recognition results, based on the current recognition result. This preset number is typically set in advance within the smart TV and is usually two. When the smart TV obtains the current recognition result, it retrieves the two previous historical recognition results and calculates the similarity between them. For example, if the smart TV obtained historical recognition results from July 20th to July 25th and then obtained the current recognition result on July 26th, with a preset number of two, the smart TV retrieves the historical recognition results from July 24th and July 25th, calculates the similarity between these historical results and the current recognition result obtained on July 26th.
[0098] Step g: If the similarity between the current recognition result and the previous preset number of historical recognition results is greater than the preset similarity threshold, then determine the user's usage habits based on the current recognition result and the previous preset number of historical recognition results.
[0099] In this step, the smart TV calculates the similarity between the current recognition result and a preset number of historical recognition results, and compares this similarity with a preset similarity threshold. If the similarity between the current recognition result and the preset number of historical recognition results is greater than the preset similarity threshold, then the user's usage habits are determined based on the current recognition result and the preset number of historical recognition results. For example, if the current recognition result is 80% for basketball programs, 10% for news programs, and 10% for music programs, and two historical recognition results are 82% for basketball programs, 10% for news programs, and 8% for music programs, and 80% for basketball programs, 12% for news programs, and 8% for music programs, the smart TV can determine that the user's usage habit is that they like watching basketball programs.
[0100] Understandably, smart TVs need to determine that the similarity between multiple consecutive recognition results is greater than a preset similarity threshold before they can determine user habits based on multiple consecutive recognition results, in order to avoid the impact of an inaccurate user habit determination due to a single accidental recognition result.
[0101] Step S30: Determine the virtual character image based on the user's usage habits, and generate a virtual character based on the virtual character image.
[0102] In this embodiment, after determining the user's usage habits, the smart TV creates a corresponding virtual character image based on the user's usage habits, and generates a virtual character based on the virtual character image.
[0103] Furthermore, since a family has more than one member, smart TVs can bind user habits to corresponding family members. For example, in a family, the father's user habit is to like watching basketball programs, while the child's user habit is to like watching cartoon programs. The smart TV can determine the corresponding virtual character image based on the user habits of different family members and generate virtual characters based on the virtual character image.
[0104] Specifically, step S30 includes:
[0105] Step h: Input the user's usage habits into the pre-created virtual character image management module, and determine the virtual character image based on the user's usage habits through the virtual character image management module;
[0106] Step i, obtain the user image corresponding to the user's usage habits;
[0107] Step j: Send the virtual character image and the user image to the cloud-based virtual character creation terminal. The cloud-based virtual character creation terminal generates a virtual character based on the virtual character image and the user image, and then sends the virtual character back to the virtual character image management module.
[0108] In steps h to j, the smart TV inputs user habits into a pre-created virtual avatar management module. This module determines the virtual avatar based on these habits; for example, if a user enjoys watching basketball programs, the virtual avatar will be a basketball image. The smart TV then uses its camera module to capture an image of the user corresponding to these habits. This image and the virtual avatar are sent to a cloud-based virtual avatar creation terminal. The terminal then uses the image to obtain the user's facial features and generates a virtual avatar based on the image and the user's features. This virtual avatar is then sent back to the virtual avatar management module for use by the smart TV. It should be noted that the virtual avatar creation terminal can also be directly integrated into the smart TV.
[0109] Further, after step S30, the following is included:
[0110] Monitor push notifications and display them to users through the virtual character.
[0111] In this step, after generating a virtual character, the smart TV can determine the current user of the TV through the camera module, call the corresponding virtual character based on the user, and when a push notification is detected, the virtual character will display the push notification to the user via voice. The virtual character can also receive the user's voice, determine the user's intention based on the voice, and then control the smart TV to complete the action corresponding to the user's intention, thus realizing the interaction between the virtual character and the user.
[0112] In this embodiment of the virtual character generation method, when a user is watching a program, the smart TV acquires the current time period and compares it with a preset time period. If the current time period is within the preset time period, the smart TV performs screenshot and audio recording operations on the program played in the current time period according to a preset acquisition frequency, obtaining screenshots and audio corresponding to the program played within the preset time period. The smart TV extracts key pixels from the screenshot using a pre-created recognition model and determines the first scene information corresponding to the screenshot based on the key pixels. The smart TV extracts voiceprint features corresponding to the audio using a pre-created recognition model and determines the second scene information corresponding to the audio based on the voiceprint features. The smart TV determines the current recognition result based on the first scene information and the second scene information. The smart TV acquires a preset number of previous historical recognition results corresponding to the current recognition result and calculates the similarity between the current recognition result and the preset number of previous historical recognition results. If the similarity between the current recognition result and the preset number of previous historical recognition results is greater than a preset similarity threshold, the user's usage habits are determined based on the current recognition result and the preset number of previous historical recognition results. The smart TV inputs user habits into a pre-created virtual avatar management module, which then determines the virtual avatar based on these habits. The smart TV acquires a user image corresponding to these habits and generates a virtual avatar based on the user image and the virtual avatar image. This invention uses a recognition model to identify screenshots and audio from programs played within a preset time period to determine user habits. Based on these habits, a virtual avatar image is determined and generated, ensuring the virtual avatar meets user needs and thus increasing user engagement and improving the user experience.
[0113] Furthermore, such as Figure 3 As shown, based on the first embodiment of the virtual character generation method of the present invention, a second embodiment of the virtual character generation method of the present invention is proposed.
[0114] The second embodiment of the virtual character generation method differs from the first embodiment in that, prior to the step of determining user habits based on the current recognition result and historical recognition results, it includes:
[0115] Step k: Obtain the scene proportion set of the current recognition results, and compare the scene proportion set with the preset confidence level;
[0116] Step 1: If there is a scene proportion in the scene proportion set that is greater than the preset confidence level, then store the current recognition result and execute the following step: determine the user's usage habits based on the current recognition result and the historical recognition result;
[0117] Step m: If there is no scene proportion in the scene proportion set that is greater than the preset confidence level, delete the current recognition result and re-execute step m: obtain the screenshots and audio corresponding to the programs played within the preset time period.
[0118] In this embodiment, after obtaining the current recognition result, the smart TV acquires the scene proportion set of the current recognition result and compares the scene proportion set with a preset confidence level. If it is determined that there is a scene proportion in the scene proportion set that is greater than the preset confidence level, the current recognition result is stored, and the user's usage habits are determined based on the current recognition result and historical recognition results, as well as subsequent steps. If it is determined that there is no scene proportion in the scene proportion set that is greater than the preset confidence level, the current recognition result is deleted, and the acquisition of screenshots and audio corresponding to the programs played within the preset time period and subsequent steps are re-executed. For example, assuming the current recognition result is 90% for basketball programs, 5% for news programs, and 5% for music programs, with a preset confidence level of 85%, the smart TV compares the set of scene proportions of the current recognition result with the preset confidence level. If there is a scene proportion in the set that is greater than the preset confidence level, the smart TV can determine that the current recognition result is valid, store the current recognition result, and proceed with subsequent steps. Conversely, assuming the current recognition result is 80% for basketball programs, 10% for news programs, and 10% for music programs, with a preset confidence level of 85%, the smart TV compares the set of scene proportions of the current recognition result with the preset confidence level. If there is no scene proportion in the set that is greater than the preset confidence level, the smart TV can determine that the current recognition result is invalid, delete the current recognition result, and wait for the next preset time period to re-execute the steps of obtaining screenshots and audio corresponding to the programs played within the preset time period. By analyzing the current recognition result and determining its validity, the accuracy of determining user habits is improved, which in turn helps improve the accuracy of determining the virtual character's image.
[0119] Furthermore, if the stored current recognition result is not used to determine user habits in subsequent steps, it is stored in the smart TV as a historical recognition result.
[0120] In this embodiment, after obtaining the current recognition result, the smart TV acquires a set of scene proportions for the current recognition result and compares this set with a preset confidence level. If it is determined that there is a scene proportion in the set that is greater than the preset confidence level, the current recognition result is stored, and the user's usage habits are determined based on the current recognition result and historical recognition results, along with subsequent steps. If it is determined that there is no scene proportion in the set that is greater than the preset confidence level, the current recognition result is deleted, and the process of acquiring screenshots and audio corresponding to programs played within a preset time period and subsequent steps is repeated. By analyzing the current recognition result and determining its validity, the accuracy of determining user usage habits is improved, which in turn helps improve the accuracy of determining the virtual character's image.
[0121] like Figure 4 As shown, the present invention also provides a virtual character generation device. The virtual character generation device of the present invention includes:
[0122] The acquisition module 101 is used to acquire screenshots and audio corresponding to programs played within a preset time period;
[0123] The recognition module 102 is used to recognize the screenshot and the audio through a pre-created recognition model, obtain the current recognition result, and determine the user's usage habits based on the current recognition result and the historical recognition result;
[0124] The generation module 103 is used to determine the virtual character image based on the user's usage habits, and generate a virtual character based on the virtual character image.
[0125] Furthermore, the acquisition module is also used for:
[0126] Get the current time period and compare it with the preset time period;
[0127] If the current time period is within the preset time period, then according to the preset acquisition frequency, screenshot and audio recording operations are performed on the program played in the current time period to obtain the screenshot and audio corresponding to the program played within the preset time period.
[0128] Furthermore, the identification module is also used for:
[0129] Key pixels in the screenshot are extracted using a pre-created recognition model, and the first scene information corresponding to the screenshot is determined based on the key pixels.
[0130] The voiceprint features corresponding to the audio are extracted by a pre-created recognition model, and the second scene information corresponding to the audio is determined based on the voiceprint features.
[0131] The current recognition result is determined based on the first scene information and the second scene information.
[0132] Furthermore, the identification module is also used for:
[0133] Obtain the scene proportion set of the current recognition results, and compare the scene proportion set with the preset confidence level;
[0134] If there is a scene proportion in the scene proportion set that is greater than the preset confidence level, then the current recognition result is stored and the following steps are executed: determine the user's usage habits based on the current recognition result and the historical recognition result;
[0135] If there is no scene proportion in the scene proportion set that is greater than the preset confidence level, then delete the current recognition result and re-execute the step: obtain the screenshots and audio corresponding to the programs played within the preset time period.
[0136] Preferably, the identification module further includes a determination module, the determination module being used to:
[0137] Obtain the previous preset number of historical recognition results corresponding to the current recognition result, and calculate the similarity between the current recognition result and the previous preset number of historical recognition results;
[0138] If the similarity between the current recognition result and the previous preset number of historical recognition results is greater than a preset similarity threshold, then the user's usage habits are determined based on the current recognition result and the previous preset number of historical recognition results. Preferably, the generation module is further used for:
[0139] The user's usage habits are input into a pre-created virtual character management module, which then determines the virtual character based on those habits.
[0140] Obtain the user image corresponding to the user's usage habits;
[0141] The virtual character image and the user image are sent to the cloud-based virtual character creation terminal, which generates a virtual character based on the virtual character image and the user image, and then sends the virtual character back to the virtual character image management module.
[0142] Preferably, the generation module further includes a display module, the display module being used for:
[0143] Monitor push notifications and display them to users through the virtual character.
[0144] The present invention also provides a virtual character generation device.
[0145] The virtual character generation device of the present invention includes: a memory, a processor, and a virtual character generation program stored in the memory and executable on the processor. When the virtual character generation program is executed by the processor, it implements the steps of the virtual character generation method described above.
[0146] The method implemented when the virtual character generation program running on the processor is executed can be referred to in various embodiments of the virtual character generation method of the present invention, and will not be repeated here.
[0147] The present invention also provides a computer-readable storage medium.
[0148] The present invention provides a computer-readable storage medium storing a virtual character generation program, which, when executed by a processor, implements the steps of the virtual character generation method described above.
[0149] The method implemented when the virtual character generation program running on the processor is executed can be referred to in various embodiments of the virtual character generation method of the present invention, and will not be repeated here.
[0150] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0151] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0153] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A virtual character generation method characterized by comprising: The virtual character generation method comprises the following steps: Obtain the screenshot and audio corresponding to the program played in a preset period; Identify the screenshot and the audio through a pre-created identification model to obtain a current identification result, and determine a user usage habit according to the current identification result and a historical identification result; The step of determining the user usage habit according to the current identification result and the historical identification result comprises: Obtain a preset number of historical identification results corresponding to the current identification result, and calculate the similarity between the current identification result and the preset number of historical identification results; If the similarity between the current identification result and the preset number of historical identification results is greater than a preset similarity threshold, determine the user usage habit according to the current identification result and the preset number of historical identification results; Determine a virtual character image according to the user usage habit, and generate a virtual character according to the virtual character image; The step of determining the virtual character image according to the user usage habit and generating the virtual character according to the virtual character image comprises: Input the user usage habit into a pre-created virtual character image management module, and determine a virtual character image according to the user usage habit through the virtual character image management module; Obtain a user image corresponding to the user usage habit; Send the virtual character image and the user image to a cloud virtual character production end, generate a virtual character according to the virtual character image and the user image through the cloud virtual character production end, and return the virtual character to the virtual character image management module; After the step of determining the virtual character image according to the user usage habit and generating the virtual character according to the virtual character image, the method comprises the following steps: Monitor push information, and display the push information to the user through the virtual character; The step of monitoring the push information and displaying the push information to the user through the virtual character comprises: Determine a user through a camera module, call a virtual character image in the virtual character image management module according to the user, and display the push information to the user through the virtual character; Receive the voice of the user, determine a user intention according to the voice of the user, and complete an interactive operation through the virtual character image based on the user intention.
2. The virtual character generation method of claim 1, wherein, The step of obtaining the screenshot and audio corresponding to the program played in a preset period comprises: Obtain a current period, and compare the current period with a preset period; If the current period is within the preset period, perform a screenshot operation and a recording operation on the program played in the current period according to a preset acquisition frequency, to obtain the screenshot and the audio corresponding to the program played in the preset period.
3. The virtual character generation method of claim 1, wherein, The step of identifying the screenshot and the audio through a pre-created identification model to obtain a current identification result comprises: Extract key pixel points in the screenshot through a pre-created identification model, and determine first scene information corresponding to the screenshot according to the key pixel points; extracting a voiceprint feature corresponding to the audio through a pre-created recognition model, and determining second scene information corresponding to the audio according to the voiceprint feature; determining a current recognition result according to the first scene information and the second scene information.
4. The virtual character generation method of claim 1, wherein, Before the step of determining the user usage habit according to the current recognition result and the historical recognition result, the method comprises: obtaining a scene proportion set of the current recognition result, and comparing the scene proportion set with a pre-set confidence; if there is a scene proportion greater than the pre-set confidence in the scene proportion set, storing the current recognition result, and performing the step of determining the user usage habit according to the current recognition result and the historical recognition result; if there is no scene proportion greater than the pre-set confidence in the scene proportion set, deleting the current recognition result, and re-performing the step of obtaining the screenshot and the audio corresponding to the program played in the preset time period.
5. A virtual character generation apparatus characterized by comprising: The virtual character generation device comprises: an obtaining module configured to obtain a screenshot and audio corresponding to a program played in a preset time period; an identification module configured to identify the screenshot and the audio through a pre-created recognition model to obtain a current recognition result, and determine a user usage habit according to the current recognition result and a historical recognition result; wherein the identification module is further configured to obtain a preset number of historical recognition results corresponding to the current recognition result, and calculate a similarity between the current recognition result and the preset number of historical recognition results; if the similarity between the current recognition result and the preset number of historical recognition results is greater than a preset similarity threshold, determining the user usage habit according to the current recognition result and the preset number of historical recognition results; a generation module configured to determine a virtual character image according to the user usage habit, and generate a virtual character according to the virtual character image; wherein the generation module is further configured to input the user usage habit into a pre-created virtual character image management module, and determine a virtual character image according to the user usage habit through the virtual character image management module; obtain a user image corresponding to the user usage habit; send the virtual character image and the user image to a cloud virtual character production end, generate a virtual character according to the virtual character image and the user image through the cloud virtual character production end, and return the virtual character to the virtual character image management module; monitor push information, and display the push information to the user through the virtual character; wherein the generation module is further configured to determine a user through a camera module, call a virtual character image in the virtual character image management module according to the user, and display the push information to the user through the virtual character; receive a voice of the user, determine a user intention according to the voice of the user, and complete an interactive operation through the virtual character image based on the user intention.
6. A virtual character generation device characterized by comprising: The virtual character generation device comprises a memory, a processor, and a virtual character generation program stored on the memory and executable on the processor, and the virtual character generation program, when executed by the processor, implements the steps of the virtual character generation method according to any one of claims 1 to 4.
7. A computer readable storage medium characterized in that, The computer readable storage medium stores a virtual character generation program, and the virtual character generation program, when executed by the processor, implements the steps of the virtual character generation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Content pushing method and system based on scene recognition and intelligent terminal
CN111416995A
Virtual character generation method and device and computer readable storage medium
CN114758040A