Multimedia information display method, device, and storage medium
By obtaining the text information of multimedia audio data and the facial information of the people in the car, a virtual character image is generated, which solves the problem of the single multimedia audio display method and improves the user experience.
Patent Information
- Application Number
- CN202111402906.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing multimedia audio display method is single, resulting in poor user experience.
By obtaining text information from multimedia audio data, extracting reference keywords, and combining it with facial photography information of people in the car, a virtual character image is generated.
The generated virtual characters are closer to users, which improves the user experience and increases the intimacy and fun of the audience.
Smart Images

Figure CN116166822B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimedia technology, and in particular to a method and device for displaying multimedia information, and a storage medium. Background Art
[0002] As the pace of modern life becomes faster and faster, people rarely have the time or mood to sit down and read traditional paper books carefully. So some people read the book aloud and record it for people who want to read. This is audiobooks. Through pre-recording, it can be played in various media, and the playback speed can be adjusted, and the stop time of playback can be automatically remembered, which greatly facilitates the majority of readers.
[0003] With the popularity of audiobooks, the types of audiobooks they cover are increasing. At the same time, the audience's appreciation level and requirements are constantly improving. However, the current simple playback methods are too homogeneous and relatively boring, making it difficult to satisfy the audience. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defect in the prior art that users have a poor experience when listening to multimedia audio such as audiobooks due to the single multimedia audio presentation method, and to provide a method, device and storage medium for presenting multimedia information.
[0005] The present invention solves the above technical problems through the following technical solutions:
[0006] The present invention provides a method for displaying multimedia information, the method comprising:
[0007] Obtain text information corresponding to multimedia audio data;
[0008] extracting reference keywords from the text information;
[0009] Obtain facial image information of people in the car;
[0010] Acquire rendering element information according to the reference keyword and the face shooting information;
[0011] A virtual character image is generated based on the rendering element information.
[0012] The present invention also provides a multimedia information display device, comprising: a display unit, one or more processing units, and a storage unit, wherein the display unit and the storage unit are respectively communicatively connected to the processing unit; the storage unit is configured to store instructions, and when the stored instructions are executed by the one or more processing units, the one or more processing units execute the following steps:
[0013] Obtain text information corresponding to multimedia audio data;
[0014] extracting reference keywords from the text information;
[0015] Obtain facial image information of people in the car;
[0016] Acquire rendering element information according to the reference keyword and the face shooting information;
[0017] generating a virtual character image based on the rendering element information;
[0018] The display unit is used to display the virtual character image.
[0019] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-mentioned multimedia information display method when executed by a processor.
[0020] The positive progressive effect of the present invention is that the multimedia information display method and device, and storage medium of the present invention obtain rendering element information and generate virtual character images by combining reference keywords in the multimedia audio data text information and facial shooting information of the people in the car, which can make the multimedia information closer to the user during the display process, and the generated virtual character image is more interesting, so that the audience can see a lifelike virtual character image that is full of intimacy while listening to the multimedia information, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flowchart of a multimedia information display method according to embodiment 1 of the present invention.
[0022] Figure 2 This is a module diagram of a multimedia information display device according to embodiment 2 of the present invention. DETAILED DESCRIPTION
[0023] The present invention is further described below by way of examples, but the present invention is not limited to the scope of the examples.
[0024] Example 1
[0025] See also Figure 1 As shown, this embodiment specifically provides a method for displaying multimedia information, and the display method includes:
[0026] S1. Get the text information corresponding to the multimedia audio data;
[0027] S2. extracting reference keywords from text information;
[0028] S3. Obtain facial information of people in the car;
[0029] S4. Obtain rendering element information based on reference keywords and face shooting information;
[0030] S5. Generate a virtual character image based on the rendering element information.
[0031] Audiobooks are a multimedia form that exists and is disseminated in audio form. They have extremely rich content and usually include chapter novels, poems and essays, documentaries, crosstalk sketches, etc. recorded in audio form; the text information corresponding to the multimedia audio data in step S1 is the corresponding text of the audiobook including the above-mentioned various forms, such as the original text of the chapter novel, the dialogue lines of the crosstalk sketch, etc.
[0032] As an optional implementation, step S1 includes:
[0033] S11. Get multimedia audio data;
[0034] S12. Perform a speech recognition operation on the multimedia audio data to obtain text information.
[0035] In this embodiment, step S1 obtains text information through voice recognition, and step S11 obtains multimedia audio data, which may include but is not limited to obtaining playback sound information in real time, such as using a directional microphone to obtain multimedia audio data played in a speaker. Step S12 performs voice recognition operations on the multimedia data, such as using NLP (Natural Language Processing) technology through word segmentation, part-of-speech tagging, entity recognition, etc. to obtain corresponding text information. Step S12 can also automatically determine the content of the multimedia audio data based on the recognition results after obtaining part of the text information, thereby obtaining the multimedia audio data from the corresponding server. For example, it is recognized that the multimedia audio data is "Romance of the Three Kingdoms". Of course, the virtual character image finally generated in this embodiment is used for synchronous display with the multimedia audio data, so the process of generating the virtual character image still needs to be based on the playback progress of the multimedia audio data.
[0036] As an optional implementation, step S1 includes:
[0037] S13. Calling an application program interface of a preset application program to obtain text information corresponding to the multimedia audio data; the application program is used to send the multimedia audio data.
[0038] This embodiment provides another method for obtaining text information corresponding to multimedia audio data. Specifically, since some applications (APPs) that play audiobooks provide an interface for providing text information corresponding to the playback content, step S13 obtains relevant text information, such as the multimedia audio corresponding text work or creative blueprint, by calling the corresponding interface in the application.
[0039] As an optional implementation, step S0 is included before step S1 to detect whether a preset application is started, and the preset application is used to send multimedia audio data; if it is started, step S1 is executed.
[0040] This embodiment pre-detects whether the application for playing audiobooks is started. If it is not started, it indicates that the current audiobook is not in the playback mode, and relevant voice recognition, text search, etc. will be invalid. If it is already started, step S1 is executed to perform the corresponding subsequent steps.
[0041] Step S2 extracts reference keywords from the text information, which can be achieved by comparing with preset keywords, and obtaining words in the text information that are identical or similar to the preset keywords or their combinations as reference keywords. The preset keywords can be set according to the rendering requirements of the virtual character image, and preferably can include name information, action information, location information, time information, preset proper noun information, etc. in the content of the multimedia audio data. Those skilled in the art will understand that based on the above-mentioned feature word information, the virtual character image can be rendered in a targeted manner to make it consistent with the playback content of the multimedia audio and the visual effects of the scene involved in the content. For example, in the audiobook "Romance of the Three Kingdoms", the reference name information reference keyword "Cao Cao"; the action information reference keyword "arrival"; the time information reference keyword "the fourteenth year of Jian'an, i.e. July 209 AD"; and the location information reference keyword "Hefei" are extracted.
[0042] Step S3 obtains facial information of the vehicle occupants. The occupants can be the driver or passengers. Video of the occupants can be captured using cameras inside the vehicle. For example, when the driver enters the vehicle, an image of the driver's head is captured to obtain the driver's head features. The video capture can be set to a preset duration, such as one minute. Cameras can be installed at the rearview mirror or on the back of the seat headrest to capture video of the driver or passenger, including headshot video information. Other methods are also possible, such as having the occupants send a video containing headshot video information to the vehicle computer. Based on the captured video of the occupants, facial information of the occupants is extracted. Specifically, several frames of facial images are captured from the video. During processing, preprocessing can be performed using techniques such as histogram equalization, and specific areas can be eliminated using image statistical characteristics to obtain facial images. Furthermore, the face is tracked based on its position and velocity within the multiple frames. The face clarity is determined using a fast discrete Fourier transform, and the facial information with the best clarity is selected.
[0043] Step S4 obtains rendering element information according to the reference keywords and facial shooting information obtained above; Step S5 generates a virtual character image based on the rendering element information.
[0044] In an optional embodiment, the rendering element information includes facial posture information and human body posture information; step S4 includes S41: according to the action information, obtaining matching human body posture information in the offline material database, and generating corresponding facial posture information according to the facial shooting information; step S5 includes S51: generating a virtual character image based on the facial posture information and human body posture information.
[0045] Facial pose information and body pose information are the primary rendering elements that contribute to the avatar's motion and gestures. Facial pose is achieved using technologies including, but not limited to, face detection, which locates key facial regions within a facial image. Head pose can be determined using pitch, yaw, and roll angles; the pitch angle represents the angle of head movement, the yaw angle represents the angle of head shaking, and the roll angle represents the angle of head turning. Furthermore, facial pose includes the angle of sight, for example, represented by the angle with the horizontal. The expression of a facial pose can be described using a facial motion coding system. Deformation units are defined based on the types and motion characteristics of facial muscles. Various facial expressions can ultimately be decomposed into corresponding deformation units and analyzed for expression feature information. These expressions are typically defined to correspond to the six basic emotions: anger, happiness, sadness, surprise, disgust, and fear. Expression recognition involves feature extraction based on the aforementioned image acquisition and preprocessing. The resulting facial deformation unit coding sequence can be used to characterize the expression of the facial pose.
[0046] Human posture can be achieved through network models including but not limited to GAN (Generative Adversarial Networks). Specifically, skeleton-assisted joint motion generation is achieved through a skeleton-guided portrait generation principle, based on a conditional GAN infrastructure, with the portrait and target skeleton as conditions. A coarse-to-fine framework for pose-guided portrait generation is adopted, which consists of a coarsening stage and a refinement stage. In addition, the human image generation model can be entangled to further improve the quality of the results by using a decomposition strategy. The deformable GAN model for pose-based human image generation attempts to alleviate the misalignment problem between different poses by using affine transformations on coarse rectangular areas, and synthesizes human images by reconstructing shapes through labels to ensure the consistency of human body structure.
[0047] The human body posture information in this embodiment can be matched and obtained in the offline material database based on the obtained action information, and the facial posture information is obtained based on the above-mentioned facial shooting information. Step S5 generates a virtual character image based on the facial posture information and the human body posture information, which can effectively combine the action information extracted from the multimedia audio data with the facial information of the occupants of the vehicle. The obtained rendering information is used to generate a virtual character image, which can bring a sense of intimacy to the occupants of the vehicle. At the same time, the dynamic virtual character image and the multimedia audio data are matched, which is interesting.
[0048] In an optional embodiment, the rendering element information includes clothing style information; step S4 includes S42: obtaining matching clothing style information in the offline material database according to the name information and / or time information; step S5 includes S52: generating a virtual character image based on the clothing style information.
[0049] Clothing style information is used to render the clothing of the virtual character, such as modern clothing, ancient clothing, men's clothing, women's clothing, etc. Step S42 obtains matching clothing style information based on the person's name information, time information, or a combination thereof. For example, if the time information is "the 14th year of Jian'an" and the person's name is "Cao Cao," then the clothing style information of ancient male officials can be matched. Specifically, this can be obtained from the clothing style database in the offline material database. Step S52 generates a virtual character image based on the clothing style information, mainly through human body analysis of the virtual character, that is, estimating the reasonable human body analysis of the target image based on the approximate shapes of the body parts and the target posture, effectively guiding the synthesis of precise areas of human body parts; secondly, through human body segmentation, that is, using a human analyzer to calculate the human segmentation map, so that the clothing corresponds to the position and shape of different body parts such as arms or torso; in addition, since changes in human posture will cause different deformations of clothing, the posture information can be explicitly modeled using a posture estimator of the part similarity field, and the human posture is calculated as the coordinates of several key points. Using spatial layout, each key point is further converted into a heat map, and the heat map is further superimposed into a heat map of the channel posture, so as to adjust the clothing deformation and generate a virtual character image.
[0050] In an optional implementation, the rendering element information includes character prototype information; step S4 includes S43: obtaining matching character prototype information in an offline material database according to the person name information; and step S5 includes S53: generating a virtual character image based on the character prototype information.
[0051] Step S43 obtains matching character prototype information from an offline material database based on the name information. The offline material database may be a rendering material database including several character prototype templates, and obtains matching character prototype information therein, such as the "Guan Yu" character prototype template; Step S53 generates a virtual character image based on the character prototype information, thereby performing targeted rendering by identifying the name information, so that the obtained virtual character image restores the character image appearing in the multimedia audio data.
[0052] The multimedia information display method of this embodiment obtains rendering element information and generates a virtual character image by combining reference keywords in the text information of the multimedia audio data and the facial shooting information of the people in the car. This can make the multimedia information closer to the user during the display process, and the generated virtual character image is more interesting, so that the audience can see a lifelike virtual character image that is friendly and lifelike while listening to the multimedia information, thereby improving the user experience.
[0053] Example 2
[0054] This embodiment specifically provides a multimedia information display device, comprising: a display unit, one or more processing units, and a storage unit, wherein the display unit and the storage unit are respectively communicatively connected to the processing unit; the storage unit is configured to store instructions, and when the stored instructions are executed by the one or more processing units, the one or more processing units execute the multimedia information display method of Example 1. The display unit is used to display the virtual character image.
[0055] It should be noted that the multimedia information display device in this embodiment can be a separate chip, chip module or network device, or a chip or chip module integrated into a network device; the various modules / units included in the communication data processing device can be software modules / units, hardware modules / units, or partly software modules / units and partly hardware modules / units. For example, for various devices and products applied to or integrated into a chip, the various modules / units included therein can all be implemented in the form of hardware such as circuits, or at least some of the modules / units can be implemented in the form of software programs, which run on the processor integrated inside the chip, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits; for various devices and products applied to or integrated into a chip module, the various modules / units included therein can all be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component of the chip module (such as a chip, circuit module, etc.) or in different components, or at least some of the modules / units can be implemented in the form of hardware such as circuits. The element can be implemented in the form of a software program, which runs on the processor integrated inside the chip module, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits; for various devices and products applied to or integrated in the terminal, the various modules / units contained therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (for example, chip, circuit module, etc.) or different components in the terminal, or, at least some modules / units can be implemented in the form of a software program, which runs on the processor integrated inside the terminal, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits.
[0056] See also Figure 2 As shown, as a preferred implementation, this embodiment specifically provides a multimedia information display device 30, including a processor 31, a memory 32, a computer program stored in the memory 32 and executable on the processor 31, and a display unit 37. When the processor 31 executes the program, the multimedia information display method in Example 1 is implemented. Figure 2The display device 30 for displaying multimedia information is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0057] The multimedia information display device 30 may be implemented as a general-purpose computing device, such as a server device. The components of the multimedia information display device 30 may include, but are not limited to, the at least one processor 31, the at least one memory 32, a display unit 37, and a bus 33 connecting various system components (including the memory 32 and the processor 31).
[0058] The bus 33 includes a data bus, an address bus, and a control bus.
[0059] The memory 32 may include a volatile memory, such as a random access memory (RAM) 321 and / or a cache memory 322 , and may further include a read-only memory (ROM) 323 .
[0060] The memory 32 may also include a program / utility 325 having a set (at least one) of program modules 324, such program modules 324 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0061] The processor 31 executes various functional applications and data processing by running computer programs stored in the memory 32, such as the multimedia information display method in embodiment 1 of the present invention.
[0062] The multimedia information presentation device 30 can also communicate with one or more external devices 34 (e.g., a keyboard, pointing device, etc.). This communication can occur via an input / output (I / O) interface 35. Furthermore, the model-generating device 30 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 36. The network adapter 36 communicates with other modules of the model-generated multimedia information presentation device 30 via a bus 33. Other hardware and / or software modules can be used in conjunction with the model-generating device 30, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, RAID (RAID) systems, tape drives, and data backup storage systems.
[0063] It should be noted that while the detailed description above refers to several units / modules or sub-units / modules of the multimedia information display device, this division is merely exemplary and not mandatory. In practice, according to embodiments of the present invention, the features and functions of two or more units / modules described above may be embodied in a single unit / module. Conversely, the features and functions of a single unit / module described above may be further divided and embodied by multiple units / modules.
[0064] The multimedia information display device of this embodiment specifically obtains feature words and user preference attribute information in media audio data, and effectively combines them to generate a virtual character image that corresponds to the content involved in the media audio data, so that the user can see the above-mentioned virtual character image while listening to the media audio data, making media audio data such as audiobooks no longer monotonous, thereby greatly improving the user's listening experience.
[0065] The multimedia information display device of this embodiment obtains rendering element information and generates a virtual character image by combining reference keywords in the text information of the multimedia audio data and the facial shooting information of the people in the car. This can make the multimedia information closer to the user during the display process, and the generated virtual character image is more interesting, so that the audience can see a lifelike virtual character image that is friendly and lifelike while listening to the multimedia information, thereby improving the user experience.
[0066] Example 3
[0067] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the multimedia information display method in Embodiment 1 is implemented.
[0068] The readable storage medium may include, but is not limited to, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0069] In a possible implementation manner, the present invention may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the multimedia information display method in Example 1.
[0070] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on the remote device, or entirely on the remote device. Although the above describes the specific embodiments of the present invention, it should be understood by those skilled in the art that this is only an example, and the scope of protection of the present invention is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, but these changes and modifications fall within the scope of protection of the present invention.
Claims
1. A method for displaying multimedia information, characterized in that: The display method includes: Obtain text information corresponding to multimedia audio data; extracting reference keywords from the text information; Obtain facial image information of people in the car; acquiring rendering element information according to the reference keyword and the face shooting information, wherein the reference keyword includes name information in the content of the multimedia audio data; generating a virtual character image based on the rendering element information; The rendering element information includes character prototype information; the step of obtaining the rendering element information according to the reference keyword and the face shooting information includes: obtaining matching character prototype information in an offline material database according to the name information; The step of generating a virtual character image based on the rendering element information includes: generating a virtual character image based on the character prototype information.
2. The method for displaying multimedia information according to claim 1, wherein: The reference keywords also include at least one of action information, time information, location information, and preset proper noun information in the content of the multimedia audio data.
3. The method for displaying multimedia information according to claim 2, wherein: The rendering element information includes clothing style information; the step of obtaining the rendering element information according to the reference keyword and the face shooting information includes: obtaining matching clothing style information from an offline material database according to the name information and / or time information; The step of generating a virtual character image based on the rendering element information includes: generating a virtual character image based on the clothing style information.
4. The method for displaying multimedia information according to claim 2, wherein: The rendering element information includes facial pose information and human body pose information; the step of obtaining the rendering element information according to the reference keyword and the facial shooting information includes: obtaining matching human body pose information from an offline material database according to the action information, and generating corresponding facial pose information according to the facial shooting information; The step of generating a virtual character image based on the rendering element information includes: generating a virtual character image based on the facial posture information and the body posture information.
5. The method for displaying multimedia information according to claim 1, wherein: The step of obtaining text information corresponding to the multimedia audio data includes: Get multimedia audio data; A speech recognition operation is performed on the multimedia audio data to obtain the text information.
6. The method for displaying multimedia information according to claim 1, wherein: The step of obtaining text information corresponding to the multimedia audio data includes: An application program interface of a preset application program is called to obtain text information corresponding to the multimedia audio data; the application program is used to send the multimedia audio data.
7. The method for displaying multimedia information according to claim 1, wherein: The step of obtaining text information corresponding to the multimedia audio data includes: detecting whether a preset application is started, the preset application being used to send the multimedia audio data; If so, execute the step of obtaining text information corresponding to the multimedia audio data.
8. A multimedia information display device, characterized in that: The display device includes: a display unit, one or more processing units, and a storage unit, wherein the display unit and the storage unit are respectively in communication with the processing unit; the storage unit is configured to store instructions, and when the stored instructions are executed by the one or more processing units, the one or more processing units execute the following steps: Obtain text information corresponding to multimedia audio data; extracting reference keywords from the text information, wherein the reference keywords include name information from the content of the multimedia audio data; Obtain facial image information of people in the car; Acquire rendering element information according to the reference keyword and the face shooting information; generating a virtual character image based on the rendering element information; The display unit is used to display the virtual character image; The rendering element information includes character prototype information; the step of obtaining the rendering element information according to the reference keyword and the face shooting information includes: obtaining matching character prototype information in an offline material database according to the name information; The step of generating a virtual character image based on the rendering element information includes: generating a virtual character image based on the character prototype information.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for presenting multimedia information according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text broadcasting method and device, electronic equipment and storage medium
CN110941954A