Interactive broadcasting method and system
By generating custom voice avatars and TTS voices through cloud-based AI synthesis technology, the problems of limited interaction options and computational resource pressure in in-vehicle human-machine interaction systems are solved, and personalized and emotional multi-scenario human-machine interaction experiences are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHERY AUTOMOBILE CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-05
AI Technical Summary
Existing in-vehicle human-machine interaction systems are limited by the diversity of passengers, making it difficult to provide a wide range of interaction options. Furthermore, localized data processing leads to computational resource pressure and latency issues.
By uploading material data and broadcasting permissions to the cloud platform, cloud-based AI is used to synthesize custom voice images and TTS voices, generating diverse asset packages. Through permission sharing and multi-scenario adaptation, the computing power and memory usage of the vehicle system are reduced.
It enriches the personalized and emotional experience of human-computer interaction, reduces the burden on in-vehicle computing resources, avoids data processing delays and loss, and is suitable for multiple application scenarios.
Smart Images

Figure CN121979474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of intelligent vehicle control, and in particular to an interactive broadcasting method and system. Background Technology
[0002] With the continuous improvement of automotive intelligence and the increasing demand from consumers for in-vehicle experiences, in-vehicle human-machine interaction systems are evolving from simple functional interactions to systems incorporating emotional design. Currently, mainstream in-vehicle systems generally adopt virtual voice assistant solutions based on computer vision and speech synthesis technologies. These systems typically collect passenger facial image data through in-vehicle cameras and use local computing units to extract and model facial features, ultimately generating a virtual avatar associated with the passenger's appearance. In terms of voice interaction, the system relies on text-to-speech (TTS) technology, generating personalized voice output by collecting the voiceprint features of specific users.
[0003] In existing technical solutions, the generation of virtual avatars heavily relies on the real-time computing power of onboard chips, requiring the processing of computationally intensive tasks such as facial recognition, feature extraction, and 3D modeling. Simultaneously, the system must continuously collect and process multimodal data from in-vehicle sensors, including but not limited to facial images captured by high-definition cameras, voice data collected by microphone arrays, and occupant distribution information collected by seat pressure sensors.
[0004] This data collection and processing mechanism leads to two significant limitations: First, the image and voice libraries generated by the system are limited by the diversity of actual passengers, and when the main passengers are fixed, the system can hardly provide a sufficiently rich selection of interactive options; second, the localized data processing flow puts continuous pressure on the vehicle's computing resources, which can easily lead to problems such as computing delays or data loss in complex driving scenarios. Summary of the Invention
[0005] The purpose of this invention is to provide an interactive broadcasting method and system to alleviate the technical problems existing in the prior art.
[0006] In a first aspect, the present invention provides an interactive broadcasting method, comprising: In response to user commands via the graphical interface, the target material data and the target user's broadcasting permissions are acquired and uploaded to the cloud platform; Based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, a target data package is generated for calling different combinations of images, voices, and permissions in the voice broadcast application. Based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, the corresponding voice broadcast image and voice broadcast data are retrieved from the target data packet for synchronous broadcasting.
[0007] In an optional implementation, the step of acquiring and uploading the target material data and the target user's broadcasting permissions to the cloud platform in response to user operation commands on the graphical interface includes: In response to the user's first operation command through the graphical interface, determine the type of target material to be uploaded; Based on the target material type, perform the following operations to obtain target material data: collect the target user's facial image and / or obtain preset image materials, and obtain preset voice materials; When the current network status meets the preset conditions, the target material data is controlled to be uploaded to the cloud platform from the corresponding data entry according to the target material type.
[0008] In an optional implementation, the step of acquiring the target material data and the target user's broadcasting permissions and uploading them to the cloud platform in response to user operation commands on the graphical interface further includes: In response to a second operation command from the user to the graphical interface, at least one of the following permission editing operations is performed to determine broadcast permissions: assigning application permissions for target material data of the corresponding type to at least one target user account associated with the user, and assigning voice broadcast permissions for the corresponding scene to at least one target user account associated with the user. The broadcasting permissions will be uploaded to the cloud platform simultaneously.
[0009] In an optional implementation, the step of generating a target data package for calling different asset combinations of image, voice, and permissions in a voice broadcasting application, based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcasting permissions, includes: Based on the driving signal determined by the audio features in the target material data and the pixel-level features extracted by the preset network model, a voice broadcast image is generated. The preset speech material in the target material data is subjected to speech recognition to extract vocal features, and then trained to generate speech broadcast data; The voice broadcast data and the voice broadcast image are associated, and then combined with the broadcast permissions to obtain the target data packet.
[0010] In an optional implementation, the step of generating a voice broadcast image based on the driving signal determined by the audio features in the target material data and the pixel-level features extracted by a preset network model includes: A pre-trained variational autoencoder is invoked to extract pixel-level features from the face images and / or preset image materials in the target material data, and the pixel-level features are encoded into feature vectors; Based on preset basic motion parameters and / or audio features extracted from preset speech materials in the target material data, a Transformer network is used to map them into a driving signal feature matrix that is adapted to the feature vector; wherein, the driving signal feature matrix is used to adjust the lip movements of the speech broadcast image based on the audio features. The feature vector and the driving signal feature matrix are aligned and fused, and then a diffusion transformer is used to generate continuous animation frames that represent the voice broadcast image.
[0011] In an optional implementation, before the step of synchronously broadcasting the corresponding voice broadcast image and voice broadcast data from the target data packet based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, the method further includes: Verify the integrity and authenticity of the target data packet, and extract it from the verified target data packet according to the current vehicle's adaptation format.
[0012] In an optional implementation, the step of synchronously broadcasting the corresponding voice broadcast image and voice broadcast data from the target data packet, based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, includes: Monitor the real-time target users who log into the current vehicle and the real-time application scenario in which the current vehicle is located; Based on the broadcast permissions, the first voice broadcast image and the first voice broadcast data that the real-time target user is allowed to use in the target data packet are determined; Based on the real-time application scenario, and / or, in response to a third operation command from the user to the graphical interface, the corresponding second voice broadcast image and second voice broadcast data are respectively called from the first voice broadcast image and the first voice broadcast data to perform synchronized audio-visual broadcast.
[0013] Secondly, the present invention provides an interactive broadcasting system, comprising: The upload module responds to user commands via the graphical interface, acquires the target material data and the target user's broadcasting permissions, and uploads them to the cloud platform. The generation module, based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, generates a target data package for calling different combinations of image, voice, and permissions assets in the voice broadcast application; The broadcast module, based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, retrieves the corresponding voice broadcast image and voice broadcast data from the target data packet for synchronous broadcasting.
[0014] Thirdly, the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and capable of running on the processor, wherein the processor executes the program to implement the method as described in any of the foregoing embodiments.
[0015] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed, implements the method described in any of the foregoing embodiments.
[0016] This invention provides an interactive broadcasting method and system. It expands material collection channels by relying on mobile / vehicle-mounted apps, supporting real-time uploading of face scanning, images, and voice materials. Data processing and AI synthesis tasks are handled through a TSP cloud platform, avoiding local computing power and memory consumption in the vehicle's infotainment system. Deep learning technologies such as diffusion models and GANs are employed to extract pixel-level features, vocal characteristics, and audio features from the materials to generate driving signals, constructing a diverse asset package combining image, voice, and permissions. A multi-level permission control mechanism is designed to support vehicle owners in allocating function and scenario permissions, adapting to multi-user collaboration needs. It breaks through the limitations of existing technologies that rely on local vehicle-mounted data collection and training, significantly enriching the selection dimensions of voice images and TTS voices, and creating differentiated emotional interaction experiences through functions such as random switching and permission sharing. Core computing tasks are handled in the cloud, significantly reducing the hardware burden on the vehicle's infotainment system and avoiding data processing delays and errors. It is compatible with multiple application scenarios such as navigation and weather broadcasts, comprehensively improving the personalization, emotionality, and fluency of human-machine interaction in intelligent vehicles.
[0017] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained through the structures particularly pointed out in the description and the drawings.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 A flowchart of an interactive broadcasting method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the functional modules of an interactive broadcasting device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the hardware architecture of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Currently, the image and voice libraries generated by the system are limited by the diversity of actual passengers. When the main passengers are fixed, the system cannot provide a sufficiently rich selection of interactive options. The localized data processing process will put continuous pressure on the vehicle's computing resources, which can easily lead to problems such as computing delays or data loss in complex driving scenarios.
[0023] Based on this, the interactive broadcasting method and system provided by the embodiments of the present invention can synthesize a custom voice image and TTS voice through cloud AI, support the uploading of multiple materials and permission sharing, enrich the emotional and personalized choices of human-computer interaction, reduce the computing power and memory occupation of the vehicle system, and adapt to the broadcasting needs of multiple scenarios.
[0024] To facilitate understanding of this embodiment, a detailed description of an interactive broadcasting method disclosed in this embodiment of the invention will be provided first. This method is applied collaboratively by the TSP cloud platform, user terminal, and TBOX vehicle terminal. The TSP cloud platform is short for Telematics Service Provider cloud platform, which is the core hub in the intelligent vehicle network system that connects user terminal (mobile phone / vehicle APP), vehicle (TBOX vehicle terminal), and backend services. It is specifically responsible for vehicle-related data transmission, processing, storage, and service scheduling.
[0025] Figure 1This is a flowchart of an interactive broadcasting method provided in an embodiment of the present invention.
[0026] Reference Figure 1 The method may include the following steps: S102, in response to user operation commands on the graphical interface, acquires the target material data and the target user's broadcasting permissions and uploads them to the cloud platform.
[0027] The graphical interface (GUI) is the user interface (mobile phone / vehicle infotainment app) that can be used for user interaction. Users can use this GUI to obtain relevant material data as target material data, acquire broadcast permissions for specific target users, and choose to upload relevant material / permission data to the cloud platform.
[0028] In some embodiments, step S102 may be implemented by the following steps: Step 1.1: In response to the user's first operation command on the graphical interface, determine the type of target material to be uploaded.
[0029] Based on user interactions with the graphical interface, such as touching specific controls, buttons, keys, and / or interacting with the graphical interface, including single clicks, double clicks, long presses, swipes, zooms, and other interactive gestures, the system can input a first operation command to determine the type of target material to be uploaded.
[0030] Step 1.2: Based on the target material type, perform the following operations to obtain target material data: collect the target user's facial image and / or obtain preset image materials, and obtain preset voice materials.
[0031] Here, user terminals such as mobile phones / vehicle-mounted apps can add three dedicated material data upload entry points, which respectively support face scanning to collect the target user's face image, obtain preset image materials to be uploaded (including newly taken pictures and images in the image library), and obtain preset voice materials to be uploaded (including voice materials in the folder and voice materials recorded in real time).
[0032] Step 1.3: When the current network status meets the preset conditions, control the target material data to be uploaded to the cloud platform from the corresponding data entry according to the target material type.
[0033] If the network is currently working properly, the user selects the target upload type and submits the corresponding material data (such as celebrity photos, family members' recorded voice messages, etc.). The app captures and temporarily stores the material data in real time and uploads the target material data directly to the TSP cloud platform. If there is a network outage or poor network signal, the app caches the material data and automatically triggers the upload once the network is restored.
[0034] Based on the foregoing embodiments, step S102 further includes: Step 1.4, in response to the user's second operation command for the graphical interface, perform at least one of the following permission editing operations to determine broadcast permissions: assign the application permission of the corresponding type of target material data to at least one target user account associated with the user, and assign the voice broadcast permission of the corresponding scene to at least one target user account associated with the user.
[0035] Users can bind their vehicle owner account to any user terminal, such as a mobile phone / vehicle infotainment app, and then input a second operation command on the graphical interface of this user terminal to edit the aforementioned broadcast permissions. For example, they can select specific types of target material data that can be used by a specific user for broadcast, or restrict certain scenarios in which a specific user can use voice broadcast permissions. Specifically, a user terminal is temporarily bound to the vehicle owner account, and the trust level of voice broadcast through this user terminal is generally low. Therefore, the user terminal's use of preset image materials (such as pre-taken photos of the vehicle owner or vehicle user) is restricted, and the user terminal's voice broadcast function usage permissions in some vehicle application scenarios are also restricted. The use of voice broadcast function in these vehicle application scenarios may include one or more of the following: navigation broadcast, voice interaction, voice control, weather broadcast, scene cube, and scenario mode.
[0036] Step 1.5: Synchronize and upload the broadcast permissions to the cloud platform.
[0037] Here, after the broadcast permissions are edited, the user terminal packages them into a permission data package and uploads it synchronously to the TSP cloud platform.
[0038] S104, based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, generate a target data package for calling different combinations of images, voices, and permissions in the voice broadcast application.
[0039] In some embodiments of the present invention, the user's broadcasting permissions may be used to determine which categories of target material data or which material data from a certain category of target material data can be selected to generate a voice broadcasting image for broadcasting. The voice broadcasting image can be generated based on the target material data. Specific steps include: Step 2.1: Based on the driving signal determined by the audio features in the target material data and the pixel-level features extracted by the preset network model, generate a voice broadcast image.
[0040] For example, firstly, a pre-trained variational autoencoder is invoked to extract pixel-level features from face images and / or preset image materials in the target material data, and the pixel-level features are encoded into feature vectors; Animation generation can be achieved using diffusion models (such as Fantasy Portrait) or generative adversarial networks (GANs) to generate animations from images / faces. First, features of the source material are extracted using a pre-trained variational autoencoder (VAE) and encoded into the latent space.
[0041] Secondly, based on preset basic motion parameters and / or audio features extracted from preset speech materials in the target material data, the Transformer network maps them into a driving signal feature matrix that matches the feature vector. Among them, the driving signal feature matrix is used to adjust the lip movements of the voice broadcast image based on audio features; Here, based on preset basic action parameters (such as natural blinking, slight head turning) or audio features, the data is mapped through a Transformer network to a driving signal feature matrix that matches the feature vector. This driving signal feature matrix is used to generate a speech image that can associate sound and image.
[0042] Next, the feature vectors and driving signal feature matrices are aligned and fused, and then a diffusion transformer (DiT) is used to generate continuous animation frames to represent the voice broadcast image.
[0043] At this point, the generated emoticon can be synchronized with the emoticon audio generated in subsequent steps. For example, the emoticon's lip movements can produce speech, and the corresponding basic movements can be adapted according to the speech content.
[0044] Step 2.2: Perform speech recognition on the preset speech materials in the target material data to extract vocal features, and then train to generate speech broadcast data.
[0045] Whisper speech recognition technology extracts audio features (such as phonemes, intonation, and vocal tone) from preset speech materials and then generates corresponding speech broadcast data.
[0046] Step 2.3: Associate the voice broadcast data with the voice broadcast image, and then combine the broadcast permissions to obtain the target data packet.
[0047] The voice broadcast image with action features such as lip movements is associated and synthesized with custom TTS broadcast voice data generated based on voice features; then, the corresponding synthesized data is selected and bound together with the broadcast permissions to form a complete target data package.
[0048] For example, if a user's desired voice is pre-defined for a certain broadcast permission, then the broadcast permission for that user is bound to the synthesized data corresponding to that desired voice. Similarly, the voice broadcast image of a user in different scenarios is also bound to the corresponding broadcast permission, which will not be elaborated here.
[0049] In practical applications, further verification operations can be performed on the synthesized and bound target data packet before step S106 to ensure the reliability of the voice broadcast image application. The method also includes: The integrity and authenticity of the target data packet are verified, and data is extracted from the verified target data packet according to the current vehicle's adaptation format; here, the extracted data is the voice broadcast image data and the broadcast voice content data after parsing the target data packet.
[0050] The TSP cloud platform sends the target data packet with associated permissions generated in the aforementioned embodiment to the vehicle's TBOX (Telematics Module Unit) via the vehicle network communication link, and simultaneously sends a data transmission completion notification, waiting for the TBOX to provide feedback on the reception status. After receiving the target data packet, the TBOX first verifies the data integrity and the validity of the permission binding; it then parses the core content in the target data packet, such as extracting the animation frame sequence of the custom voice image, the audio encoding data of the TTS broadcast voice, and the permission control information; according to the hardware adaptation requirements of the vehicle system, it converts the animation frame sequence into a format that the vehicle system can render, and converts the TTS audio data into a format that the vehicle system can play.
[0051] S106, based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, retrieves the corresponding voice broadcast image and voice broadcast data from the target data packet for synchronous broadcasting.
[0052] In some embodiments, step S106, which calls the corresponding data to implement synchronized audio-visual voice broadcasting, includes: Step 3.1: Monitor the real-time target users who are logging into the current vehicle and the real-time application scenario in which the current vehicle is located.
[0053] In practical applications, the vehicle infotainment system can monitor the current usage scenario (such as the user starting navigation, checking the weather, triggering voice control, etc.).
[0054] Step 3.2: Based on the broadcast permissions, determine the first voice broadcast image and the first voice broadcast data that the real-time target user is allowed to use in the target data packet.
[0055] Based on the user's broadcast permissions, we can know the resources that the user has the authority to generate broadcast avatars and the scene mapping relationships that broadcasts are allowed to appear in, and confirm that the user has the authority to use custom voice avatars and TTS broadcast voices.
[0056] Step 3.3, based on the real-time application scenario, and / or, in response to the user's third operation command for the graphical interface, call the corresponding second voice broadcast image and second voice broadcast data from the first voice broadcast image and the first voice broadcast data respectively to perform synchronized audio-visual broadcast.
[0057] Here, based on the current real-time application scenario in which the user is located, the user can know from the aforementioned pre-set mapping relationship the second voice broadcast image and second voice broadcast data that can be selected and called in this real-time application scenario, and / or, the user interacts with the graphical interface again to input a third operation command, so as to accurately determine the second voice broadcast image and second voice broadcast data to be selected and called in the current real-time application scenario, and then render the custom voice image on the display interface (synchronously matching the lip movements of the TTS broadcast), and play the custom TTS broadcast voice to complete the human-computer interaction.
[0058] Key scenarios for customizing voice avatars and TTS broadcasting in practical applications include: 1. Reminder Scenario (Customizable on User Terminal): When driving home from get off work or on holiday, family members or friends may remind the driver of things, such as what to buy or carry. At this time, the in-car voice VPA will display the cartoon virtual avatar of the person who reminded you and read out the reminder in their corresponding voice, for example: "XXX, please pick up my package when you get home."
[0059] 2. Blessing Scene (Automatically Presented): While the vehicle is in motion, the VPA avatar can display birthday / holiday wishes from family members; especially wishes using the driver's child's cartoon virtual avatar and voice: Happy Birthday, Dad, good luck at work.
[0060] 3. Long driving time reminder (automatic): When the system detects that the driver has been driving for an extended period of time, the system displays the parents' VPA image and TTS voice, reminding them to drive safely, find a place to rest when necessary, and prioritize safety.
[0061] 4. Pet travel scenarios (automatically displayed): If a family travels with their pet, upon arrival at their destination, the VPA displays a cartoon image of their beloved pet and reminds them in its corresponding voice: Remember to bring a leash, skateboard, or other tools for interacting with your pet.
[0062] 5. Vehicle safety reminders: When a vehicle experiences a safety issue or is due for maintenance, the system can analyze the user's regular and preferred behaviors to identify their favorite celebrities or figures. Based on these preferences, when the system detects a preset problem or maintenance condition, it automatically plays corresponding reminders using the image and voice of that favorite person. For example, if you frequently listen to Eason Chan's songs, the system will capture key information from your listening or video viewing. When the vehicle requires maintenance, it will generate an image and voice of Eason Chan to remind you to change the air filter, tires, etc.
[0063] In some embodiments, if the user triggers a scene switch (such as switching from navigation to weather broadcast), steps 3.1-3.3 above are repeated, and the resources corresponding to the latest scene are re-called from the mapping relationship based on the latest scene to generate a voice broadcast image and voice data that meet the requirements of the latest scene.
[0064] This invention allows users to experience different scenario modes by switching between different voice avatars and TTS (Text-to-Speech) voice prompts. Users can periodically upload and update their favorite characters or other avatars, as well as their voice prompts, to obtain different emotional human-computer interactions. Compared to existing single hardware voice avatars, pre-installed software avatars, or paid unlockable software avatars, this application maximizes the level of emotional engagement. Users can upload multiple source files to generate multiple voice avatars and voice prompts, which will appear randomly at specified times. Users can also transfer editing rights for voice avatars to their family and friends, creating a surprising and differentiated voice emotional interaction experience. This interaction experience can be applied to schedule announcements, weather reports, navigation announcements, and other scenarios involving TTS broadcasts.
[0065] In some embodiments, such as Figure 2 As shown, this embodiment of the invention also provides an interactive broadcasting system, including: The upload module responds to user commands via the graphical interface, acquires the target material data and the target user's broadcasting permissions, and uploads them to the cloud platform. The generation module, based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, generates a target data package for calling different combinations of image, voice, and permissions assets in the voice broadcast application; The broadcast module, based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, retrieves the corresponding voice broadcast image and voice broadcast data from the target data packet for synchronous broadcasting.
[0066] In response to user commands to the graphical interface, this invention acquires target material data and the target user's broadcasting permissions and uploads them to the cloud platform. Based on the vocal features, audio features, driving signals, pixel-level features, and broadcasting permissions determined in the target material data, a target data package with different combinations of images, voices, and permissions that can be called in the voice broadcasting application is generated. According to the current real-time application scenario of the vehicle and the broadcasting permissions corresponding to the real-time target user, the corresponding voice broadcasting image and voice broadcasting data are called from the target data package for synchronous broadcasting, enriching the selection of images and voices for voice broadcasting, reducing the computing power and memory usage of the vehicle system, and enhancing the emotional and personalized experience of human-computer interaction.
[0067] The present invention provides an embodiment of an electronic device. In this embodiment, the electronic device may be, but is not limited to, a personal computer (PC), a laptop computer, a monitoring device, a server, or other computer device with analysis and processing capabilities.
[0068] As an exemplary embodiment, see [link to example]. Figure 3 The electronic device 110 includes a communication interface 111, a processor 112, a memory 113, and a bus 114. The processor 112, the communication interface 111, and the memory 113 are connected via the bus 114. The memory 113 is used to store a computer program that supports the processor 112 in executing the above-described method. The processor 112 is configured to execute the program stored in the memory 113.
[0069] The machine-readable storage medium mentioned in this article can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For example, machine-readable storage media can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0070] Non-volatile media can be non-volatile memory, flash memory, storage drives (such as hard disk drives), any type of storage disk (such as optical discs, DVDs, etc.), or similar non-volatile storage media, or combinations thereof.
[0071] It is understood that the specific operation methods of each functional module in this embodiment can be referred to the detailed description of the corresponding steps in the above method embodiment, and will not be repeated here.
[0072] The computer-readable storage medium provided in the embodiments of the present invention stores a computer program. When the computer program code is executed, it can implement the method described in any of the above embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0073] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and apparatus described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0074] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0075] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0076] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention.
Claims
1. An interactive broadcasting method, characterized in that, include: In response to user commands via the graphical interface, the target material data and the target user's broadcasting permissions are acquired and uploaded to the cloud platform; Based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, a target data package is generated for calling different combinations of images, voices, and permissions in the voice broadcast application. Based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, the corresponding voice broadcast image and voice broadcast data are retrieved from the target data packet for synchronous broadcasting.
2. The method according to claim 1, characterized in that, The steps for acquiring and uploading target material data and the target user's broadcasting permissions to the cloud platform in response to user commands via the graphical interface include: In response to the user's first operation command through the graphical interface, determine the type of target material to be uploaded; Based on the target material type, perform the following operations to obtain target material data: collect the target user's facial image and / or obtain preset image materials, and obtain preset voice materials; When the current network status meets the preset conditions, the target material data is controlled to be uploaded to the cloud platform from the corresponding data entry according to the target material type.
3. The method according to claim 2, characterized in that, The steps of acquiring and uploading target material data and the target user's broadcasting permissions to the cloud platform in response to user operation commands via the graphical interface also include: In response to a second operation command from the user to the graphical interface, at least one of the following permission editing operations is performed to determine broadcast permissions: assigning application permissions for target material data of the corresponding type to at least one target user account associated with the user, and assigning voice broadcast permissions for the corresponding scene to at least one target user account associated with the user. The broadcasting permissions will be uploaded to the cloud platform simultaneously.
4. The method according to claim 1, characterized in that, Based on the driving signals determined by the vocal and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, the steps for generating a target data package for calling different asset combinations of image, voice, and permissions in a voice broadcast application include: Based on the driving signal determined by the audio features in the target material data and the pixel-level features extracted by the preset network model, a voice broadcast image is generated. The preset speech material in the target material data is subjected to speech recognition to extract vocal features, and then trained to generate speech broadcast data; The voice broadcast data and the voice broadcast image are associated, and then combined with the broadcast permissions to obtain the target data packet.
5. The method according to claim 4, characterized in that, The step of generating a voice broadcast image based on the driving signal determined by the audio features in the target material data and the pixel-level features extracted by the preset network model includes: A pre-trained variational autoencoder is invoked to extract pixel-level features from the face images and / or preset image materials in the target material data, and the pixel-level features are encoded into feature vectors; Based on preset basic motion parameters and / or audio features extracted from preset speech materials in the target material data, a Transformer network is used to map them into a driving signal feature matrix that is adapted to the feature vector; wherein, the driving signal feature matrix is used to adjust the lip movements of the speech broadcast image based on the audio features. The feature vector and the driving signal feature matrix are aligned and fused, and then a diffusion transformer is used to generate continuous animation frames that represent the voice broadcast image.
6. The method according to claim 1, characterized in that, Before the step of synchronously broadcasting the corresponding voice broadcast image and voice broadcast data from the target data packet based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, the method further includes: Verify the integrity and authenticity of the target data packet, and extract it from the verified target data packet according to the current vehicle's adaptation format.
7. The method according to claim 1 or 6, characterized in that, Based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, the steps of retrieving the corresponding voice broadcast image and voice broadcast data from the target data packet for synchronous broadcasting include: Monitor the real-time target users who log into the current vehicle and the real-time application scenario in which the current vehicle is located; Based on the broadcast permissions, the first voice broadcast image and the first voice broadcast data that the real-time target user is allowed to use in the target data packet are determined; Based on the real-time application scenario, and / or, in response to a third operation command from the user to the graphical interface, the corresponding second voice broadcast image and second voice broadcast data are respectively called from the first voice broadcast image and the first voice broadcast data to perform synchronized audio-visual broadcast.
8. An interactive broadcasting system, characterized in that, include: The upload module responds to user commands via the graphical interface, acquires the target material data and the target user's broadcasting permissions, and uploads them to the cloud platform. The generation module, based on the driving signal determined by the vocal features and audio features in the target material data, the pixel-level features in the target material data, and the broadcast permissions, generates a target data package for calling different combinations of image, voice, and permissions assets in the voice broadcast application; The broadcast module, based on the current real-time application scenario of the vehicle and the broadcast permissions corresponding to the real-time target user, retrieves the corresponding voice broadcast image and voice broadcast data from the target data packet for synchronous broadcasting.
9. An electronic device, characterized in that, It includes a memory, a processor, and a program stored in the memory and capable of running on the processor, wherein the processor executes the program to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed, implements the method described in any one of claims 1-7.