An interactive method for interactive podcasts based on smart glasses

By combining smart glasses with cloud-based AI models, personalized podcast content can be collected and generated in real time, solving the problem of traditional podcasts lacking interactivity, fulfilling users' needs to obtain specific information instantly, and improving the flexibility and interactivity of the podcast experience.

CN119132298BActive Publication Date: 2025-10-31BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411212393.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-10-31
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Traditional podcasts lack interactivity and personalization, failing to meet users' needs for real-time access to specific information.

Method used

By combining smart glasses with a cloud-based AI model, user voice information is collected in real time to generate and play personalized podcast content. The ASR and TTS modules are used for speech-to-text and text-to-speech processing to achieve an interactive podcast experience.

Benefits of technology

It enhances the personalization and immediacy of podcast content, allowing users to trigger content generation through voice, thus increasing the flexibility and interactivity of the user experience. It breaks the traditional one-way reception mode of podcasts, adding convenience and fun to fast-paced life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132298B_ABST
    Figure CN119132298B_ABST
Patent Text Reader

Abstract

This disclosure provides an interactive podcast method based on smart glasses, belonging to the field of interactive technology between smart devices and podcast playback. The method includes: responding to receiving first interactive information, generating first text information based on the first interactive information, sending the first text information to a second device, wherein the first text information is used to instruct the second device to generate first podcast text information matching the first text information based on a target AI model; and responding to receiving the first podcast text information sent by the second device, converting the first podcast text information into first podcast audio information, and playing the first podcast audio information. This interactive podcast method based on smart glasses provides improved interactivity and personalization, meeting users' needs for real-time access to specific information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure pertains to the field of interactive technology of smart devices and podcast playback, and more specifically, relates to an interactive method for interactive podcasts based on smart glasses. Background Technology

[0002] Traditional podcasts typically consist of pre-recorded audio content, requiring users to select content of interest on their phones before listening. Therefore, existing interactive methods lack interactivity and personalization, failing to meet users' needs for real-time access to specific information. Summary of the Invention

[0003] The purpose of this disclosure is to provide an interactive method for interactive podcasts based on smart glasses, in order to improve interactivity and personalization and meet users' needs for obtaining specific information in real time.

[0004] A first aspect of this disclosure provides an interactive method for an interactive podcast based on smart glasses, applied to a first device, comprising:

[0005] In response to receiving the first interaction information, the device generates first text information based on the first interaction information and sends the first text information to the second device. The first text information is used to instruct the second device to generate first podcast text information that matches the first text information based on the target AI big model.

[0006] In response to receiving the first podcast text information sent by the second device, the first podcast text information is converted into first podcast voice information and played;

[0007] The second device is a device that has established a connection with the first device, and the first interactive information includes the voice information of the first user, who is a user wearing the first device.

[0008] A second aspect of this disclosure provides an interactive device for an interactive podcast based on smart glasses, applied to a first device, comprising:

[0009] The first podcast text information module is used to respond to receiving the first interaction information, generate the first text information based on the first interaction information, and send the first text information to the second device. The first text information is used to instruct the second device to generate first podcast text information that matches the first text information based on the target AI big model.

[0010] The first podcast voice information module is used to respond to receiving first podcast text information sent by the second device, convert the first podcast text information into first podcast voice information, and play the first podcast voice information; wherein, the second device is a device that has established a connection with the first device, and the first interaction information includes the voice information of the first user, and the first user is a user wearing the first device.

[0011] A third aspect of this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the interactive method for an interactive podcast based on smart glasses described above.

[0012] A fourth aspect of this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the interactive method for an interactive podcast based on smart glasses described above.

[0013] The beneficial effects of the interactive method for an interactive podcast based on smart glasses provided in this disclosure are as follows:

[0014] This disclosure significantly enhances the personalization and immediacy of podcast content. Users can trigger content generation via voice, allowing podcast content to dynamically adjust according to the user's immediate needs or interests, thus enhancing the flexibility and interactivity of the user experience.

[0015] On the one hand, by using smart glasses as an interactive medium, the traditional one-way receiving mode of podcasts has been broken, enabling interactive content consumption anytime and anywhere, adding convenience and fun to a fast-paced life.

[0016] On the other hand, thanks to the powerful capabilities of the target AI model, the generated podcast content is of high quality and highly relevant, which can meet the diverse information needs of users and promote the innovative application and popularization of AI technology in daily life.

[0017] Therefore, this embodiment improves interactivity and personalization, meeting users' needs to obtain specific information in real time. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating an interactive method for an interactive podcast based on smart glasses, provided as an embodiment of this disclosure;

[0020] Figure 2 This is a schematic diagram of the signaling interaction process between devices provided in an embodiment of the present disclosure;

[0021] Figure 3A structural block diagram of an interactive device for an interactive podcast based on smart glasses, provided as an embodiment of this disclosure;

[0022] Figure 4 This is a schematic block diagram of an electronic device provided according to an embodiment of the present disclosure. Detailed Implementation

[0023] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, so as to provide a thorough understanding of the embodiments of this disclosure. However, those skilled in the art will understand that this disclosure may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this disclosure with unnecessary detail.

[0024] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0025] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an interactive podcasting method based on smart glasses, provided as an embodiment of the present disclosure. The method is applied to a first device and may include steps S101 and S102.

[0026] S101: In response to receiving the first interaction information, generate first text information based on the first interaction information, and send the first text information to the second device. The first text information is used to instruct the second device to generate first podcast text information that matches the first text information based on the target AI big model.

[0027] In this embodiment, the first device can be a wearable device capable of real-time, direct information interaction with the user, and the first user is the user wearing the first device. For example, the first device can be smart glasses, which have the functions of collecting and playing voice and communicating and interacting with other devices.

[0028] The first interactive information includes the voice information of the first user, for example, the first user speaking out a podcast topic that interests them.

[0029] For example, the first user speaks out a podcast topic of interest, such as "explain Journey to the West," and the smart glasses can collect the first user's voice information through the built-in microphone, that is, the first device receives the first interaction information.

[0030] The first device includes a microphone, speaker, processor, wireless communication module (such as BLE or WIFI), automatic speech recognition (ASR) module, text-to-speech (TTS) module, background noise detection and voice activity detection (VAD) module.

[0031] After the first device responds to the received first interaction information, it converts the first interaction information into first text information through the first device's ASR module, and the first text information can be sent to the second device through a wireless communication module (such as BLE or WIFI).

[0032] The second device is one that has established a connection with the first device. It is responsible for receiving the first text information from the first device and performing subsequent processing. For example, the second device could be a cloud-based AI large model server, capable of receiving text information, understanding and expanding the text information, and generating podcast information.

[0033] This embodiment can utilize an ASR module to convert the first interactive information into first text information. The ASR module can incorporate natural language processing technology.

[0034] Specifically, the steps to convert the first interactive information into the first text information can be as follows:

[0035] First, speech signal acquisition and preprocessing.

[0036] Voice signal acquisition: The voice signal of the first user is acquired through the microphone built into the smart glasses as the first interaction information.

[0037] Preprocessing: Since the speech signal is a continuous analog signal, it contains the voice characteristics of the first user and various environmental noises. Therefore, preprocessing of the acquired speech signal is necessary.

[0038] For example, noise reduction (such as filtering) can be applied to the speech signal to remove or reduce the interference of environmental noise on the speech signal and improve the clarity of the speech.

[0039] Alternatively, pre-emphasis can be applied to boost the high-frequency components of the speech signal, compensating for the high-frequency attenuation of the speech signal during transmission, resulting in a flatter speech signal spectrum, which is beneficial for subsequent feature extraction.

[0040] Second, feature extraction is performed on the speech signal, that is, extracting parameters that can represent speech features from each frame of speech signal.

[0041] For example, the Mel-Frequency Cepstral Coefficients (MFCCs) of the speech signal can be extracted. The Mel-Frequency Cepstral Coefficients simulate the human ear's perception of sounds at different frequencies and can effectively represent the spectral characteristics of speech.

[0042] In this embodiment, the spectrum can be obtained by performing a fast Fourier transform on the speech signal, and then processed by a Mel filter and logarithmic operations to obtain MFCCs features.

[0043] Third, the processed first interaction information can be sequentially input into the pre-built acoustic model and language model to decode the first interaction information and obtain the first text information.

[0044] The acoustic model can be used to calculate the probabilistic relationship between speech features and phonemes (or other speech units). For example, the acoustic model can be a Hidden Markov Model (HMM). An HMM model treats the speech signal as generated by a series of hidden states, each state corresponding to a phoneme or subphoneme. In this embodiment, the parameters of the HMM can be determined in advance by training a large amount of speech data, thereby obtaining a trained HMM model.

[0045] Language models are used to calculate the probability of a word sequence occurring. Taking into account the grammatical, semantic, and contextual information of language, language models can further constrain and optimize the phoneme sequences output by acoustic models to obtain more accurate text results. For example, language models utilize deep neural networks to learn the relationships between words, which can better capture long-distance dependencies and semantic information in language.

[0046] For example, after the smart glasses receive the first interactive information (such as "explain Journey to the West"), the ASR module can convert the first user's voice information into first text information (such as "explain Journey to the West"), that is, the first interactive information generates first text information. Then, the first text information (such as "explain Journey to the West") is sent to the cloud AI large model server via WIFI, that is, the first text information is sent to the second device.

[0047] The first text information instructs the second device to generate a first podcast text that matches the first text information based on the target AI model. Specifically, the second device generates the first podcast text (a detailed explanation of Journey to the West) that matches the first text information based on the target AI model. For example, the target AI model could be a pre-trained Generative Adversarial Network (GAN) that understands and expands the content of the first text information to generate the matching first podcast text.

[0048] Specifically, the steps for converting the first interaction information into the first text information using a pre-trained GAN model can be as follows:

[0049] First, data preparation.

[0050] First, a large amount of speech data and corresponding text data pairs need to be collected as the basic dataset for training the GAN model. This dataset should cover various speakers, accents, speech rates, environmental noise, etc., to ensure that the model can learn a wide range of speech features and text correspondences.

[0051] Second, GAN model architecture design.

[0052] The generator converts the initial interactive information (speech signal) into initial text information. It typically consists of neural network layers, such as Convolutional Neural Networks (CNN) layers for extracting features from the speech signal, and Recurrent Neural Networks (RNN) layers for processing sequential data and generating text sequences. After receiving the speech signal, the generator performs spectral analysis and feature extraction to transform it into an intermediate representation. Then, it uses the RNN to progressively generate text characters, one character at a time, until the complete initial text information is generated.

[0053] The discriminator determines whether the generated text is authentic. It receives both the generated text and real text data as input and outputs a probability value representing the likelihood that the input text is authentic.

[0054] Discriminators are also composed of neural networks, typically including multilayer perceptrons (MLPs). They determine the authenticity of text by analyzing its grammatical, semantic, and statistical features.

[0055] Third, model training.

[0056] Alternately train the generator and discriminator:

[0057] First, with the discriminator's parameters fixed, the generator is trained. The generator's goal is to produce text that is as realistic as possible, making it difficult for the discriminator to distinguish between generated and real text. The generator's parameters are then updated by minimizing the difference between the generated and real text, for example, using a cross-entropy loss function.

[0058] Then, with the generator's parameters fixed, the discriminator is trained. The discriminator's goal is to accurately distinguish between real text and text generated by the generator. The discriminator's parameters are updated by maximizing the probability of correct discrimination—that is, maximizing the probability of real text being judged as real and the probability of generated text being judged as fake.

[0059] This process is repeated multiple times until the generator and discriminator reach an equilibrium, where the generator can generate sufficiently realistic text, while the discriminator has difficulty distinguishing between the generated text and the real text.

[0060] Introducing adversarial losses:

[0061] During training, in addition to using traditional loss functions (such as cross-entropy loss), adversarial loss is also introduced. Adversarial loss measures the adversarial relationship between the generator and the discriminator, prompting the generator to generate more realistic text to deceive the discriminator, while also prompting the discriminator to more accurately judge the authenticity of the text.

[0062] Fourth, text generation.

[0063] After training, when the first device receives the first interactive information (speech signal), it inputs it into the generator. The generator then generates the first text information based on the learned correspondence between speech and text. The generated text can undergo post-processing steps, such as removing redundant punctuation marks and correcting spelling errors, to improve the readability and accuracy of the text.

[0064] For example, the first interactive information could be a question or the starting point for a discussion. The generated first text information could be a further elaboration on the question or a guiding description of the topic. After the second device receives the first text information, the target AI model can generate a detailed podcast script, including the host's opening remarks, the guests' viewpoints, case analyses, etc., to meet the podcast requirements indicated by the first text information.

[0065] S102: In response to receiving the first podcast text information sent by the second device, convert the first podcast text information into first podcast voice information and play the first podcast voice information;

[0066] The second device is a device that has established a connection with the first device, and the first interactive information includes the voice information of the first user, who is a user wearing the first device.

[0067] In this embodiment, the first device receives the first podcast text information (detailed explanation of Journey to the West) sent by the second device, and then the first device converts the first podcast text information into first podcast voice information, which can be done using a TTS module.

[0068] For example, by calling the TTS module, text content is converted into a playable audio signal. During this process, voice parameters can be adjusted according to the needs and preferences of the first user, such as selecting different voice timbre, speech rate, and pitch, to provide a more personalized audio playback experience.

[0069] Next, the first device plays the first podcast audio message. This is played through the first device's speaker or a connected wireless headset, allowing the first user wearing the first device to hear podcast content on a topic of interest.

[0070] For example, if the first interactive information includes the voice information of the first user, it means that the first user expresses their needs through voice, triggering the entire podcast generation and playback process. As the main user wearing the first device, the first user can enjoy convenient and personalized podcast services, obtaining the information and entertainment content they need through voice interaction without manual operation.

[0071] For example, when the first user says "Explain Journey to the West," the first device converts it into text and sends it to the second device. The second device then generates a podcast text explaining Journey to the West and sends it back to the first device. The first device then converts the text into audio and plays it to the first user, allowing the user to listen to the explanation of Journey to the West anytime, anywhere.

[0072] As can be seen from the above, this embodiment greatly enhances the personalization and immediacy of podcast content. Users can trigger content generation through voice, allowing podcast content to be dynamically adjusted according to the user's immediate needs or interests, thus enhancing the flexibility and interactivity of the user experience.

[0073] On the one hand, by using smart glasses as an interactive medium, the traditional one-way receiving mode of podcasts has been broken, enabling interactive content consumption anytime and anywhere, adding convenience and fun to a fast-paced life.

[0074] On the other hand, thanks to the powerful capabilities of the target AI model, the generated podcast content is of high quality and highly relevant, which can meet the diverse information needs of users and promote the innovative application and popularization of AI technology in daily life.

[0075] Therefore, this embodiment improves interactivity and personalization, meeting users' needs to obtain specific information in real time.

[0076] In one embodiment of this disclosure, after playing the first podcast audio information, the method further includes:

[0077] In response to receiving the second interaction information, the second text information is generated based on the second interaction information and sent to the second device. The second text information is used to instruct the second device to expand the first podcast text information based on the target AI big model and the second text information to obtain the second podcast text information.

[0078] In response to receiving a second podcast text message from a second device, the second podcast text message is converted into a second podcast voice message and played.

[0079] The second interactive information includes the voice information of the first user.

[0080] In this embodiment, the second device can extract multiple first keywords from the second text information and multiple second keywords from the first podcast text information. Then, the relevance of the multiple first keywords and multiple second keywords is calculated respectively, and multiple second keywords that are related to each of the first keywords (with a relevance greater than a preset relevance) are selected from the multiple second keywords as multiple target keywords.

[0081] First, define the first keyword set as The second set of keywords is .

[0082] Define the relevance function , indicating the first keyword With the second keyword The relevance. The preset relevance is... .

[0083] First, calculate the importance weight of the keywords.

[0084] For keywords in the first keyword set Its importance weight It can be calculated using the following formula:

[0085]

[0086] in, Keywords Word frequency in the second text information Keywords Inverse document frequencies throughout the corpus, and It is an adjustable weighting coefficient, and .

[0087] For keywords in the second keyword set Its importance weight It can be calculated using the following formula:

[0088]

[0089] in, Keywords Word frequency in the second text information Keywords Inverse document frequencies throughout the corpus, and It is an adjustable weighting coefficient, and .

[0090] Second, calculate the correlation function.

[0091] The relevance is calculated by combining the vector space model with semantic similarity.

[0092] Let the word vector of the keyword be represented as The semantic similarity function is A pre-trained language model can be used to calculate the semantic similarity between two keywords.

[0093] The correlation function is defined as:

[0094]

[0095] in, and It is an adjustable weighting coefficient, and .

[0096] Third, determine the target keyword set.

[0097] The formula for calculating the target keyword set is as follows:

[0098]

[0099] This embodiment comprehensively considers multiple factors such as keyword frequency, inverse document frequency, word vector similarity, and semantic similarity, enabling more accurate identification of target keywords. In practical applications, the weighting coefficients and preset relevance can be adjusted according to specific circumstances. To achieve the best results.

[0100] Multiple primary keywords and multiple target keywords are input into the target AI model to expand the first podcast text information and obtain the second podcast voice information that matches the second interactive information.

[0101] After receiving the second podcast audio information, the second device can send the second podcast audio information to the first device.

[0102] The second interactive information includes the voice information of the first user. At this time, the voice information of the first user is a further question to the first interactive information, such as "Please explain Sun Wukong to me in more detail", which is based on the first interactive information (explain Journey to the West).

[0103] While playing the first podcast audio message, the first device (smart glasses) will dynamically respond and update based on further interactions from the first user.

[0104] While the first device is playing the first podcast audio information, if the first user provides second interactive information, this second interactive information is also received by the first device in the form of audio. The first device processes the second interactive information in the same way as it processed the first interactive information. First, it generates second text information based on the second interactive information. This process can still use ASR technology, including steps such as audio signal acquisition and preprocessing, feature extraction, and pre-construction of acoustic and language models, to convert the first user's audio information into accurate second text content.

[0105] Then, the first device sends the second text information to the second device. Upon receiving the second text information, the second device expands the first podcast text information based on the target AI model (such as a pre-trained GAN model) and the second text information. The target AI model further mines relevant content based on the new second text information and the previously generated first podcast text information, conducting more in-depth analysis and expansion to obtain the second podcast text information. During this process, the target AI model fully considers the context of the first podcast text information and the new needs and directions provided by the second text information, generating even newer second podcast text information.

[0106] Next, when the first device receives the second podcast text message from the second device, it converts it into second podcast audio message. This step is similar to converting the first podcast text message into first podcast audio message, using the TTS module to convert the new second podcast text message into playable second podcast audio message.

[0107] Finally, the first device plays the second podcast audio message, allowing the first user to hear the expanded podcast content.

[0108] For example, while a first user is listening to the first podcast audio message about "Journey to the West," they say, "Please explain Sun Wukong in more detail," as a second interactive message. The first device converts this into second text information and sends it to a second device. The second device uses a target AI model to expand on the first podcast text message about "Journey to the West," expanding on the personality traits of Sun Wukong, generating a second podcast text message, and sending it back to the first device. The first device then converts this into second podcast audio information and plays it back to the first user.

[0109] As can be seen from the above, this interactive podcast interaction method further enhances the coherence and depth of the first-user experience. The first user can provide immediate feedback and guide the expansion of podcast content, achieving a natural transition from a single podcast to in-depth, conversational content. Dynamic content expansion not only satisfies the first user's curiosity to explore the unknown but also promotes a more personalized and immersive information acquisition process, making the podcast experience richer and more engaging.

[0110] In one embodiment of this disclosure, before playing the second podcast audio information, the method further includes:

[0111] Stop playing the first podcast audio message.

[0112] In this embodiment, stopping the playback of the first podcast audio message before playing the second podcast audio message avoids the confusion and interference caused by simultaneous playback of different audio content. If the first and second podcast audio messages are played simultaneously, the first user will have difficulty hearing either content clearly, greatly affecting the efficiency and experience of the first user in obtaining information. By stopping the playback of the first podcast audio message, the first user can focus on the new extended content, namely the second podcast audio message, ensuring that the first user can clearly receive and understand the current podcast content.

[0113] When the second interactive message triggers the expansion of the first podcast text message and generates a second podcast text message, the first device needs to clearly indicate that the current playback focus should shift to the new content. Stopping playback of the first podcast audio message is a clear signal switching mechanism, indicating that the first device is responding to a new user request and preparing to play the second podcast audio message.

[0114] For example, when a first user is listening to the first podcast audio message about "Journey to the West," they might send a second interactive message requesting further information about Sun Wukong. At this point, stopping the first podcast audio message allows the first user's attention to shift away from the previous explanation and better prepare to receive the second podcast audio message about Sun Wukong, thus achieving a smoother and more coherent interactive experience.

[0115] As can be seen from the above, this embodiment ensures the smoothness of content switching, avoids information overlap, allows the first user to focus on the current podcast content, and improves the coherence and clarity of the overall listening experience.

[0116] In one embodiment of this disclosure, after playing the second podcast audio information, the method further includes:

[0117] In response to the completion of the second podcast audio message, a first voice prompt is output, which asks the user whether to continue playing the first podcast audio message;

[0118] In response to receiving the third interaction information, a third text information is generated based on the third interaction information and sent to the second device. The third text information is used to instruct the second device to expand the second podcast text information based on the target AI big model and the third text information to obtain the third podcast text information.

[0119] In response to receiving a third podcast text message from a second device, the third podcast text message is converted into a third podcast voice message and played.

[0120] In response to receiving the fourth interactive message, continue playing the first podcast audio message.

[0121] In this embodiment, after the second podcast audio message has finished playing, the first device outputs a first voice prompt. The purpose of this voice prompt is to ask the first user whether to continue playing the first podcast audio message. For example, the voice prompt could be, "Do you want to continue playing the previous explanation of Journey to the West?"

[0122] Next, if the first user responds to the voice prompt, namely, "Tell me about Sun Wukong's personality traits," a third interactive message is generated. The first device generates a third text message based on this third interactive message and sends it to the second device. The second device expands the second podcast text message based on the target AI model and the third text message, generating a third podcast text message. This process of expanding the first podcast text message is similar; the target AI model will conduct in-depth analysis and expansion based on the new third text message and the existing second podcast text message to meet the first user's further needs.

[0123] Then, after receiving the text message from the third podcaster, the first device converts it into the third podcaster's audio message and plays it, providing users with new extended content and enabling them to continuously access richer information.

[0124] On the other hand, if the user, after hearing the voice prompt, sends a fourth interactive message requesting continued playback of the first podcast audio message, the first device will respond to the fourth interactive message and continue playing the first podcast audio message. At this time, the user can return to the previous content at any time for review or further reflection, improving the flexibility and adaptability of the first device.

[0125] For example, after playing the second podcast audio message explaining Sun Wukong, the first device issues a voice prompt asking the first user whether to continue playing the explanation of "Journey to the West". If the first user answers "yes", the first device continues playing the first podcast audio message; if the first user raises a new question or request as a third interactive message, the first device will expand and play the third podcast audio message accordingly.

[0126] As can be seen from the above, interacting with the first user through voice prompts after playing the second podcast audio message enhances the flexibility of the first user's participation. Allowing users to choose whether to continue the first podcast or explore the third podcast content expanded by the target AI model not only satisfies personalized needs but also enriches the listening experience and increases the diversity of podcasts.

[0127] In one embodiment of this disclosure, when playing the first podcast audio information, the method further includes:

[0128] Upon receiving the target voice information, determine whether the target voice information matches the first user;

[0129] If the target voice information does not match the first user, then control the first device to enter conversation mode;

[0130] Specifically, when the first device is in talk mode, the first device stops playing podcast audio messages.

[0131] In this embodiment, when the first device receives target voice information while playing the first podcast voice information, the first device will determine whether the target voice information matches the first user. This matching process can be carried out in various ways, such as analyzing the timbre, pitch, and speech rate of the voice, or identifying the speaker's identity information through specific ASR technology. If the determination result shows that the target voice information does not match the first user, it means that someone else is talking to the first user at this time, and the first device automatically enters the conversation mode.

[0132] When the first device is in talk mode, it will stop playing podcast audio messages. This prevents podcast audio from interfering with the current conversation and ensures that the conversation can proceed clearly. Enabling talk mode allows the first device to better adapt to different usage scenarios, improving its flexibility and usability.

[0133] For example, when a user is playing a podcast about Journey to the West through smart glasses, and someone nearby suddenly speaks, the first device receives this target audio message, determines that it doesn't match the user's voice, and immediately enters conversation mode, stopping the podcast playback so the user can focus on communicating with the person next to them. After the conversation ends, the user can restart the podcast playback as needed.

[0134] As can be seen from the above, when playing the first podcast audio information, the device intelligently switches to conversation mode by recognizing the matching degree between the target audio information and the user, effectively avoiding interference from audio information and ensuring smooth communication between users. This embodiment not only improves the user experience but also demonstrates the first device's keen perception and flexible response capability to the first user's context, creating a more user-friendly and convenient usage environment for the first user. At the same time, it also enhances the applicability of the first device in various usage scenarios.

[0135] In one embodiment of this disclosure, an interactive podcast method based on smart glasses further includes:

[0136] In response to the first device being in conversation mode, if no voice information is detected within the target duration, the first device is controlled to exit conversation mode and continue playing podcast voice information.

[0137] In this embodiment, when the first device is in conversation mode, it indicates that someone other than the first user is interacting with the environment surrounding the first device, and the playback of podcast audio is paused to avoid interfering with the conversation. If no audio is detected within the target duration, it means that the surrounding conversation has ended or a relatively quiet phase has begun. To ensure that the first user can continue to enjoy the podcast content, the first device can exit conversation mode and resume playing the podcast audio.

[0138] The target duration can be adjusted based on the actual usage scenario and the needs of the first user. For example, it can be set to tens of seconds or a few minutes. If the target duration is set shorter, it can respond more quickly to the end of the conversation, allowing the first user to return to podcast listening mode as soon as possible; if the target duration is set longer, it can avoid frequent switching between conversation mode and podcast playback mode due to brief silences, improving the stability of the first device.

[0139] For example, a user is listening to a podcast using smart glasses. Suddenly, someone speaks to the user. The smart glasses enter conversation mode and stop playing the podcast audio. If no voice is detected within the next minute, the smart glasses automatically exit conversation mode and resume playing the previously paused podcast audio, allowing the user to continue enjoying the podcast content without manual intervention.

[0140] As can be seen from the above, the introduction of the automatic exit from talk mode significantly improves the continuity of the first-user experience. When the talk mode automatically ends without any voice activity within the target duration, podcast playback resumes, avoiding unnecessary waiting and ensuring a seamless transition for the first user back to the original podcast content. This not only reduces manual operation for the first user but also demonstrates the device's deep understanding and precise response to the user's behavioral patterns, further enhancing the smoothness and convenience of the podcast experience.

[0141] In one embodiment of this disclosure, determining whether the target voice information matches the first user includes:

[0142] Extract multiple speech features from the target speech information;

[0143] Multiple speech features are fused to obtain the target fused feature;

[0144] Calculate the matching degree between the target fusion feature and the standard fusion feature;

[0145] The matching degree is used to determine whether the target voice information matches the first user.

[0146] In this embodiment, the target speech information includes multiple speech segments, and multiple speech features of the target speech information are extracted, including:

[0147] Extract multiple segment speech features from the target segment speech information, where the target segment speech information is one segment of multiple speech information.

[0148] Specifically, the target speech information can be divided into multiple segments according to a preset duration. For each segment, the information content of that segment is calculated. If the information content of a segment is less than the preset information content, the segment is marked as invalid. If the information content of a segment is greater than or equal to the preset information content, the segment is marked as valid.

[0149] For example, the steps for calculating the information content of each speech segment based on speech feature statistics are as follows:

[0150] First, spectral characteristic analysis.

[0151] Perform a Fast Fourier Transform (FFT) on each segment of speech information to obtain its spectrum.

[0152] The entropy of the spectrum is used as a measure of information content. Spectral entropy reflects the uniformity of signal distribution at different frequencies. The more uniform the distribution, the larger the entropy value, and the relatively smaller the information content; the more concentrated the distribution, the smaller the entropy value, and the relatively larger the information content.

[0153] For example, for a relatively single-frequency speech signal, its spectrum will be concentrated around a specific frequency, with a low entropy value, indicating a large amount of information and conveying specific speech content, such as a clear word or phrase.

[0154] Second, temporal characteristic analysis.

[0155] Calculate the zero-crossing rate of a speech signal. The zero-crossing rate refers to the number of times a signal crosses zero per unit time. Speech signals with a higher zero-crossing rate usually contain more high-frequency components and may have a relatively larger amount of information.

[0156] Analyzing the energy distribution of a speech signal allows us to calculate its energy over different time periods and observe its variations. Regions with larger energy changes contain more information.

[0157] In this embodiment, all the segment speech features extracted from the target segment speech information are valid speech information. The target segment speech information can be a segment of speech information from multiple valid speech information segments, and the preset information content can be set according to the actual situation. For example, it can be determined based on the information content of speech information in the first user's historical conversation scenario, and multiple valid speech information segments can be calculated and then averaged.

[0158] As can be seen from the above, by meticulously dividing speech information and evaluating its information content, invalid speech segments are effectively eliminated, while information-rich valid speech segments are retained, significantly improving the accuracy and efficiency of speech processing.

[0159] Accordingly, multiple speech features are fused to obtain the target fused features, including:

[0160] The speech features of multiple segments corresponding to the speech information of the target segment are fused to obtain the target segment fused features.

[0161] The target fusion features are determined based on all the target segment fusion features.

[0162] In this embodiment, multiple segment speech features corresponding to the target segment speech information are fused to obtain the target segment fused features:

[0163] Determine the feature weights of each speech segment.

[0164] The speech features of each segment are weighted and calculated based on their respective feature weights to obtain the target segment fusion features.

[0165] Multiple speech features can include timbre, intonation, and speech rate. First, we need to unify the dimensions of each speech feature. For example, we can perform standardization processing, as follows:

[0166] (1) For numerical features (such as speech rate), calculate the mean and standard deviation of the feature across all speech data. For each speech rate value, standardize it using the formula "(original value - mean) / standard deviation" so that the speech rate feature value is distributed within a specific interval, usually an interval with a mean of 0 and a standard deviation of 1. Therefore, speech rate values ​​of different ranges can be unified to the same scale.

[0167] (2) For the quantification of non-numerical features (such as timbre), timbre is a relatively complex feature, which can be converted into a numerical representation through some feature extraction methods. For example, the Mel frequency cepstral coefficients (MFCCs) method can be used to extract the numerical vector of timbre features.

[0168] Then, the extracted timbre feature vectors are standardized, similar to the method for numerical features, by calculating the mean and standard deviation and standardizing them so that different timbre feature vectors are within the same scale range.

[0169] (3) For encoding categorical features (such as intonation), if intonation can be divided into several different types, one-hot encoding can be used to convert it into a numerical vector. For example, if there are three intonation types, they can be represented by vectors [1, 0, 0], [0, 1, 0], and [0, 0, 1], respectively. Therefore, different intonation types are converted into numerical vectors of the same dimension, which can be further processed with other features.

[0170] Optionally, the target segment fusion features can be calculated using the first formula.

[0171] The first formula is:

[0172]

[0173] in, For target segment fusion features, For the first Weights of speech features in each segment For the first Segment speech features The number of speech features in a segment.

[0174] This embodiment can unify the vector dimension of the fused features of each target segment. For example, the features of timbre, intonation, and speech rate can be represented as vectors respectively. , , They can be combined into a new vector. ,in , , The dimensions are respectively , 1. It is necessary to ensure that different features have reasonable positions and orders in the combined vector for subsequent processing and analysis.

[0175] In this embodiment, after obtaining all the target fusion features, the target fusion features can be calculated using the second formula.

[0176] The second formula is:

[0177]

[0178] in, To achieve the goal of feature fusion, For the first fusion features of target segments For the first The weights of the fused features of each target segment The number of features in the target segment. It can be determined based on the information content of the target segment speech corresponding to the speech features of that target segment.

[0179]

[0180] Among them, a <b<c, For the first Information content of speech features in each target segment As the primary information content, This is the second piece of information.

[0181] For example, a, b, and c are 0.2, 0.6, and 1.2 respectively, with the greater the amount of information, the greater the weight.

[0182] In this embodiment, the standard fusion feature can be calculated based on the first user's speech information over a certain period of time, using the same calculation method as described above. For example, by using cosine similarity calculation, the similarity between the target segment fusion feature and the standard fusion feature is calculated as the matching degree between the two.

[0183] The specific calculation method is as follows:

[0184] First, calculate the dot product. Calculate the dot product between the target segment fused feature vector and the standard fused feature vector. The dot product reflects the sum of the products of two vectors across all dimensions, and can measure the common features of two vectors in the same dimensions. If two vectors have similar values ​​in some dimensions, their dot product will be larger, indicating a high correlation in those dimensions.

[0185] Second, calculate the modulus. Calculate the modulus of the target segment fused feature vector and the standard fused feature vector. The modulus represents the length of the vector, reflecting its overall size across all dimensions. The modulus can be calculated by summing the squares of the values ​​in each dimension of the vector and then taking the square root.

[0186] Third, calculate the similarity. Divide the dot product by the product of the magnitudes of the two vectors to obtain the cosine similarity value. This value ranges from -1 to 1, where 1 indicates that the two vectors are exactly the same, 0 indicates that the two vectors are completely unrelated, and -1 indicates that the two vectors are completely opposite.

[0187] In speech feature matching, a higher cosine similarity value indicates a higher degree of matching between the target segment fusion feature and the standard fusion feature, meaning that the two speech segments have a high degree of similarity in terms of timbre, intonation, and speech rate.

[0188] In this embodiment, determining whether the target voice information matches the first user based on the matching degree may include:

[0189] If the matching degree is greater than or equal to the first matching degree, the target voice information is determined to match the first user.

[0190] If the matching degree is less than the first matching degree, the target voice information is determined not to match the first user.

[0191] As can be seen from the above, this embodiment significantly improves the accuracy and personalization of speech recognition. Through multi-dimensional feature fusion, such as timbre, intonation, and speech rate, the comprehensiveness and representativeness of speech features are ensured. Feature weight calculation enhances the influence of important features and improves the effectiveness of feature fusion. Simultaneously, the cosine similarity-based matching method can quickly and accurately assess the similarity between speech segments, effectively distinguishing the speech features of different users. This embodiment not only enhances the reliability of user identification but also provides strong support for personalized voice services.

[0192] refer to Figure 2 , Figure 2 This is a schematic diagram of the signaling interaction process between devices according to an embodiment of the present disclosure. The signaling interaction process between the first device and the second device is as follows:

[0193] A101: Receive the first interactive information.

[0194] The first device receives the first interactive information.

[0195] A102: Generate the first text information.

[0196] The first device generates the first text information based on the first interaction information.

[0197] A103: Send the first text message.

[0198] The first device sends the first text information to the second device.

[0199] A104: Generate the first podcast text information.

[0200] The first text information is used to instruct the second device to generate a first podcast text information that matches the first text information based on the target AI large model.

[0201] A105: Send the first podcast text message.

[0202] The second device sends the first podcast text message to the first device.

[0203] A106: Generate the first podcast audio message.

[0204] The first device converts the text information of the first podcast into the audio information of the first podcast.

[0205] A107: Play the first podcast audio message.

[0206] The first device plays the first podcast audio message.

[0207] A108: Receive the second interactive information.

[0208] The first device receives the second interactive information.

[0209] A109: Generate the second text information.

[0210] The first device generates second text information based on the second interactive information.

[0211] A110: Send the second text message.

[0212] The first device sends the second text information to the second device.

[0213] A111: Generate the second podcast text information.

[0214] The second text information is used to instruct the second device to expand the first podcast text information based on the target AI big model and the second text information to obtain the second podcast text information.

[0215] A112: Send the second podcast text message.

[0216] The first device receives a second podcast text message sent by the second device.

[0217] A113: Generate second podcast audio information.

[0218] The first device converts the text information of the second podcast into the audio information of the second podcast.

[0219] A114: Stop playing the first podcast audio message.

[0220] Before playing the second podcast audio message, the first device stops playing the first podcast audio message.

[0221] A115: Play the second podcast audio message.

[0222] The first device plays the second podcast audio message.

[0223] A116: The second podcast audio message has finished playing.

[0224] The first device finished playing the second podcast audio message.

[0225] A117: Output the first voice prompt.

[0226] The first device outputs a first voice prompt, which asks the user whether to continue playing the first podcast audio message.

[0227] A118: Receive third interactive information.

[0228] The first device receives the third interactive information.

[0229] A119: Generate third text information.

[0230] The first device generates third text information based on the third interactive information.

[0231] A120: Send a third text message.

[0232] The first device sends the third text information to the second device.

[0233] A121: Generate third podcast text information.

[0234] The third text information is used to instruct the second device to expand the second podcast text information based on the target AI large model and the third text information, thereby obtaining the third podcast text information.

[0235] A122: Send a text message to the third podcaster.

[0236] The first device receives a third podcast text message sent by the second device.

[0237] A123: Generate third podcast audio information.

[0238] The first device converts the text information of the third podcast into the audio information of the third podcast.

[0239] A124: Play the third podcast audio message.

[0240] The first device plays audio messages from a third podcast.

[0241] A125: Receive the fourth interactive information.

[0242] The first device receives the fourth interactive information.

[0243] A126: Continue playing the first podcast audio message.

[0244] The first device continues playing the first podcast audio message.

[0245] Corresponding to the interactive podcasting method based on smart glasses in the above embodiment, Figure 3 This is a structural block diagram of an interactive device for an interactive podcast based on smart glasses, according to one embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiment of this disclosure are shown. References Figure 3 The interactive podcast device 20 based on smart glasses is applied to the first device and includes: a first podcast text information module 201 and a first podcast voice information module 202.

[0246] Among them, the first podcast text information module 201 is used to respond to receiving the first interaction information, generate first text information based on the first interaction information, and send the first text information to the second device. The first text information is used to instruct the second device to generate first podcast text information that matches the first text information based on the target AI big model.

[0247] The first podcast voice information module 202 is used to respond to receiving first podcast text information sent by the second device, convert the first podcast text information into first podcast voice information, and play the first podcast voice information; wherein, the second device is a device that has established a connection with the first device, and the first interaction information includes the voice information of the first user, and the first user is a user wearing the first device.

[0248] In one embodiment of this disclosure, an interactive podcast device 20 based on smart glasses further includes: a second podcast text information module and a second podcast voice information module;

[0249] The second podcast text information module is used to respond to receiving the second interaction information, generate second text information based on the second interaction information, and send the second text information to the second device. The second text information is used to instruct the second device to expand the first podcast text information based on the target AI big model and the second text information to obtain the second podcast text information.

[0250] The second podcast voice information module is used to respond to receiving second podcast text information sent by the second device, convert the second podcast text information into second podcast voice information, and play the second podcast voice information; the second interactive information includes the voice information of the first user.

[0251] In one embodiment of this disclosure, an interactive podcast device 20 based on smart glasses further includes: a stop module;

[0252] The stop module is used to stop playing the first podcast audio message.

[0253] In one embodiment of this disclosure, an interactive podcast device 20 based on smart glasses further includes: a first voice prompt module, a third podcast text information module, a third podcast voice information module, and a continue playback module;

[0254] The first voice prompt module is used to output a first voice prompt in response to the completion of the second podcast audio information playback. The first voice prompt is used to ask the user whether to continue playing the first podcast audio information.

[0255] The third podcast text information module is used to respond to the receipt of the third interaction information, generate third text information based on the third interaction information, and send the third text information to the second device. The third text information is used to instruct the second device to expand the second podcast text information based on the target AI big model and the third text information to obtain the third podcast text information.

[0256] The third podcast voice information module is used to respond to receiving third podcast text information sent by the second device, convert the third podcast text information into third podcast voice information, and play the third podcast voice information;

[0257] The Continue Play module is used to continue playing the first podcast audio message in response to receiving the fourth interactive message.

[0258] In one embodiment of this disclosure, the first podcast voice information module 202 is specifically used for:

[0259] Upon receiving the target voice information, determine whether the target voice information matches the first user;

[0260] If the target voice information does not match the first user, then control the first device to enter conversation mode;

[0261] Specifically, when the first device is in talk mode, the first device stops playing podcast audio messages.

[0262] In one embodiment of this disclosure, the first podcast voice information module 202 is further configured to:

[0263] In response to the first device being in conversation mode, if no voice information is detected within the target duration, the first device is controlled to exit conversation mode and continue playing podcast voice information.

[0264] In one embodiment of this disclosure, the first podcast voice information module 202 is further configured to:

[0265] Extract multiple speech features from the target speech information;

[0266] Multiple speech features are fused to obtain the target fused feature;

[0267] Calculate the matching degree between the target fusion feature and the standard fusion feature;

[0268] The matching degree is used to determine whether the target voice information matches the first user.

[0269] See Figure 4 , Figure 4 This is a schematic block diagram of an electronic device provided according to an embodiment of the present disclosure. Figure 4 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above-described device embodiments, for example... Figure 3 The functions of modules 21 and 22 shown.

[0270] It should be understood that, in the embodiments of this disclosure, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0271] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0272] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.

[0273] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this disclosure can execute the implementation methods described in the first and second embodiments of the interactive podcast based on smart glasses provided in the embodiments of this disclosure, or they can execute the implementation methods of the electronic devices described in the embodiments of this disclosure, which will not be repeated here.

[0274] In another embodiment of this disclosure, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to implement these processes. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0275] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0276] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0277] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0278] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.

[0279] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this disclosure, depending on actual needs.

[0280] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0281] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this disclosure, and these modifications or substitutions should all be covered within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. An interactive method for an interactive podcast based on smart glasses, characterized in that, Applied to the first device, including: In response to receiving the first interaction information, the device generates first text information based on the first interaction information and sends the first text information to the second device. The first text information is used to instruct the second device to generate first podcast text information that matches the first text information based on the target AI big model. In response to receiving the first podcast text information sent by the second device, the first podcast text information is converted into first podcast voice information and played. Wherein, the second device is a device that has established a connection with the first device, and the first interaction information includes the voice information of the first user, and the first user is a user wearing the first device; When playing the first podcast audio information, the method further includes: in response to receiving the target audio information, extracting multiple speech features of the target audio information; fusing the multiple speech features to obtain a target fused feature; calculating the matching degree between the target fused feature and the standard fused feature; if the matching degree is greater than or equal to a first matching degree, determining that the target audio information matches the first user; if the matching degree is less than the first matching degree, determining that the target audio information does not match the first user; if the target audio information does not match the first user, controlling the first device to enter a talk mode; wherein, when the first device is in talk mode, the first device stops playing the podcast audio information; The extraction of multiple speech features from the target speech information includes: dividing the target speech information into multiple speech segments according to a preset duration; calculating the information content of each speech segment; if the information content of the speech segment is less than a preset information content, then marking the speech segment as invalid speech information; if the information content of the speech segment is greater than or equal to the preset information content, then marking the speech segment as valid speech information; extracting multiple segment speech features from the target speech segment, wherein the target speech segment is one of the multiple speech segments; and all extracted segment speech features from the target speech segment are valid speech information. The target fusion feature is obtained by fusing multiple speech features, including: fusing multiple speech features corresponding to the target segment speech information to obtain the target segment fusion feature; and determining the target fusion feature based on all the target segment fusion features. The step of fusing multiple speech features corresponding to the target segment speech information to obtain the target segment fused features includes: determining the feature weights of each speech feature; and performing weighted calculations on each speech feature based on its feature weights to obtain the target segment fused features.

2. The interactive method for an interactive podcast based on smart glasses as described in claim 1, characterized in that, After playing the first podcast audio message, the program also includes: In response to receiving the second interaction information, a second text information is generated based on the second interaction information, and the second text information is sent to the second device. The second text information is used to instruct the second device to expand the first podcast text information based on the target AI big model and the second text information to obtain the second podcast text information. In response to receiving the second podcast text information sent by the second device, the second podcast text information is converted into second podcast voice information and played. The second interactive information includes the voice information of the first user.

3. The interactive method for an interactive podcast based on smart glasses as described in claim 2, characterized in that, Before playing the second podcast audio message, it also includes: Stop playing the first podcast audio message.

4. The interactive method for an interactive podcast based on smart glasses as described in claim 3, characterized in that, After playing the second podcast audio message, it also includes: In response to the completion of playback of the second podcast audio information, a first voice prompt is output, which asks the user whether to continue playing the first podcast audio information; In response to receiving third interaction information, third text information is generated based on the third interaction information and sent to the second device. The third text information is used to instruct the second device to expand the second podcast text information based on the target AI big model and the third text information to obtain the third podcast text information. In response to receiving the third podcast text information sent by the second device, the third podcast text information is converted into third podcast voice information and played. Upon receiving the fourth interactive message, continue playing the first podcast audio message.

5. The interactive method for an interactive podcast based on smart glasses as described in claim 1, characterized in that, Also includes: In response to the first device being in conversation mode, if no voice information is detected within the target duration, the first device is controlled to exit conversation mode and continue playing podcast voice information.

6. An interactive device for an interactive podcast based on smart glasses, characterized in that, Applied to the first device, including: The first podcast text information module is used to respond to receiving the first interaction information, generate the first text information based on the first interaction information, and send the first text information to the second device. The first text information is used to instruct the second device to generate the first podcast text information that matches the first text information based on the target AI big model. The first podcast voice information module is used to respond to receiving the first podcast text information sent by the second device, convert the first podcast text information into first podcast voice information, and play the first podcast voice information; wherein, the second device is a device that has established a connection with the first device, and the first interaction information includes the voice information of the first user, and the first user is a user wearing the first device; The first podcast voice information module is specifically configured to, in response to receiving target voice information, extract multiple voice features of the target voice information; fuse the multiple voice features to obtain target fused features; calculate the matching degree between the target fused features and the standard fused features; if the matching degree is greater than or equal to a first matching degree, determine that the target voice information matches the first user; if the matching degree is less than the first matching degree, determine that the target voice information does not match the first user; if the target voice information does not match the first user, control the first device to enter a talk mode; wherein, when the first device is in talk mode, the first device stops playing podcast voice information; The extraction of multiple speech features from the target speech information includes: dividing the target speech information into multiple speech segments according to a preset duration; calculating the information content of each speech segment; if the information content of the speech segment is less than a preset information content, then marking the speech segment as invalid speech information; if the information content of the speech segment is greater than or equal to the preset information content, then marking the speech segment as valid speech information; extracting multiple segment speech features from the target speech segment, wherein the target speech segment is one of the multiple speech segments; and all extracted segment speech features from the target speech segment are valid speech information. The step of fusing multiple speech features to obtain target fusion features includes: fusing multiple speech features corresponding to the target segment speech information to obtain target segment fusion features; and determining the target fusion features based on all target segment fusion features. The step of fusing multiple speech features corresponding to the target segment speech information to obtain the target segment fused features includes: determining the feature weights of each speech feature; and performing weighted calculations on each speech feature based on its feature weights to obtain the target segment fused features.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • In-car mute system and playing equipment with same

    CN101582700A

  • Speech interaction method and device, equipment and storage medium

    CN109256133A

  • Voice interaction system and method, electronic equipment and storage medium

    CN116417003A