Processing circuit of electronic device and related processing method

By designing processing circuits in the audio and video playback system, user hot spot detection and audio/video content recognition are solved, and the problem of how to improve the user experience is provided is provided. A detailed analysis of user responses and content is helped to optimize streaming media services.

CN120020949APending Publication Date: 2025-05-20REALTEK SEMICON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410290787.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-03-14
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

How to improve the user's experience in the audio and video playback system, especially in streaming media services, by detecting the user's hot spots and identifying audio/video content in order to provide reference for subsequent playback.

Method used

A processing circuit for an electronic device is designed, including an audio/video content generation module, a user hotspot detection module and an output module. When playing audio and video, the system receives sound through the microphone, performs acoustic echo cancellation, voice activity detection and emotion analysis, generates user hot spot detection results, and recognizes audio/video content through the audio/video content recognition module.

Benefits of technology

It realizes the response and content recognition of users during audio and video playback, and can provide user hot spot detection results and audio/video content recognition results for streaming media platforms, helping to optimize playback experience and recommendation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020949A_ABST
    Figure CN120020949A_ABST
Patent Text Reader

Abstract

The invention discloses a processing circuit of an electronic device. The processing circuit comprises an audio / video content generation module, a user hotspot detection module and an output module. The audio / video content generation module is used for generating audio data and video data and transmitting the audio data and the video data to the loudspeaker and the display panel respectively. The user hotspot detection module is used for receiving microphone input by a microphone of the electronic device when the loudspeaker plays the audio data and the display panel displays the video data, and detecting the microphone input to generate a user hotspot detection result. And the output module is used for storing the user hotspot detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio - video playback system. Background Art

[0002] In recent years, media entertainment has become a part of most people's lives, and people spend more time on streaming media services such as Youtube, Tiktok, Netflix, etc. Therefore, how to improve the user experience has become an important issue. Summary of the Invention

[0003] Therefore, an object of the present invention is to provide a control method for an electronic device having an audio - video playback system, which can obtain user hot - spot detection results and / or audio / video recognition results on the time axis of a program for reference by a streaming media platform when playing subsequent programs, so as to solve the above problems.

[0004] In an embodiment of the present invention, a processing circuit of an electronic device is disclosed, which includes an audio / video content generation module, a user hot - spot detection module, and an output module. The audio / video content generation module is used to generate audio data and video data and transmit them to a speaker and a display panel respectively. The user hot - spot detection module is used to receive a microphone input by a microphone of the electronic device and detect the microphone input to generate a user hot - spot detection result when the speaker plays the audio data and the display panel displays the video data. The output module is used to store the user hot - spot detection result.

[0005] In an embodiment of the present invention, a processing method of an electronic device is disclosed, which includes: generating audio data and video data to a speaker and a display panel respectively; receiving a microphone input from a microphone of the electronic device; detecting the microphone input to generate a user hot - spot detection result when the speaker plays the audio data and the display panel displays the video data; and storing the user hot - spot detection result. Brief Description of the Drawings

[0006] Figure 1 It is a schematic diagram of an electronic device according to an embodiment of the present invention.

[0007] Figure 2 It is a schematic diagram of a user hot - spot detection module according to an embodiment of the present invention.

[0008] Figure 3 A schematic diagram of the structure of an artificial intelligence model according to an embodiment of the present invention.

[0009] Figure 4 It is a schematic diagram of a user hot - spot detection result generated by a user hot - spot detection module according to an embodiment of the present invention.

[0010] Figure 5Schematic diagram of an audio / video content recognition module according to an embodiment of the present invention.

[0011] Figure 6 Schematic diagram of the audio content recognition result generated by the audio / video content recognition module according to an embodiment of the present invention.

[0012] Among them, the reference numerals are explained as follows:

[0013] 100 Electronic device

[0014] 110 Processing circuit

[0015] 112 Audio / video content generation circuit

[0016] 114 User hotspot detection module

[0017] 116 Audio / video content recognition module

[0018] 118 Output module

[0019] 120 Microphone

[0020] 130 Speaker

[0021] 140 Display panel

[0022] 210 Acoustic echo cancellation module

[0023] 220 Residual echo suppression module

[0024] 230 Emotion detection module

[0025] 232 MFCC feature acquisition module

[0026] 234 Artificial intelligence model

[0027] 236 Judgment module

[0028] 240 Voice activity detection module

[0029] 302 Convolutional layer

[0030] 304 Rectified linear unit layer

[0031] 306 Polling layer

[0032] 310 Residual block

[0033] 311 Convolutional layer

[0034] 312 Rectified linear unit layer

[0035] 313 Batch normalization layer

[0036] 314 Convolutional layer

[0037] 315 Modified Linear Unit Layer

[0038] 316 Adder

[0039] 320 Reshaping Layer

[0040] 322 Fully Connected Layer

[0041] 502 Difference Operator

[0042] 504 Difference-Difference Operator

[0043] 510 MFCC Feature Acquisition Module

[0044] 520 Artificial Intelligence Model

[0045] 530 Artificial Intelligence Model

[0046] 540 Judgment Module Detailed Implementation Manner

[0047] Figure 1 It is a schematic diagram of the electronic device 100 according to an embodiment of the present invention. As Figure 1 shown, the electronic device 100 includes a processing circuit 110, a microphone 120, a speaker 130, and a display panel 140. The processing circuit 110 includes an audio / video content generation circuit 112, a user hotspot detection module 114, an audio / video content recognition module 116, and an output module 118. In this embodiment, the electronic device 100 can be any type of device with an audio / video playback system, such as a television, a notebook computer, a tablet computer, a smart phone, or a desktop computer. In addition, the microphone 120 or the speaker 130 can be externally connected to the electronic device 100.

[0048] In the processing circuit 110 of the electronic device 100, the audio / video content generation circuit 112 is configured to generate audio data and video data for the speaker 130 and the display panel 140 respectively, so that the speaker 130 plays the audio data and the display panel 140 displays the video data. The user hotspot detection module 114 can be implemented by a processor executing program code (i.e., an algorithm) or a circuit. The user hotspot detection module 114 is used to receive the microphone input from the microphone 120, and when the speaker 130 plays the audio data and the display panel 140 displays the video data, it detects the human voice input by the microphone to generate a user hotspot detection result. The audio / video content recognition module 116 can be implemented by a processor executing program code or a circuit. The audio / video content recognition module 116 is used to recognize the audio / video content of the audio / video data to generate an audio / video content recognition result, where the audio content or the video content can be obtained from the audio / video content generation circuit 112. The output module 118 can be implemented by a processor executing program code or a circuit, and the output module 118 is used to receive the user hotspot detection result and the audio / video content recognition result to generate output information, where the output information can be stored in the storage device in the electronic device 100, or the output information can be transmitted to the server through the Internet.

[0049] Figure 2 It is a schematic diagram of the user hotspot detection module 114 according to an embodiment of the present invention. As Figure 2 shown, the user hotspot detection module 114 includes an Acoustic Echo Cancellation (AEC) module 210, a Residual Echo Suppression (RES) module 220, an emotion detection module 230, and a Voice Activity Detection (VAD) module 240. The emotion detection module 230 includes a Mel-scale Frequency Cepstral Coefficients (MFCC) feature acquisition module 232, an Artificial Intelligence (AI) model 234, and a judgment module 236.

[0050] In the operation of the user hotspot detection module 114, since the microphone input includes human voices, speaker sounds (echoes), and ambient noise, the acoustic echo cancellation module 210 and the residual echo suppression module 220 eliminate or reduce echoes and ambient noise to generate a clean microphone input. For example, the acoustic echo cancellation module 210 is based on an adaptive finite impulse response (FIR) filter, and the acoustic echo cancellation module 210 also calculates a residual signal that includes non-linear acoustic artifacts, and this signal is sent to the residual echo suppression module 220 to recover the microphone input signal. Since the design and detailed operation of the acoustic echo cancellation module 210 and the residual echo suppression module 220 are known to those skilled in the art, further description is omitted here.

[0051] The voice activity detection module 240, also known as speech activity detection or speech detection, is the detection of the presence or absence of human speech or human language. In the operation of the voice activity detection module 240, the clean microphone input is divided into many parts, and each part is calculated to obtain many features, and classification rules are applied to determine whether the part is speech or non-speech based on whether the value calculated based on the features exceeds a critical value, so as to classify the part as speech or non-speech. In addition, when the voice activity detection module 240 determines that the microphone input contains human speech or human voice, the voice activity detection module 240 also determines the intensity of the human voice. In this embodiment, the voice activity detection result may include information on whether the clean microphone input is a non-speech signal, a low-intensity human voice, a medium-intensity human voice, or a high-intensity human voice. Since the design and detailed operation of the voice activity detection module 240 are known to those skilled in the art, further description is omitted here.

[0052] In the operation of the emotion detection module 230, the MFCC feature acquisition module 232 performs some operations on the clean microphone input, such as Fourier transform, Mel-scale mapping, discrete cosine transform, etc., to generate MFCC features. Since the design and detailed operation of the MFCC feature acquisition module 232 are known to those skilled in the art, further description is omitted here. Then, the artificial intelligence model 234 is configured to receive the MFCC features to generate corresponding human emotions, where the artificial intelligence model 234 can analyze the MFCC features to determine whether the clean microphone input corresponds to anger, happiness, neutrality, sadness, silence, or other emotions. Then, the judgment module 236 generates a user emotion detection result indicating whether the clean microphone input corresponds to a positive emotion, a neutral emotion, or a negative emotion, where positive emotions include happiness, negative emotions include anger and sadness, and neutral emotions include emotions such as neutrality and silence.

[0053] Figure 3 is a schematic diagram of the structure of the artificial intelligence model 234 according to an embodiment of the present invention. AsFigure 3 As shown, the artificial intelligence model 234 includes a convolutional layer 302, a rectified linear unit (ReLU) layer 304, a pooling layer 306, a residual block 310, a reshaping layer 320, and a fully-connected (FC) layer 322. The residual block 310 includes a convolutional layer 311, a rectified linear unit layer 312, a batch normalization (BN) layer 313, a convolutional layer 314, a rectified linear unit layer 315, and an adder 316. The structure of the artificial intelligence model 234 is known to those skilled in the art, and this embodiment focuses on the training and use of the artificial intelligence model 234. During the training phase of the artificial intelligence model 234, engineers provide many audio segments for training, testing, and validation. The audio segments include the above-mentioned emotions, such as anger, happiness, neutrality, sadness, or silence. In addition, these audio segments can be processed under some data augmentation techniques to increase diversity, such as track length normalization, pitch shifting, time stretching, data shifting in the time domain, and / or noise addition.

[0054] Figure 4 Schematic diagram of the user hot spot detection result generated by the user hot spot detection module 114 according to an embodiment of the present invention. As Figure 4 shown, the user hot spot detection result includes the voice activity detection result generated by the voice activity detection module 240 and the user emotion detection result generated by the emotion detection module 230. Specifically, the user hot spot detection result has information about the user emotion and the corresponding time sequence. For example, the user hot spot detection result indicates that the user has a positive emotion at 00:30:00 in the video, the user has a positive emotion at 00:35:03 in the video, the user has a neutral emotion at 00:40:09 in the video, and the user has a negative emotion at 00:40:15 in the video.

[0055] Figure 5 Diagram of the audio / video content recognition module 116 according to an embodiment of the present invention. As Figure 5 shown, the audio / video content recognition module 116 includes an MFCC feature acquisition module 510, two artificial intelligence models 520 and 530, a difference operator 502, a difference-difference operator 504, and a judgment module 540.

[0056] In the operation of the audio / video content recognition module 116, the MFCC feature acquisition module 510 performs some operations on the audio content, such as Fourier transform, Mel scale mapping, discrete cosine transform, etc., to generate MFCC features. Then, the artificial intelligence model 520 is configured to receive the MFCC features to generate corresponding audio content recognition, where the artificial intelligence model 520 can analyze the MFCC features to determine whether the audio content to be played by the speaker 130 corresponds to happy, sad, angry or other content. The structure and training steps of the artificial intelligence model 520 are similar to Figure 3 the artificial intelligence model 234 shown, so the details will not be elaborated. In addition, to improve the recognition accuracy, the artificial intelligence model 530 is provided to use additional features, such as the outputs of the difference operator 502 and the difference-difference operator 504, to determine whether the audio content played by the speaker 130 corresponds to happy, sad, angry or other content. Then, the judgment module 540 generates an audio / video content recognition result indicating whether the audio content corresponds to happy, sad, angry or other content according to the outputs of the artificial intelligence models 520 and 530.

[0057] In addition, Figure 5 the audio / video content recognition module 116 shown is only for illustration and not a limitation of the present invention. In other embodiments of the present invention, the difference operator 502, the difference-difference operator 504 and the artificial intelligence model 530 can be removed from the audio / video content recognition module 116, and the judgment module 540 generates an audio content recognition result only according to the output of the artificial intelligence model 520. Such a design change should belong to the scope of the present invention.

[0058] Figure 6 FIG. is a schematic diagram of the audio content recognition result generated by the audio / video content recognition module 116 according to an embodiment of the present invention. As Figure 6 shown, the audio content recognition result has information on the audio content and the corresponding time sequence. For example, the audio content recognition result indicates that the audio content has happy content at the time point t1 of the video, sad content at the time point t2 of the video, and angry content at the time point t3 of the video.

[0059] In summary, by obtaining the user hotspot detection result through the user hotspot detection module 114, the electronic device 100 can learn about the user's reaction when watching a video. Additionally, by obtaining the audio / video content recognition result using the audio / video content recognition module 116, the electronic device 100 can learn about the classification of video segments. The user hotspot detection result and / or the audio / video content recognition result can be used by video application companies to provide better services for users. For example, video application companies or streaming services can insert appropriate advertisements or personalized recommendation information at some specific times of the video.

[0060] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A processing circuit of an electronic device, characterized in that: An audio / video content generation module, used for generating audio data and video data and transmitting them to a speaker and a display panel respectively; A user hotspot detection module, configured to receive a microphone input from a microphone of the electronic device and detect the microphone input to generate a user hotspot detection result when the speaker plays the audio data and the display panel displays the video data; as well as The output module is used to store the user hotspot detection result.

2. The processing circuit according to claim 1, wherein: The user hotspot detection module includes: an acoustic echo cancellation module for canceling or reducing the echo or ambient noise of the microphone input to produce a clean microphone input; and The emotion detection module is used to generate a user emotion detection result indicating which user emotion the clean microphone input corresponds to.

3. The processing circuit according to claim 2, wherein: The emotion detection module comprises: A Mel-scale frequency cepstral coefficient feature acquisition module, which is used to receive the clean microphone input to generate MFCC features; An artificial intelligence model, for receiving the MFCC features to generate corresponding user emotions; as well as A judgment module is used to generate a user emotion detection result based on the user emotion judged by the artificial intelligence model.

4. The processing circuit according to claim 2, wherein: The user hotspot detection module also includes: The voice activity detection module is used to detect whether the clean microphone input contains human voice or human speech to generate a voice activity detection result.

5. The processing circuit of claim 4, wherein: The user hotspot detection module generates the user hotspot detection result according to the user emotion detection result and the voice activity detection result.

6. The processing circuit of claim 5, wherein: The user hotspot detection result includes information about the user's emotions and the timing of the corresponding audio / video content.

7. The processing circuit of claim 1, further comprising: An audio / video content recognition module, used to recognize the audio content corresponding to the audio data to generate an audio / video content recognition result; The output module also stores the audio / video content recognition result.

8. The processing circuit of claim 7, wherein: The audio / video content recognition module comprises: An MFCC feature acquisition module, used for receiving the audio content to generate MFCC features; An artificial intelligence model, for receiving the MFCC features to determine corresponding content; as well as A judgment module is used to generate the audio / video content recognition result according to the content judged by the artificial intelligence model.

9. A processing method for an electronic device, characterized in that: Generating audio data and video data to a speaker and a display panel respectively; receiving a microphone input from a microphone of the electronic device; When the speaker plays the audio data and the display panel displays the video data, detecting the microphone input to generate a user hotspot detection result; as well as The user hotspot detection result is stored.

10. The processing method according to claim 9, wherein: When the speaker plays the audio data and the display panel displays the video data, the step of detecting the microphone input to generate the user hotspot detection result includes: Eliminating or reducing echo or ambient noise of the microphone input to produce a clean microphone input; as well as A user emotion detection result is generated indicating which user emotion the clean microphone input corresponds to.