Remote AI control method and device for multimedia, and storage medium

By using a microphone array and deep learning model on the multimedia control box to initially distinguish sounds, and then combining it with the AI ​​model of the cloud computing control system, the problem of the audio system being unable to distinguish between its own output sound and ambient sound is solved, achieving efficient and accurate remote audio control and management.

CN120673769APending Publication Date: 2025-09-19CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510976611.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing audio systems cannot effectively distinguish between their own output sound and ambient sound, resulting in low control accuracy. Remote control methods also have problems such as limited distance, low security, poor compatibility, poor transmission quality, and inefficient equipment management.

Method used

A microphone array is used to collect sound signals and audio output signals of multimedia devices, which are initially distinguished at the multimedia control box end through a deep learning model. Further analysis is carried out in combination with the AI ​​model of the cloud computing control system to achieve accurate control of multimedia devices.

Benefits of technology

It improves the accuracy of distinguishing ambient sound from device output sound, enhances the adaptability of the audio system and the stability of remote control, and enhances the efficiency of device management and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673769A_ABST
    Figure CN120673769A_ABST
Patent Text Reader

Abstract

The invention discloses a remote AI control method and device for multimedia and a computer readable storage medium, and belongs to the technical field of multimedia equipment control. The control method comprises the following steps: acquiring a sound signal acquired by a microphone array and an audio output signal of multimedia equipment; respectively extracting features of the sound signal and the audio output signal; the extracted features are input into a deep learning model, a classification result is obtained, and the classification result comprises environment sound or sound output by the multimedia equipment; the extracted features and the classification result are sent to a cloud computing control system, playing of the multimedia device is controlled according to a first feedback instruction of the cloud computing control system, and the first feedback instruction is obtained by the cloud computing control system through analyzing the received features and the classification result based on an AI model. According to the method, the problems that the output sound of the sound equipment and the environment sound cannot be effectively distinguished and the control accuracy is low in the related technology can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimedia device control, and specifically relates to a multimedia remote AI control method and device, and a computer-readable storage medium. Background Art

[0002] The rapid development of multimedia technology has led to higher demands for the intelligence and adaptability of audio systems. Advances in sensor technology, computing power, and communications technology have transformed audio systems from mere sound playback devices to complex systems capable of intelligently interacting with the environment and users.

[0003] When it comes to sound processing, traditional audio systems often require manual adjustments to parameters like volume and sound effects, and are unable to automatically adapt to changes in the environment. In recent years, some research has begun to explore the use of sound sensors to detect ambient noise and adjust audio parameters accordingly. However, these solutions suffer from low detection accuracy and an inability to effectively distinguish between the speaker's own sound output and ambient sound, resulting in low control accuracy for the audio system. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to address the above-mentioned shortcomings of the existing technology and provide a multimedia remote AI control method and device, and a computer-readable storage medium, which can effectively distinguish the sound output by the multimedia device itself and the ambient sound, thereby accurately controlling the multimedia device.

[0005] In a first aspect, the present invention provides a remote AI control method for multimedia, which is applied to a multimedia control box. The method includes: obtaining sound signals collected by a microphone array and audio output signals of a multimedia device; extracting features of the sound signal and the audio output signal respectively; inputting the extracted features into a deep learning model to obtain classification results, which include ambient sounds or the output sounds of the multimedia device itself; sending the extracted features and classification results to a cloud computing control system, and controlling the playback of the multimedia device according to a first feedback instruction of the cloud computing control system, wherein the first feedback instruction is obtained by the cloud computing control system based on the AI ​​model analysis of the received features and classification results.

[0006] In some embodiments, after inputting the extracted features into the deep learning model to obtain the classification results, and before sending the extracted features and classification results to the cloud computing control system, the remote AI control method of multimedia also includes: calculating the feature differences between the sound signal and the audio output signal; based on the feature differences between the sound signal and the audio output signal, as well as the sound wave propagation characteristics, judging whether the sound signal is ambient sound or the sound output by the multimedia device itself, to verify the classification results.

[0007] In some embodiments, the deep learning model includes a discriminative function D(S, E).

[0008]

[0009] Where S=(s1,s2,…,s n ) is the feature vector of the multimedia device’s own sound output, E=(e1,e2,…,e m ) is the characteristic vector of the ambient sound, w si and w ej are the characteristics of the sound output by the multimedia device itself i and the characteristics of the ambient sound j The weight of .

[0010] In response to D(S, E) being greater than 0, the classification result is that the multimedia device itself outputs the sound;

[0011] In response to D(S, E) being less than 0, the classification result is environmental sound.

[0012] In some embodiments, the remote AI control method for multimedia also includes: obtaining audio and video signals of the current environment, and using a deep learning model to extract a first key feature of the audio and video signal; matching a first audio and video set based on the first key feature, wherein the first audio and video set refers to an audio and video set associated with the content type; sending the first audio and video set and the first key feature to the cloud computing control system to receive a second feedback instruction from the cloud computing control system; determining a second audio and video set based on the second feedback instruction and the first audio and video set; and controlling the playback of the multimedia device based on the second audio and video set.

[0013] In some embodiments, matching a first audio and video set based on a first key feature specifically includes: matching a third audio and video set based on the first key feature; evaluating the audio and video in the third audio and video set based on a machine learning model; prioritizing the audio and video in the third audio and video set according to the evaluation results; and screening out at least one audio and video having a priority greater than a threshold to generate a first audio and video set.

[0014] In some embodiments, the remote AI control method for multimedia also includes: obtaining audio and video signals of the current environment, and using a deep learning model to extract a second key feature of the audio and video signal; estimating a first frequency range based on the second key feature and the machine learning model, wherein the first frequency range refers to the frequency range corresponding to the content type; sending the first frequency range and the second key feature to the cloud computing control system to receive a third feedback instruction from the cloud computing control system; determining a second frequency range based on the third feedback instruction and the first frequency range; and controlling the playback of the multimedia device according to the second frequency range.

[0015] In some embodiments, the remote AI control method for multimedia also includes: sending the performance indicators of the deep learning model and the machine learning model during the calculation process to the cloud computing control system to receive the model parameters issued by the cloud computing control system, and communicating with the cloud computing control system based on the communication network, wherein the communication network includes 5G network and 6G network.

[0016] In the second aspect, the present invention also provides a remote AI control method for multimedia, which is applied to a cloud computing control system. The remote AI control method for multimedia includes: receiving features and classification results extracted by a multimedia control box, wherein the classification results are obtained by the multimedia control box performing feature extraction on the sound signals collected by the microphone array and the audio output signals of the multimedia device, and inputting the extracted features into a deep learning model, and the classification results include ambient sounds or the sound output by the multimedia device itself; analyzing the received features and classification results based on the AI ​​model to obtain a first feedback instruction; and sending the first feedback instruction to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

[0017] In some embodiments, the remote AI control method of multimedia also includes: receiving current environmental data and user preference parameters sent by the multimedia control box, wherein the environmental data includes at least one of the following: environmental photos, light intensity, and number of people; analyzing the current environmental data and the user preference parameters based on the AI ​​model to obtain audio parameters and / or digital people that match the current environment; sending the audio parameters and / or digital people to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

[0018] In a third aspect, the present invention also provides a multimedia remote AI control device, comprising an acquisition module, an edge computing module, and a first communication module.

[0019] An acquisition module is used to acquire the sound signal collected by the microphone array and the audio output signal of the multimedia device. An edge computing module is connected to the acquisition module and is used to extract the features of the sound signal and the audio output signal respectively, and input the extracted features into the deep learning model to obtain a classification result, which includes the environmental sound or the sound output by the multimedia device itself. A first communication module is connected to the edge computing module and is used to send the extracted features and classification results to the cloud computing control system, and control the playback of the multimedia device according to the first feedback instruction of the cloud computing control system, wherein the first feedback instruction is obtained by the cloud computing control system based on the AI ​​model analysis of the received features and classification results.

[0020] In a fourth aspect, the present invention also provides a multimedia remote AI control device, comprising a second communication module and a computing module.

[0021] The second communication module is configured to receive features and classification results extracted by the multimedia control box. The classification results are obtained by the multimedia control box extracting features from the sound signals collected by the microphone array and the audio output signals of the multimedia device, and inputting the extracted features into a deep learning model. The classification results include ambient sound or the sound output by the multimedia device itself. A computing module, connected to the second communication module, is configured to analyze the received features and classification results based on the AI ​​model to obtain a first feedback instruction. The second communication module is also configured to send the first feedback instruction to the multimedia control box so that the multimedia control box can control the playback of the multimedia device.

[0022] In a fifth aspect, the present invention also provides a multimedia remote AI control device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the multimedia remote AI control method as described in the first aspect or the second aspect.

[0023] In a sixth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, it implements the remote AI control method for multimedia as in the first aspect or the second aspect.

[0024] The present invention provides a multimedia remote AI control method and device, and a computer-readable storage medium. The method obtains sound signals collected by a microphone array, including ambient sound and / or the self-output sound of a multimedia device, and uses the audio output signal of a known multimedia device as a reference signal. A deep learning model is used on the multimedia control box to distinguish whether it is ambient sound or the self-output sound of the multimedia device. The collected sound signal features, the audio output signal features of the multimedia device, and the distinction results are then sent to a cloud computing control system for further analysis using an AI model to determine the final control instructions to control the playback of the multimedia device. Since the audio output signals of the multimedia devices are obtained simultaneously for comparative analysis, and an artificial intelligence model is applied on the multimedia control box to perform a preliminary distinction between the sound source (ambient sound or the multimedia device's own audio output), the effect of quickly and accurately distinguishing the sound is achieved. Then, a more complex artificial intelligence model and global data in the cloud are used to enhance decision-making, further improving the accuracy of the distinction, which is conducive to accurate control of the multimedia device. In short, through the collaborative mechanism of local deep feature classification and cloud global optimization, the accuracy of distinguishing between ambient sound and device output sound can be improved in complex acoustic scenarios (spectral overlap, low signal-to-noise ratio, dynamic reverberation). BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flowchart of a multimedia remote AI control method according to embodiment 1 of the present invention;

[0026] Figure 2 This is a schematic diagram of a smart home environment according to embodiment 3 of the present invention;

[0027] Figure 3 This is a schematic diagram of the integration of a multimedia control box and a cloud computing control system according to embodiment 3 of the present invention;

[0028] Figure 4 This is a schematic diagram of information processing of a cloud computing control system according to embodiment 3 of the present invention;

[0029] Figure 5 This is a structural diagram of a multimedia control box according to embodiment 3 of the present invention;

[0030] Figure 6 This is a schematic diagram of a model for distinguishing device sound from ambient sound according to embodiment 3 of the present invention;

[0031] Figure 7 Schematic diagram of a process for calculating A(C) according to embodiment 3 of the present invention;

[0032] Figure 8 Schematic diagram of a process for calculating F(C) according to embodiment 3 of the present invention. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0034] It should be understood that the specific embodiments and drawings described herein are only used to explain the present invention rather than to limit the present invention.

[0035] It is understood that, in the absence of conflict, the various embodiments of the present invention and the various features in the embodiments may be combined with each other.

[0036] It can be understood that, for the convenience of description, the drawings of the present invention only show parts related to the present invention, while parts unrelated to the present invention are not shown in the drawings.

[0037] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may be integrated into one physical structure.

[0038] It will be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.

[0039] It is understood that the flowcharts and block diagrams of the present invention illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to various embodiments of the present invention. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented using a hardware-based system that implements the specified functions, or may be implemented using a combination of hardware and computer instructions.

[0040] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented by software or hardware. For example, the units and modules can be located in a processor.

[0041] Example 1:

[0042] Multimedia devices include audio equipment, wireless headphones, portable speakers, smart TVs, cameras, video conferencing equipment, virtual reality / augmented reality equipment, in-vehicle multimedia equipment, etc.

[0043] The multimedia device of this embodiment is described by taking a stereo or a smart TV as an example.

[0044] In some related technologies, the control method of the audio has the following problems:

[0045] (1) Problems based on Bluetooth connection:

[0046] ① Limited distance: The effective transmission distance of Bluetooth technology is relatively short, generally about 10 meters. When the distance is exceeded, the connection signal will become weak, resulting in poor control effect or even disconnection, making it impossible to achieve stable long-distance control.

[0047] ② Low security: Any device with Bluetooth function can connect to the speaker through search, and the connection information will be automatically saved. There is a risk of connection and control by unauthorized devices, resulting in problems such as tampering of the speaker playback content or leakage of personal privacy.

[0048] ③ Compatibility issues: There may be differences in Bluetooth compatibility between devices of different brands and models, resulting in unstable connections or certain functions not working properly.

[0049] ④ Slow data transmission: In actual applications, it is difficult to achieve due to the compatibility between devices and other factors. It is usually around 100KB and can only be used for audio control but not video control.

[0050] (2) Issues related to Wi-Fi connection:

[0051] ① Network dependence: The speaker needs to be connected to a Wi-Fi network to achieve remote control. If the network is unstable or fails, the speaker will not be able to be controlled normally, affecting the user experience.

[0052] ② Complex configuration: Compared with Bluetooth connection, the setup process of Wi-Fi connection is relatively complicated, requiring users to perform network configuration, enter passwords, and other operations, which is not user-friendly for users who are not familiar with technology.

[0053] ③ Security risks: Although Wi-Fi networks can be password-protected, if the password is cracked or the network is hacked, the audio equipment may also be attacked, posing certain security risks.

[0054] (3) Audio transmission quality issues:

[0055] ①Signal interference: When connected via Bluetooth and Wi-Fi, the wireless transmission process is easily interfered with by other wireless signals in the surrounding environment, such as microwave ovens, wireless routers, etc., which may cause the audio signal to be distorted, stuck or interrupted, affecting the sound quality and playback smoothness.

[0056] ②Data compression: To improve transmission efficiency, some remote control technologies compress audio data, which may lead to a decline in audio quality, making the sound less clear and full, and losing some details and dynamic range.

[0057] ③ Intelligent management and control: Currently, control devices are unable to intelligently select appropriate audio content output based on the on-site environment, thus affecting user satisfaction.

[0058] (4) Equipment compatibility and standardization issues:

[0059] ① Lack of unified standards: Currently, there is no unified standard or specification for remote audio and video control technology. Products from different manufacturers differ in functions, protocols, interfaces, etc. This brings inconvenience to users and limits the interoperability and collaboration between devices of different brands.

[0060] ② Difficulty in system integration: When building multi-device audio and large-screen devices, the compatibility issues of different remote control boxes may increase the difficulty of system integration, requiring more time and effort for debugging and configuration to achieve coordination between various devices.

[0061] (5) Equipment control issues: When multimedia control is required at multiple locations, multiple main control rooms and management personnel are required, which results in high construction and maintenance costs.

[0062] (6) Inefficient equipment status management: When audio and video are playing, operators cannot know whether the controlled equipment is playing normally. They need to arrange personnel to patrol the site, which leads to low management efficiency and inability to correct problems in a timely manner.

[0063] (7) Low accuracy in environmental noise detection: Noise detection sensors in related technologies are often not sensitive and accurate enough to accurately capture subtle changes in environmental noise, resulting in inaccurate adjustments to the sound system.

[0064] (8) Lack of comprehensive consideration of environmental factors: Most of them only focus on sound factors, ignoring the impact of other environmental factors such as light and personnel distribution on the sound effects, and cannot provide a fully optimized user experience.

[0065] (9) Poor sound differentiation: In terms of distinguishing the sound output by the speaker itself from the ambient sound, the relevant technology is single and not effective enough, which can easily lead to misjudgment and inaccurate adjustments.

[0066] (10) Poor coordination between cloud and local devices: The coordination between cloud computing and local devices is imperfect. Network delays and data transmission limitations affect the real-time and accuracy of control, making it impossible to respond quickly to environmental changes.

[0067] (11) Limited adaptive capabilities: The adaptive adjustment functions of existing audio systems are relatively simple and limited, and cannot be flexibly and intelligently optimized according to complex and changing environments and user needs.

[0068] like Figure 1 As shown, this embodiment provides a remote AI (Artificial Intelligence) control method for multimedia, which is applied to a multimedia control box to solve the problems existing in the above-mentioned related technologies one by one.

[0069] Among them, the remote AI control method of multimedia includes:

[0070] S11, obtaining a sound signal collected by a microphone array and an audio output signal of a multimedia device.

[0071] S12, extracting features of the sound signal and the audio output signal respectively.

[0072] S13, inputting the extracted features into a deep learning model to obtain a classification result, which includes the ambient sound or the sound output by the multimedia device itself.

[0073] S14, sending the extracted features and classification results to the cloud computing control system, and controlling the playback of the multimedia device according to the first feedback instruction of the cloud computing control system, wherein the first feedback instruction is obtained by the cloud computing control system based on the AI ​​model analysis of the received features and classification results.

[0074] In this embodiment, the multimedia control box is located on the multimedia device side, the cloud computing control system is located in the cloud, and the cloud computing control system is wirelessly connected to the user terminal, which includes a smartphone, computer, etc. The user can connect to the cloud computing control system through the client to achieve remote control of the multimedia device. It should be noted that the multimedia remote AI control method of this embodiment can also automatically remotely control the multimedia device in real time based on the current environment of the multimedia device, that is, the user does not need to control it through the client.

[0075] The following is an example of a multimedia control box: it uses a brand-new card motherboard based on the micro Linux system, integrating a 5G module, an audio module, an HDMI output module (which can output audio and video signals), a sound acquisition module, a video acquisition module, and an edge AI chip (i.e., an edge computing module). The integrated design makes the smart multimedia control box only the size of a palm, and the collaboration between the various modules is more efficient, improving overall performance. It is used to overcome the above-mentioned defects in the control methods and functions of audio and large-screen devices in related technologies. Through the innovative design of connecting cloud computing resources and AI models with 5G modules and edge computing modules, efficient remote transmission control of multimedia devices and AI intelligent output of appropriate content can be achieved.

[0076] The microphone array in this embodiment includes at least one microphone, and the sound signals collected by the microphone array include ambient sounds and / or signals output by the multimedia device itself. The audio output signal of the multimedia device is known content. The extracted features include frequency features, time difference features, and phase difference features, such as the time difference and phase difference features of the sound signals arriving at different microphones. The deep learning layer of the deep learning model uses a convolutional neural network (CNN) or a recurrent neural network (RNN) to learn and classify the extracted features, and trains the neural network to learn and distinguish the patterns of the multimedia device's own output sound and the ambient sound. During the training process of the deep learning model, a large amount of sample data including the multimedia device's own output sound and various ambient sounds needs to be collected. The first feedback instruction includes echo cancellation (such as eliminating echoes in remote meetings), noise suppression (such as enhancing human voices and suppressing background noise), voiceprint recognition (used for security or multimedia device wake-up, such as making the multimedia device respond only to ambient human voices to avoid false wake-ups), turning up the volume, etc.

[0077] In this embodiment, the audio output signal of the multimedia device is obtained synchronously for reference and comparative analysis, and an artificial intelligence model (such as a deep learning model) is applied on the multimedia control box to perform a preliminary distinction between the sound source (ambient sound or the multimedia device's own audio output), thereby achieving the effect of quickly and accurately distinguishing the sound. Then, the more complex artificial intelligence AI model and global data on the cloud are used for enhanced decision-making, further improving the accuracy of the distinction and facilitating accurate control of the multimedia device. In short, through the collaborative mechanism of local deep feature classification and cloud-based global optimization, the accuracy of distinguishing between ambient sound and device output sound is improved in complex acoustic scenarios (spectral overlap, low signal-to-noise ratio, dynamic reverberation).

[0078] In some embodiments, after S11 and before S12, the multimedia remote AI control method further includes: preprocessing the sound signal collected by the microphone array, the preprocessing including at least one of the following: filtering, noise reduction, and normalization. In S12, the sound signal is converted to the frequency domain using a fast Fourier transform (FFT) to extract the frequency characteristics of the sound signal.

[0079] In some embodiments, after inputting the extracted features into the deep learning model to obtain the classification results, and before sending the extracted features and classification results to the cloud computing control system, the remote AI control method of multimedia also includes: calculating the feature differences between the sound signal and the audio output signal; based on the feature differences between the sound signal and the audio output signal, as well as the sound wave propagation characteristics, judging whether the sound signal is ambient sound or the sound output by the multimedia device itself, to verify the classification results.

[0080] In this embodiment, if the judgment result is inconsistent with the classification result of the deep learning model, the judgment result is adopted or adaptive calibration is triggered (such as readjusting the model weight). The judgment result obtained based on the feature difference can provide physical layer verification, which is a comprehensive decision, which relies on the classification result of the deep learning model and the threshold verification of the feature difference value. Since this embodiment adopts double verification, misjudgment can be avoided (for example, deep learning may not accurately classify some edge samples, and the difference value is required to assist in decision-making). For example, when the speaker is playing low-frequency music, there is air-conditioning noise of similar frequency in the environment. If only the deep learning model is used, it may be misjudged as the speaker output. However, combined with the delay difference (the arrival time of the air-conditioning noise is inconsistent with the speaker output), the misjudgment can be corrected, thereby improving the accuracy of control.

[0081] In some embodiments, the deep learning model includes a discriminative function D(S, E).

[0082]

[0083] Where S=(s1,s2,…,s n) is the feature vector of the multimedia device’s own sound output, E=(e1,e2,…,e m ) is the characteristic vector of the ambient sound, w si and w ej are the characteristics of the sound output by the multimedia device itself i and the characteristics of the ambient sound j The weight of

[0084] In response to D(S, E) being greater than 0, the classification result is that the multimedia device itself outputs the sound;

[0085] In response to D(S, E) being less than 0, the classification result is environmental sound.

[0086] In some embodiments, the remote AI control method for multimedia further includes:

[0087] S21, obtaining the audio and video signals of the current environment, and using the deep learning model to extract the first key features of the audio and video signals.

[0088] S22: Match a first audio and video set according to the first key feature, where the first audio and video set refers to an audio and video set associated with the content type.

[0089] S23: Send the first audio and video set and the first key feature to the cloud computing control system to receive a second feedback instruction from the cloud computing control system.

[0090] S24: Determine a second audio and video set according to the second feedback instruction and the first audio and video set.

[0091] S25: Control the playback of the multimedia device according to the second audio and video set.

[0092] In some embodiments, matching a first audio and video set based on a first key feature specifically includes: matching a third audio and video set based on the first key feature; evaluating the audio and video in the third audio and video set based on a machine learning model; prioritizing the audio and video in the third audio and video set according to the evaluation results; and screening out at least one audio and video having a priority greater than a threshold to generate a first audio and video set.

[0093] In this embodiment, the multimedia control box requires more complex content analysis and a richer audio and video database. With the help of the edge computing module, deep learning technology is used to more accurately understand the content and match audio and video, and output appropriate audio and video and audio and video frequencies. Assume that an audio and video a∈A(C) has a frequency range f=F(C). In actual implementation, the cloud computing control system obtains A(C) and F(C) in a more complicated way, and needs to be determined through cloud resource database query, model calculation, etc. The following is a process for the multimedia control box to speed up the calculation of A(C). Specifically, if Figure 7 As shown, the process includes:

[0094] S31, the sound acquisition module and the video acquisition module obtain the audio and video signals of the current environment (that is, the edge computing module completes data input). The edge computing module uses a deep learning model to analyze the obtained audio and video signals (feature extraction and analysis) to obtain content descriptions (such as recognizing sound or images as descriptions of text content), labels (such as strong / weak light, cheerful emotions, number of boys / number of girls in the personnel distribution), and metadata.

[0095] S32, the edge computing module uses a deep learning model to extract the first key features of the audio and video signals (model calculation), and the first key features include keywords, emotional tendencies, and themes.

[0096] S33, local database query: In the local database of the edge computing module, a quick query is performed based on the first key feature to match the third audio and video set associated with content type C, where content type C includes one of the following: happy, sad, tense, and calm.

[0097] S34, Local Model Evaluation: The edge computing module uses the locally pre-trained machine learning model to evaluate the audio and video in the third audio and video set. Specifically, the audio and video features are scored based on their matching degree with the content features of the current environment to obtain an evaluation result.

[0098] S35, priority sorting: based on the evaluation result, the audio and video in the third audio and video set are prioritized, and at least one audio and video with a priority greater than a threshold (ie, a high matching degree) is screened out to generate a first audio and video set.

[0099] S36, interacting with the cloud: The 5G module sends the first audio and video set and the first key feature calculated by the edge computing module to the cloud computing control system (i.e., the cloud), and at the same time obtains the latest audio and video data and model update parameters from the cloud and passes them to the edge computing module.

[0100] S37, Final Selection and Output: The edge computing module integrates the second feedback instruction returned by the cloud and the local calculation results (i.e., the first audio and video set) to determine the second audio and video set (i.e., the final suitable audio A (C)), and outputs it to the multimedia device via the HDMI output interface for playback. Examples of the second feedback instruction include: deleting individual audio and video in the first audio and video set, adding at least one audio and video to the first audio and video set, or adjusting the melody and timbre of individual audio and video in the first audio and video set.

[0101] S38, Learning and Optimization: The edge computing module records the process and results of each calculation for subsequent local model training and optimization to improve the accuracy and speed of future calculations.

[0102] By utilizing the fast processing capabilities and local resources of the edge computing module through the S31-S38 process, the calculation and selection process of A(C) can be significantly accelerated, reducing reliance on the cloud and improving the response speed and efficiency of multimedia device control. Even if there is a communication interruption between the multimedia control box and the cloud computing control system, the multimedia device can still be controlled to play normally, improving the stability of control. It should be noted that A(C) represents an audio and video collection associated with a specific content type C. For example, if A(C) is an audio collection, it is an ordered or unordered collection composed of multiple elements. Each element is a string that uniquely identifies a specific audio file, which can be in the form of a file path (such as "C:\music\happy.mp3") or a specific audio identifier (such as "ID12345"). For different content types, such as "happy content," "sad content," and "stressful content," the composition of A(C) varies. Taking "happy content" as an example, the audio in the collection may have the following characteristics:

[0103] (1) The rhythm is brisk, usually with a high beats per minute (BPM), such as between 120-180.

[0104] (2) The melody mostly uses major keys and bright note combinations.

[0105] (3) The choice of musical instruments may include piano, guitar, drums and other dynamic instruments.

[0106] The process of determining the elements in A(C) in the edge computing module is as follows:

[0107] (1) Conduct large-scale data collection, including samples of various content types and their matching audio.

[0108] (2) Use audio feature extraction technology to extract features such as rhythm, melody, harmony, and timbre.

[0109] (3) Audios with similar features are classified into corresponding content type sets through manual labeling or machine learning-based classification algorithms.

[0110] (4) Continuously evaluate and optimize, and adjust the audio elements in the collection based on user feedback and actual application effects to improve matching accuracy and satisfaction.

[0111] A(C) is not static; instead, it is continuously updated and optimized as new audio is generated, user needs change, and technology advances. New audio can be added to the appropriate set through the same feature extraction and classification process, while some audio that is no longer suitable or performs poorly may be removed from the set.

[0112] In some embodiments, the remote AI control method for multimedia further includes:

[0113] S41, obtaining the audio and video signals of the current environment, and using the deep learning model to extract the second key features of the audio and video signals.

[0114] S42: Estimate a first frequency range based on the second key feature and the machine learning model, where the first frequency range refers to a frequency range corresponding to the content type.

[0115] S43: Send the first frequency range and the second key feature to the cloud computing control system to receive a third feedback instruction from the cloud computing control system.

[0116] S44: Determine a second frequency range according to the third feedback instruction and the first frequency range.

[0117] S45: Control the playback of the multimedia device according to the second frequency range.

[0118] In this embodiment, the edge computing module of the multimedia control box is used to speed up the calculation of F(C). The first frequency range refers to the frequency range corresponding to each audio and video, that is, each audio and video has a corresponding first frequency range. It should be noted that S21-S25 can be executed simultaneously with S41-S45. Figure 8 As shown, S41-S45 specifically include:

[0119] S51, the sound acquisition module and the video acquisition module obtain the audio and video signals of the current environment (that is, the edge computing module completes data input). The edge computing module uses a deep learning model to analyze the obtained audio and video signals to obtain audio features, content labels, etc. (feature extraction).

[0120] S52, the edge computing module uses a deep learning model to extract the second key features of the audio and video signal. The second key features include the main frequency components, rhythm characteristics, melody patterns, etc. of the audio and video.

[0121] S53, model calculation: A trained lightweight machine learning model is pre-loaded in the edge computing module, which is used to quickly estimate the first frequency range of each audio and video in the first audio and video set based on the extracted features.

[0122] S54, interacting with the cloud computing control system: uploading the first frequency range and the second key feature to the cloud (ie, sending them to the cloud computing control system), and obtaining the latest model parameters and optimization strategy from the cloud.

[0123] S55, local adjustment: Further adjust and optimize the first frequency range locally based on the third feedback instruction returned by the cloud. The third feedback instruction includes adjusting the first frequency range of individual audio and video.

[0124] S56, result output: controlling the audio and video adjustment and playback of the multimedia device according to the determined second frequency range.

[0125] S57, real-time monitoring and feedback: The edge computing module continuously monitors performance indicators during the computing process, such as computing time and accuracy, and uploads this feedback information to the cloud to help improve models and algorithms.

[0126] The S51-S57 process primarily leverages the local processing power of the edge computing module, combined with collaboration with the cloud. This significantly accelerates F(C) calculations while ensuring accuracy, enabling more timely and efficient audio and video processing. Even when communication between the multimedia control box and the cloud computing control system is interrupted, normal playback of multimedia devices can be controlled, improving control stability. F(C) represents the frequency range corresponding to a specific content type (C). The frequency range covered by F(C) varies significantly for different content types, based on an understanding of the relationship between human perception and emotional response and audio frequencies. For example, if content type C is "upbeat," F(C) might be defined as a higher frequency range, such as 500 Hz to 2000 Hz. This frequency range is often associated with a bright and lively feeling, enhancing a cheerful atmosphere. In this case, higher-frequency sound components are relatively prevalent, potentially including high-pitched instruments and clear, bright vocals, creating a vibrant and positive listening experience. For "sad content," F(C) might be set to a lower frequency range, such as from 200 Hz to 800 Hz. Audio in this range often gives people a deep, oppressive feeling, echoing the emotion of sadness. It may include elements such as low strings and heavy vocals to trigger emotional resonance in the listener. Determining the specific range of F(C) usually requires comprehensive consideration of many factors, including psychological research results, music theory, extensive user feedback, and actual audio effect testing. Through continuous adjustment and optimization, F(C) can accurately reflect the emotions and atmosphere that different content types hope to convey, providing users with an audio experience that is more in line with the content.

[0127] In some embodiments, the remote AI control method for multimedia also includes: sending the performance indicators of the deep learning model and the machine learning model during the calculation process to the cloud computing control system to receive the model parameters issued by the cloud computing control system.

[0128] In this embodiment, the multimedia control box receives model parameters issued by the cloud computing control system, which can optimize the local artificial intelligence model and ultimately improve the calculation speed, thereby improving the response speed of the multimedia control box to the control of the multimedia device.

[0129] In some embodiments, the remote AI control method for multimedia also includes: communicating with a cloud computing control system based on a communication network, wherein the communication network includes a 5G network and a 6G network.

[0130] In this embodiment, 5G or 6G or other new generation communication networks are used to transmit data, which has the effects of fast transmission speed, low latency, anti-interference, and no data compression.

[0131] In some embodiments, the remote AI control method for multimedia also includes: inferring the user's potential needs based on the current human activity and time information in the environment, and then controlling the playback of the multimedia device. For example, if it is evening and there are fewer people, soothing music and a lower volume may be recommended.

[0132] In some embodiments, the remote AI control method for multimedia also includes presenting a digital human service on the display interface to enable user interaction. Specifically, the cloud selects an appropriate digital human image and voice style based on the user's preferences and the current scenario. For example, in home entertainment mode, a lively and interesting digital human image might be presented, interacting with the user in a cheerful tone, and then transmitted to the multimedia control box.

[0133] The multimedia remote AI control method of this embodiment, because the audio output signal of the multimedia device is obtained synchronously for comparative analysis, and an artificial intelligence model is applied on the multimedia control box to perform a preliminary distinction between the sound source (ambient sound or the multimedia device's own audio output), it achieves the effect of quickly and accurately distinguishing the sound, and then uses the more complex artificial intelligence model and global data on the cloud to enhance the decision-making, further improving the accuracy of the distinction, which is conducive to the accurate control of the multimedia device. In short, through the collaborative mechanism of local deep feature classification and cloud global optimization, the accuracy of distinguishing between ambient sound and device output sound is improved in complex acoustic scenarios (spectral overlap, low signal-to-noise ratio, dynamic reverberation). In addition, this control method can improve the precision and accuracy of environmental noise detection, enabling the cloud computing control system to make precise adjustments based on subtle changes in environmental noise. Furthermore, the control method comprehensively considers multiple environmental factors, such as light, personnel distribution, etc., to provide a more comprehensive and optimized sound effect and enhance the user experience. It also uses 5G or 6G network communication to optimize the collaborative work between cloud computing and local devices, reduce the impact of network delay, and achieve fast, real-time and accurate control. Finally, this control method also enhances the adaptive capability of the cloud computing control system, enabling it to flexibly and intelligently respond to various complex environments and user needs, and provide personalized high-quality multimedia services.

[0134] Example 2:

[0135] This embodiment provides a multimedia remote AI control method, which is applied to a cloud computing control system. The method includes:

[0136] S61, receiving the features and classification results extracted by the multimedia control box, wherein the classification results are obtained by the multimedia control box extracting features from the sound signals collected by the microphone array and the audio output signals of the multimedia device, and inputting the extracted features into the deep learning model, and the classification results include environmental sounds or the sounds output by the multimedia device itself.

[0137] S62, analyzing the received features and classification results based on the AI ​​model to obtain a first feedback instruction.

[0138] S63: Send the first feedback instruction to the multimedia control box, so that the multimedia control box controls the playback of the multimedia device.

[0139] In some embodiments, the remote AI control method for multimedia further includes:

[0140] S64, receiving current environment data and user preference parameters sent by the multimedia control box, wherein the environment data includes at least one of the following: environment photos, light intensity, and number of people.

[0141] S65, based on the AI ​​model, current environment data and user preference parameters are divided to obtain audio parameters and / or digital humans that match the current environment.

[0142] S66: Send the audio parameters and / or the digital human to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

[0143] In some embodiments, the remote AI control method for multimedia further includes: backing up data of the multimedia control box.

[0144] In this embodiment, the cloud computing control system will regularly back up the settings and user data of the multimedia control box. If the multimedia control box fails or data is lost, the cloud computing control system can quickly restore the latest backup data to the multimedia control box to ensure that user use is not affected.

[0145] In some embodiments, the multimedia remote AI control method further includes: continuous learning and optimization of algorithms. Specifically, based on long-term user usage habits and feedback, the judgment of scene modes and the adjustment of audio parameters are continuously improved to provide more personalized and accurate services.

[0146] The rest of the contents are the same as those in Example 1 and will not be repeated here.

[0147] Example 3:

[0148] This embodiment provides a multimedia remote AI control device, which is applied to a multimedia control box, or can be understood as a multimedia control box. The multimedia control box is generally rectangular in shape and palm-sized. The outer shell is made of high-strength plastic with good insulation and drop resistance. The front of the control box is equipped with a power indicator, a network connection indicator, and operation buttons, allowing users to intuitively understand the device's operating status and perform basic operations. The back of the multimedia control box is equipped with a power port, a network port, and a camera port. The power port is used to connect to an external power adapter to provide a stable power supply to the device. The network port supports wired network connection as a supplement and backup for wireless network. The integrated NVMe storage device has a read and write speed of 800MB / s, which facilitates local storage and playback of audio and video files. In terms of software, a dedicated control program has been developed based on the micro-Linux system to implement the multimedia remote AI control method. Users can remotely control and configure the multimedia control box through a mobile phone app or a webpage. The control program has a simple and intuitive user interface, allowing users to easily select audio sources, adjust playback effects, set scene modes, etc.

[0149] exist Figure 2 In the smart home environment shown, the smart home environment 100 includes multiple intelligent, multi-sensing, network-connected devices. These devices can communicate with the multimedia control box and be integrated together. Smart home devices may include one or more infrared human body sensors 101, omnidirectional microphones 102, multimedia control boxes 103, cameras 104, speakers 105, and TVs 106. The multimedia control box includes a 5G module, an audio module, an HDMI output module, a WIFI module, a Bluetooth module, an RJ45 module, an edge AI chip (i.e., an edge computing module) and multiple communication interfaces. The above communication interfaces connect the smart home devices (i.e., multimedia devices) to the local multimedia control box, and the multimedia control box can communicate with the cloud computing control system via the Internet. Data communication is usually carried out using any of a variety of different types of communication media and protocols, including various wired protocols (such as Ethernet, USB, etc.) or wireless protocols (such as Bluetooth, Wi-Fi, 5G, etc.).

[0150] like Figure 3In the schematic diagram of the integration of the multimedia control box and the cloud computing control system shown, the multimedia control box 201 in another smart home environment 200 can be connected to the Internet 203 by means of a 5G network 202, and the multimedia control box data can be stored in the cloud computing control system 204 and can be retrieved from the control system 204, which includes a cloud-based control system. The control system may include various types of statistics, data analysis, intelligent decision-making, audio processing, image rendering, security protection, resource scheduling, model training and interface adaptation engines 209, which are used for data processing and control of rules related to the smart home environment. Typically, the cloud computing control system 204 is operated by users associated with the smart home. Therefore, the control system 204 can collect and process information collected by the devices connected to the control box 201, and the engine 209 can process the information to generate intelligent multimedia control strategies. The control system 204 includes an AI large model (i.e., AI model) 206, a database 207, an OSS storage 205, and a console 208.

[0151] like Figure 4 As shown, given Figure 3 Another information processing diagram of the cloud computing control system. Various processing engines within 209 can output more personalized multimedia content based on the environment of the intelligent control box 312 (i.e., multimedia control box). This includes obtaining personnel location and information 301, data collected by various connected devices 313, setting optimal volume levels 302, invoking a large AI model for personnel communication 303, and obtaining communication instructions 304. The intelligent control box can also obtain additional external information through internet information 310.

[0152] like Figure 5 The structure of a multimedia control box shown in FIG. Figure 5 Unfilled circles (such as unfilled circle 410) represent sensors and acquisition devices, and arrows (such as arrow 401) represent data input. The multimedia control box 400 can receive data from device 410, analyze and process the data through the internal processor and edge computing module, and access cloud computing resources, AI models, and Internet resources in real time to process and store data. The edge computing module then calculates the appropriate volume and sound quality parameters, and finally outputs the signal to the audio or TV 411 to control playback, creating a comfortable atmosphere for the user.

[0153] The multimedia remote AI control device includes an acquisition module (including a sound acquisition module and a video acquisition module), an edge computing module, and a first communication module (such as a 5G module).

[0154] The acquisition module is used to obtain the sound signal collected by the microphone array and the audio output signal of the multimedia device. For example, Figure 6The sound signal collected by the microphone and the audio output characteristics of the audio equipment are shown.

[0155] The edge computing module, connected to the acquisition module, is used to extract features of the sound signal and audio output signal (feature extraction and analysis), and input the extracted features into the deep learning model (deep learning layer) to obtain classification results, which include environmental sounds or the multimedia device's own output sound. Before feature extraction and analysis, a preprocessing layer is also included to preprocess the acquired sound signal.

[0156] The first communication module is connected to the edge computing module and is used to send the extracted features and classification results to the cloud computing control system, and control the playback of the multimedia device according to the first feedback instruction of the cloud computing control system, wherein the first feedback instruction is obtained by the cloud computing control system based on the AI ​​model analysis of the received features and classification results.

[0157] In some embodiments, the edge technology module is also used to calculate the characteristic differences between the sound signal and the audio output signal; and based on the characteristic differences between the sound signal and the audio output signal, as well as the sound wave propagation characteristics, determine whether the sound signal is environmental sound or the sound output by the multimedia device itself, to verify the classification result.

[0158] In some embodiments, the deep learning model in the edge computing module includes a discriminability function D(S, E).

[0159]

[0160] Where S=(s1,s2,…,s n ) is the feature vector of the multimedia device’s own sound output, E=(e1,e2,…,e m ) is the characteristic vector of the ambient sound, w si and w ej are the characteristics of the sound output by the multimedia device itself i and the characteristics of the ambient sound j The weight of .

[0161] In response to D(S, E) being greater than 0, the classification result is that the multimedia device itself outputs the sound.

[0162] In response to D(S, E) being less than 0, the classification result is environmental sound.

[0163] like Figure 7As shown, in some embodiments, the edge computing module is also used to obtain audio and video signals of the current environment, and use a deep learning model to extract the first key feature of the audio and video signal, and match a first audio and video set based on the first key feature, wherein the first audio and video set refers to an audio and video set associated with the content type, and is also used to send the first audio and video set and the first key feature to the cloud computing control system to receive a second feedback instruction from the cloud computing control system, and to determine the second audio and video set based on the second feedback instruction and the first audio and video set, and control the playback of the multimedia device according to the second audio and video set.

[0164] In some embodiments, the edge computing module is also used to match a third audio and video set based on the first key feature, and evaluate the audio and video in the third audio and video set based on the machine learning model. It is also used to prioritize the audio and video in the third audio and video set according to the evaluation results, and filter out at least one audio and video with a priority greater than a threshold to generate a first audio and video set.

[0165] like Figure 8 As shown, in some embodiments, the edge computing module is also used to obtain audio and video signals of the current environment, and use a deep learning model to extract a second key feature of the audio and video signal, and estimate a first frequency range based on the second key feature and the machine learning model, wherein the first frequency range refers to the frequency range corresponding to the content type, and is used to send the first frequency range and the second key feature to the cloud computing control system to receive a third feedback instruction from the cloud computing control system, and determine a second frequency range based on the third feedback instruction and the first frequency range, for controlling the playback of the multimedia device according to the second frequency range.

[0166] In some embodiments, the edge computing module is also used to send the performance indicators of the deep learning model and the machine learning model during the calculation process to the cloud computing control system to receive the various model parameters issued by the cloud computing control system.

[0167] In some embodiments, the first communication module is used to communicate with the cloud computing control system based on a communication network, wherein the communication network includes a 5G network and a 6G network.

[0168] It should be noted that the use of a multimedia control box can achieve compatibility between different multimedia devices and enable the control of multiple multimedia devices, thereby reducing costs and improving management efficiency.

[0169] Example 4:

[0170] This embodiment provides a multimedia remote AI control device, which is applied to a cloud computing control system and includes a second communication module and a computing module.

[0171] The second communication module is used to receive the features and classification results extracted by the multimedia control box, wherein the classification results are obtained by the multimedia control box extracting features from the sound signals collected by the microphone array and the audio output signals of the multimedia device, and inputting the extracted features into the deep learning model. The classification results include environmental sounds or the sounds output by the multimedia device itself.

[0172] The computing module is connected to the second communication module and is used to analyze the received features and classification results based on the AI ​​model to obtain a first feedback instruction.

[0173] The second communication module is further configured to send the first feedback instruction to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

[0174] The cloud computing intelligent control system achieves remote management and optimization of the multimedia control box through real-time communication with the control box. The cloud computing intelligent control system can analyze the audio environment, such as room size and noise level, based on the image information collected by the camera, and automatically adjust the playback parameters of the multimedia device to achieve the best audio effect. At the same time, the cloud computing intelligent control system can also set different scene modes for the control box according to user needs and preferences, such as shopping malls, stores, homes, offices, etc. In the home use scenario, the specific working method of the cloud computing intelligent control system is as follows:

[0175] When the multimedia control box is started and connected to the network, the control box uploads its hardware information, current environmental data (such as the living room layout, light intensity, number of people, etc. obtained by the camera) and the user's initial settings (such as preferred music type, default volume, etc.) to the cloud computing control system. The cloud computing control system first analyzes and processes these data. Using a pre-trained machine learning model, the audio parameters suitable for the current environment are determined, such as volume, sound effect mode (surround sound, stereo, etc.). At the same time, based on the activities of people in the living room and the time information, the user's possible demand scenarios are inferred. For example, if it is at night and there are fewer people, soothing music and a lower volume may be recommended. The other contents are the same as in Example 2 and will not be repeated here.

[0176] The cloud computing control system in this embodiment connects to the multimedia control box, eliminating the need for the box to store large amounts of audio and video files locally, saving device storage space. Furthermore, the richness and real-time updates of cloud resources provide users with a more diverse audio and video selection, allowing them to access the latest and most popular content at any time. Cloud resources also provide a backup for the control box system, facilitating rapid recovery in the event of a device failure, ensuring the stable operation of the AI ​​multimedia control box. The deep learning and machine learning models in the edge computing module enhance the control box's artificial intelligence capabilities. It can sense its environment and determine whether the sound being played is normal by analyzing connected camera images. For example, it can automatically adjust the volume and sound quality based on ambient noise to provide a higher-quality multimedia experience. Furthermore, the cloud computing control system can set scenes for the multimedia control box, analyzing the audio environment using camera images and dynamically adjusting and optimizing them. The cloud computing intelligent control system also provides digital human services, providing intelligent services for different scenarios.

[0177] Example 5:

[0178] This embodiment provides a multimedia remote AI control device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to implement the multimedia remote AI control method as described in Example 1 or Example 2.

[0179] Example 6:

[0180] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the remote AI control method for multimedia as described in Embodiment 1 or 2 is implemented.

[0181] It will be understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present invention, and the present invention is not limited thereto. Those skilled in the art will appreciate that various modifications and improvements can be made without departing from the spirit and substance of the present invention, and such modifications and improvements are also considered to be within the scope of protection of the present invention.

Claims

1. A multimedia remote AI control method, applied to a multimedia control box, characterized in that: The method comprises: Acquire the sound signal collected by the microphone array and the audio output signal of the multimedia device; extracting features of the sound signal and the audio output signal respectively; The extracted features are input into a deep learning model to obtain classification results, which include environmental sounds or the output sounds of the multimedia device itself; The extracted features and classification results are sent to the cloud computing control system, and the playback of the multimedia device is controlled according to the first feedback instruction of the cloud computing control system, wherein the first feedback instruction is obtained by the cloud computing control system based on the AI ​​model analysis of the received features and classification results.

2. The method according to claim 1, characterized in that After inputting the extracted features into the deep learning model to obtain the classification results, and before sending the extracted features and the classification results to the cloud computing control system, the method further includes: calculating a characteristic difference between the sound signal and the audio output signal; Based on the characteristic difference between the sound signal and the audio output signal, as well as the sound wave propagation characteristics, it is determined that the sound signal is environmental sound or the sound output by the multimedia device itself, so as to verify the classification result.

3. The method according to claim 2, characterized in that The deep learning model includes the discrimination function D(S,E), Where S=(s1,s2,…,s n ) is the feature vector of the multimedia device’s own sound output, E=(e1,e2,…,e m ) is the characteristic vector of the ambient sound, w si and w ej are the characteristics of the sound output by the multimedia device itself i and the characteristics of the ambient sound j The weight of In response to D(S, E) being greater than 0, the classification result is that the multimedia device itself outputs the sound; In response to D(S, E) being less than 0, the classification result is environmental sound.

4. The method according to claim 1, wherein Also includes: Acquire audio and video signals of the current environment, and extract a first key feature of the audio and video signals using a deep learning model; Matching a first audio and video set according to the first key feature, wherein the first audio and video set refers to an audio and video set associated with the content type; Sending the first audio and video set and the first key feature to the cloud computing control system to receive a second feedback instruction from the cloud computing control system; Determining a second audio and video set according to the second feedback instruction and the first audio and video set; The playback of the multimedia device is controlled according to the second audio and video set.

5. The method according to claim 4, characterized in that The matching of the first audio and video set according to the first key feature specifically includes: Matching a third audio and video set according to the first key feature; evaluating the audio and video in the third audio and video set based on the machine learning model; Prioritizing the audio and video in the third audio and video set according to the evaluation result; At least one audio or video with a priority greater than a threshold is screened out to generate a first audio or video set.

6. The method according to claim 1, characterized in that Also includes: Acquire audio and video signals of the current environment, and extract a second key feature of the audio and video signals using a deep learning model; estimating a first frequency range based on the second key feature and the machine learning model, wherein the first frequency range refers to a frequency range corresponding to the content type; sending the first frequency range and the second key feature to the cloud computing control system to receive a third feedback instruction from the cloud computing control system; determining a second frequency range according to the third feedback instruction and the first frequency range; Playback of the multimedia device is controlled according to the second frequency range.

7. The method according to any one of claims 1 to 6, characterized in that Also includes: The performance indicators of the deep learning model and the machine learning model during the calculation process are sent to the cloud computing control system to receive the model parameters issued by the cloud computing control system. Communicate with the cloud computing control system based on a communication network, wherein the communication network includes a 5G network and a 6G network.

8. A multimedia remote AI control method, applied to cloud computing control systems, characterized in that: The method comprises: receiving features extracted by the multimedia control box and classification results, wherein the classification results are obtained by the multimedia control box performing feature extraction on the sound signals collected by the microphone array and the audio output signals of the multimedia device and inputting the extracted features into the deep learning model, and the classification results include ambient sound or the sound output by the multimedia device itself; Analyzing the received features and classification results based on the AI ​​model to obtain a first feedback instruction; The first feedback instruction is sent to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

9. The method according to claim 8, characterized in that Also includes: Receiving current environment data and user preference parameters sent by the multimedia control box, wherein the environment data includes at least one of the following: an environment photo, light intensity, and the number of people; Analyzing the current environment data and the user preference parameters based on the AI ​​model to obtain audio parameters and / or a digital human that matches the current environment; The audio parameters and / or the digital human are sent to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

10. A multimedia remote AI control device, characterized in that: Including acquisition module, edge computing module, first communication module, The acquisition module is used to obtain the sound signal collected by the microphone array and the audio output signal of the multimedia device. The edge computing module is connected to the acquisition module and is used to extract the features of the sound signal and the audio output signal respectively, and input the extracted features into the deep learning model to obtain classification results, which include environmental sounds or the output sounds of the multimedia device itself. The first communication module is connected to the edge computing module and is used to send the extracted features and classification results to the cloud computing control system, and control the playback of the multimedia device according to the first feedback instruction of the cloud computing control system, wherein the first feedback instruction is obtained by the cloud computing control system based on the AI ​​model analysis of the received features and classification results.

11. A multimedia remote AI control device, characterized in that: including a second communication module and a computing module, The second communication module is used to receive the features and classification results extracted by the multimedia control box, wherein the classification results are obtained by the multimedia control box extracting features from the sound signals collected by the microphone array and the audio output signals of the multimedia device, and inputting the extracted features into the deep learning model. The classification results include environmental sounds or the audio output of the multimedia device itself. The computing module is connected to the second communication module and is used to analyze the received features and classification results based on the AI ​​model to obtain a first feedback instruction. The second communication module is further configured to send the first feedback instruction to the multimedia control box so that the multimedia control box controls the playback of the multimedia device.

12. A multimedia remote AI control device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement the remote AI control method for multimedia according to any one of claims 1 to 7, or 8 to 9.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the remote AI control method for multimedia according to any one of claims 1 to 7, or 8 to 9 is implemented.