Intelligent terminal adaptive audio and video optimization method and device based on terahertz sensing and multi-mode fusion, equipment and storage medium

By using terahertz sensing and multimodal fusion technology, a dynamic environment model is generated, which solves the problems of single-dimensional environment perception, insufficient accuracy, and lack of synergy in audio and video optimization of smart terminals, and realizes high-precision, multi-dimensional audio and video collaborative optimization.

CN121547552APending Publication Date: 2026-02-17SHENZHEN JIUZHOU ELECTRIC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511489683.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing smart terminals suffer from limited environmental perception dimensions and insufficient precision, and their audio and video optimization lacks synergy, resulting in a fragmented user experience.

Method used

Employing terahertz sensing and multimodal fusion technology, the system acquires user space data through terahertz sensors, combines it with multimodal environmental data collected from cameras, microphone arrays, and ambient light sensors, generates a dynamic environmental model, and optimizes audio and video output based on this model.

Benefits of technology

It achieves high-precision, multi-dimensional, and non-contact environmental perception, accurately identifying user distance, location distribution, and micro-movement behavior, and realizing dynamic collaborative control of functions such as precise audio beam pointing, adaptive volume adjustment, intelligent noise reduction, and video local backlighting and resolution optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547552A_ABST
    Figure CN121547552A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent terminal adaptive audio and video optimization method and device based on terahertz sensing and multi-mode fusion, equipment and a storage medium, and relates to the technical field of intelligent terminals, and the method comprises the steps: obtaining the spatial data, collected by a terahertz sensor, of a user relative to terminal equipment; acquiring multi-modal environment data acquired by a camera, a microphone array and an environment light sensor; fusing the spatial data and the multi-modal environment data to generate a dynamic environment model; based on a dynamic environment model, audio output and video output are optimized, high-precision, multi-dimensional and non-contact environment perception capability is realized, user distance, position distribution and micro-motion behaviors can be accurately identified, and dynamic cooperative control of audio and video output is realized under the guidance of a unified environment model. The technical problems that an existing intelligent terminal depends on a traditional sensor, so that the environment perception dimension is single, precision is insufficient, and audio and video optimization lacks collaboration are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent terminals, and in particular to an intelligent terminal adaptive audio and video optimization method and device based on terahertz sensing and multi-modal fusion, an intelligent terminal adaptive audio and video optimization equipment and a storage medium. BACKGROUND

[0002] The current environmental adaptation function of intelligent terminals mainly relies on traditional sensors, such as ambient light sensors, infrared sensors or cameras. For example, a television can adjust the screen brightness according to the ambient light intensity, and a video conference system can detect the face position through a camera and adjust the picture composition or the direction of a beamforming microphone.

[0003] At present, traditional optical sensors are susceptible to environmental interference. The ambient light sensor can only sense the overall light intensity and cannot distinguish the light distribution of different areas, nor can it identify the position, distance or number of the audience. The performance of the camera decreases under low light conditions, and there is a risk of privacy leakage. The existing system lacks non-contact accurate perception ability of user distance. For example, a television cannot dynamically adjust the display resolution or local backlight according to the viewing distance, and a sound box cannot real-time optimize the sound field directivity according to the audience distribution. The optimization of audio and video is usually carried out independently, and there is a lack of collaborative control mechanism based on a unified environmental model, resulting in a fragmented user experience.

[0004] Therefore, how to comprehensively improve the dimension and accuracy of environmental perception, and on this basis to realize the collaborative optimization of audio and video output, is a problem to be solved at present.

[0005] The above content is only used to assist in understanding the technical solutions of the present application, and does not represent the acknowledgement of the above content as prior art. SUMMARY

[0006] The main purpose of the present application is to provide an intelligent terminal adaptive audio and video optimization method and device based on terahertz sensing and multi-modal fusion, and an intelligent terminal adaptive audio and video optimization equipment and a storage medium, aiming at solving the technical problem of how to comprehensively improve the dimension and accuracy of environmental perception, and on this basis to realize the collaborative optimization of audio and video output.

[0007] To achieve the above purpose, the present application provides an intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multi-modal fusion, which comprises: The first independent right technical solution.

[0008] In an embodiment, the antenna array and the waveform modulation circuit of the terahertz sensor are integrated in the frame or the back of the terminal equipment, and the spatial data includes distance data and position distribution data; The spatial data collected by the terahertz sensor relative to the terminal equipment is obtained, comprising: The waveform modulation circuit is controlled to generate a frequency-modulated continuous wave; Control the antenna array to transmit the frequency-modulated continuous wave and receive the echo signal of the frequency-modulated continuous wave; Based on the echo signal, the flight time and Doppler frequency shift of the echo signal are obtained; Based on the flight time and the Doppler frequency shift, the distance data between the user and the terminal device is obtained; The antenna array is controlled to perform beam scanning, and the energy distribution of the echo signal in different scanning directions is analyzed to obtain the user's location distribution data.

[0009] In one embodiment, the multimodal environmental data includes texture data, contour data, background noise data, sound field structure data, spectral data, and brightness data; The acquisition of multimodal environmental data collected by the camera, microphone array, and ambient light sensor includes: Based on the image data captured by the camera, texture data and contour data are obtained; Based on the sound field data collected by the microphone array, background noise data and sound field structure data are obtained. Based on the optical data collected by the ambient light sensor, spectral data and brightness data are obtained.

[0010] In one embodiment, fusing the spatial data with the multimodal environment data to generate a dynamic environment model includes: The spatial data and the multimodal environment data are time-stamped and aligned to a coordinate system to obtain spatiotemporally aligned data; Feature-level fusion is performed on the spatiotemporal aligned data to generate a dynamic environment model.

[0011] In one embodiment, optimizing the audio output based on the dynamic environment model includes: Based on the location distribution data of the dynamic environment model, the microphone array is controlled to perform beamforming. Based on the distance data from the dynamic environment model, the microphone array is controlled to adjust the output volume and sound field width. Based on the background noise data from the dynamic environment model, a selective noise reduction algorithm is used to suppress the background noise data, thereby optimizing the audio output.

[0012] In one embodiment, optimizing the video output based on the dynamic environment model includes: Based on the spectral and brightness data of the dynamic environment model, ambient lighting is dimmed in different zones. Based on the distance and location distribution data of the dynamic environment model, the pixel rendering strategy is adjusted; Based on the distance data from the dynamic environment model, the image sharpness, super-resolution, and sharpening algorithms are adjusted to optimize the video output.

[0013] In one embodiment, after controlling the antenna array to transmit the frequency-modulated continuous wave and receiving the echo signal of the frequency-modulated continuous wave, the method further includes: Based on the echo signal, extract the phase change features or frequency change features caused by the user's micro-motion; The phase change feature or the frequency change feature is matched with a preset gesture pattern to identify the user's gesture; Based on the user's gestures, contactless operation of the terminal device can be achieved.

[0014] Furthermore, to achieve the above objectives, this application also proposes an adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion, wherein the adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion includes: The first acquisition module is used to acquire spatial data of the user relative to the terminal device collected by the terahertz sensor. The second acquisition module is used to acquire multimodal environmental data collected by the camera, microphone array and ambient light sensor; The modeling module is used to fuse the spatial data with the multimodal environment data to generate a dynamic environment model; The optimization module is used to optimize the audio and video outputs based on the dynamic environment model.

[0015] Furthermore, to achieve the above objectives, this application also proposes an adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion. The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion as described above.

[0016] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multimodal fusion as described above.

[0017] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multimodal fusion as described above.

[0018] This application acquires spatial data of the user relative to the terminal device collected by a terahertz sensor; acquires multimodal environmental data collected by a camera, microphone array, and ambient light sensor; fuses the spatial data with the multimodal environmental data to generate a dynamic environmental model; and optimizes audio and video output based on the dynamic environmental model, achieving high-precision, multi-dimensional, and non-contact environmental perception capabilities. It can accurately identify user distance, location distribution, and micro-movement behavior, and under the guidance of a unified environmental model, achieves dynamic collaborative control of functions such as precise audio beam pointing, adaptive volume adjustment, intelligent noise reduction, and video local backlighting, resolution optimization, and color temperature adaptation. This effectively solves the technical problems of existing smart terminals relying on traditional sensors, such as single-dimensional environmental perception, insufficient accuracy, and lack of synergy in audio and video optimization. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating an embodiment of the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion provided in this application. Figure 2 This is a simplified flowchart of an embodiment of the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion provided in this application. Figure 3 This is a schematic diagram of the module structure of the intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion according to an embodiment of this application; Figure 4 This is a schematic diagram of the device structure of the hardware operating environment involved in the intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multimodal fusion in the embodiments of this application.

[0022] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0023] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0024] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0025] The main solution of this application embodiment is: to acquire spatial data of the user relative to the terminal device collected by the terahertz sensor; to acquire multimodal environmental data collected by the camera, microphone array and ambient light sensor; to fuse the spatial data and multimodal environmental data to generate a dynamic environmental model; and to optimize the audio output and video output based on the dynamic environmental model.

[0026] Traditional optical sensors are susceptible to environmental interference. Ambient light sensors can only sense overall light intensity and cannot distinguish the light distribution in different areas, let alone identify the location, distance, or number of viewers. Cameras experience performance degradation in low-light conditions and pose a risk of privacy breaches. Existing systems lack the ability to accurately sense user distance without contact. For example, televisions cannot dynamically adjust display resolution or local backlighting based on viewing distance, and speakers cannot optimize sound field directivity in real time based on audience distribution. Audio and video optimization are often performed independently, lacking a collaborative control mechanism based on a unified environmental model, resulting in a fragmented user experience.

[0027] This application provides a solution that achieves high-precision, multi-dimensional, and non-contact environmental perception capabilities. It can accurately identify user distance, location distribution, and micro-movement behavior. Under the guidance of a unified environmental model, it realizes dynamic collaborative control of functions such as precise audio beam pointing, adaptive volume adjustment, intelligent noise reduction, video local backlighting, resolution optimization, and color temperature adaptation. It effectively solves the technical problems of existing smart terminals relying on traditional sensors, which result in single-dimensional environmental perception, insufficient accuracy, and lack of synergy in audio and video optimization.

[0028] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions. The following description uses a smart terminal as an example to illustrate this embodiment and the subsequent embodiments.

[0029] Based on this, embodiments of this application provide an adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion according to this application.

[0030] In this embodiment, the intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multimodal fusion includes steps S10 to S40: Step S10: Obtain spatial data of the user relative to the terminal device collected by the terahertz sensor; It should be noted that a terahertz sensor is a device that uses electromagnetic waves in the terahertz band for sensing; terminal devices refer to intelligent devices such as televisions, smart speakers, and video conferencing terminals that require optimized audio and video output; and spatial data refers to information used to describe the spatial relationship between the user and the terminal device.

[0031] Understandably, traditional optical sensors suffer from performance degradation and privacy risks in low-light environments, and cannot accurately sense user distance and location. Therefore, step S10 avoids the limitations of traditional sensors, thereby providing a reliable data foundation for improving the accuracy and coordination of audio and video optimization.

[0032] In one feasible implementation, step S10 may include: controlling the waveform modulation circuit to generate a frequency-modulated continuous wave; controlling the antenna array to transmit the frequency-modulated continuous wave and receiving the echo signal of the frequency-modulated continuous wave; obtaining the flight time and Doppler frequency shift of the echo signal based on the echo signal; obtaining distance data between the user and the terminal device based on the flight time and the Doppler frequency shift; controlling the antenna array to perform beam scanning and analyzing the energy distribution of the echo signal in different scanning directions to obtain the user's location distribution data.

[0033] It should be noted that the terahertz sensor's antenna array and waveform modulation circuitry are integrated into the bezel or back of the terminal device. Spatial data includes distance data and location distribution data. The waveform modulation circuitry is an electronic circuit used to generate terahertz signals with specific frequency variation patterns. Frequency-modulated continuous wave (FM-CW) is a continuous electromagnetic wave whose frequency varies linearly with time. The antenna array is a component composed of multiple antenna elements used for directional transmission and reception of electromagnetic waves. The echo signal is the signal received after the terahertz wave is transmitted and reflected back from the user. Time of flight is the time it takes for the electromagnetic wave to travel from transmission to reception. Doppler shift is the frequency change of the echo signal caused by the relative motion between the user and the terminal device. The terminal device is the physical hardware carrier that performs this method, such as a smart TV. Distance data is a numerical value characterizing the straight-line length between the user and the terminal device. Beam scanning is the process of controlling the phase of the antenna array to move the transmitted beam according to a specific pattern within a spatial range. The energy distribution of the echo signal refers to the intensity information of the received echo signal at different scanning angles. Location distribution data is information used to describe the orientation and distribution of the user in the space in front of the terminal device.

[0034] For example, the terahertz sensor on top of the television can accurately calculate the presence of three viewers in front of it and their respective distances and positions through the above steps: viewer A is 3.5 meters away in the center, viewer B is 3.7 meters away in the center slightly to the right, and viewer C is 2.8 meters away in the left.

[0035] In this embodiment, by using terahertz waves for non-contact active detection, the shortcomings of traditional sensors in accurately sensing user spatial information are solved, laying the foundation for subsequent audio and video collaborative optimization based on accurate spatial data.

[0036] The above are merely feasible implementations of step S10 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S10.

[0037] Step S20: Acquire multimodal environmental data collected by the camera, microphone array, and ambient light sensor; It should be noted that multimodal environmental data refers to a collection of data describing various aspects of the environment, acquired from different types of sensors. This includes texture and contour data provided by cameras, background noise and sound field structure data provided by microphone arrays, and spectral and brightness data provided by ambient light sensors. Texture data refers to information representing the detailed features of an object's surface in an image; contour data refers to information representing the edges of an object's shape in an image; background noise data refers to the characteristic information of interference sounds from non-target sound sources present in the environment; sound field structure data refers to the energy distribution and propagation characteristics of sound in space; spectral data refers to the intensity distribution of light at different wavelengths in the environment; and brightness data refers to the overall or local brightness of light in the environment.

[0038] Understandably, since a single type of data cannot fully and accurately describe the complex real-world usage environment, step S20 can avoid optimization decision bias caused by missing or one-sided environmental information, thereby providing a rich and complementary multi-dimensional information foundation for building a high-precision dynamic environment model.

[0039] In one feasible implementation, step S20 may include: obtaining texture data and contour data based on image data acquired by a camera; obtaining background noise data and sound field structure data based on sound field data acquired by a microphone array; and obtaining spectral data and brightness data based on optical data acquired by an ambient light sensor.

[0040] It should be noted that an ambient light sensor is a photoelectric sensor used to detect the intensity and color temperature characteristics of ambient light; a microphone array is a system composed of multiple microphones arranged in a certain geometric structure, used to collect spatial sound field information.

[0041] For example, the camera can confirm the approximate location of the audience and provide facial contour information; the microphone array can monitor the low-frequency noise spectrum (background noise data) from the air conditioner and determine the direction of arrival of the main sound waves (sound field structure data); the ambient light sensor can detect that the room is generally dark, but there is a local bright light area caused by the desk lamp on the right side of the screen (brightness data) and its warm yellow temperature characteristics (spectral data).

[0042] In this embodiment, by simultaneously collecting multi-dimensional environmental data such as visual, acoustic, and optical data, the problem of the single dimension of environmental perception in traditional systems is solved, providing comprehensive data input for subsequent multimodal fusion and collaborative optimization.

[0043] In one feasible implementation, after step S20, the method may further include: extracting phase change features or frequency change features caused by user micro-movements based on the echo signal; matching the phase change features or frequency change features with a preset gesture pattern to identify the user gesture; and realizing contactless operation of the terminal device based on the user gesture.

[0044] It should be noted that user micro-movements refer to small, non-rigid movements made by the user's hand or body parts, such as waving or sliding; the phase change characteristics refer to the specific pattern of echo signal phase shift caused by the change in distance between the user and the sensor due to the micro-movement; the frequency change characteristics refer to the specific pattern of echo signal Doppler frequency shift caused by the relative velocity generated by the user's micro-movement; the preset gesture pattern is a standard phase or frequency change template corresponding to a specific control command (such as volume adjustment or menu switching) that is pre-defined and stored in the system through machine learning or rule definition; contactless operation means that the user can control the device function by gestures without physically touching the terminal device or its remote control.

[0045] For example, when a user waves their hand to the right in the air, the terahertz sensor captures a series of specific phase-continuous change features. The system matches these features with the "wave right" pattern in a preset gesture library. If the match is successful, the system sends a "next channel" control command to the smart TV to complete the channel switching.

[0046] In this embodiment, gesture recognition is achieved by utilizing the high-precision sensing capability of terahertz waves for micro-motions, which solves the problems of traditional visual gesture recognition failure in low-light environments and privacy concerns, thereby realizing a more natural and reliable contactless human-computer interaction.

[0047] The above are merely feasible implementations of step S20 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S20.

[0048] Step S30: The spatial data and the multimodal environment data are fused to generate a dynamic environment model; It should be noted that the dynamic environment model is a digital representation generated through data fusion technology that can reflect the comprehensive state of the environment in which the terminal device is located in real time. It includes at least the spatial distribution of users, the zoning of ambient lighting, and the distribution characteristics of noise sources.

[0049] Understandably, since the raw data from different sensors are independent and heterogeneous in time and space, direct use can lead to model errors and inaccurate decisions. Therefore, step S30 can avoid environmental perception bias caused by data asynchrony and coordinate inconsistency, thereby improving the accuracy and synergy of subsequent audio and video optimization strategies.

[0050] In one feasible implementation, step S30 may include: synchronizing the spatial data and the multimodal environment data with timestamps and coordinate system unification to obtain spatiotemporally aligned data; and performing feature-level fusion on the spatiotemporally aligned data to generate a dynamic environment model.

[0051] It should be noted that spatiotemporally aligned data refers to a multi-source data set that has been processed so that all data points correspond consistently on the timeline and have a unified reference benchmark in the spatial coordinate system.

[0052] For example, the system integrates the precise audience coordinates obtained by the terahertz sensor, the audience outline confirmed by the camera, the brightness of the strong light area on the right side of the screen detected by the ambient light sensor, and the direction of the air conditioning noise source identified by the microphone array into a three-dimensional environment map to form a dynamic environment model that is updated in real time.

[0053] In this embodiment, by aligning and fusing multi-source heterogeneous data in a time and space and using features, the problem of isolated perceptual information and the inability to form a unified environmental cognition is solved, thereby providing a precise and unified decision-making basis for achieving integrated audio and video collaborative optimization.

[0054] The above are merely feasible implementations of step S30 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S30.

[0055] Step S40: Optimize the audio and video outputs based on the dynamic environment model.

[0056] It should be noted that audio output refers to the sound signal generated by the speaker or audio system of the terminal device; video output refers to the image signal displayed on the screen of the terminal device.

[0057] Understandably, since audio and video optimization strategies are difficult to provide a consistent immersive experience if they are independent of each other, step S40 can avoid the disconnect between audio and video optimization, thereby achieving coordinated control based on a unified dynamic environment model and significantly improving the overall user experience.

[0058] In one feasible implementation, step S40 may include: controlling the microphone array to perform beamforming based on the location distribution data of the dynamic environment model; controlling the microphone array to adjust the output volume and sound field width based on the distance data of the dynamic environment model; and using a selective noise reduction algorithm to suppress the background noise data based on the background noise data of the dynamic environment model to optimize the audio output.

[0059] It should be noted that user location distribution data in the dynamic environment model refers to information describing the user's spatial orientation; beamforming is a technique that focuses sound energy in a specific direction by controlling the phase and amplitude of each unit in a speaker array; user distance data in the dynamic environment model refers to the straight-line distance between the user and the terminal device; output volume refers to the loudness level of the sound; sound field width refers to the perceived range of stereo or surround sound effects; selective noise reduction algorithm is a signal processing technique that suppresses specific noise frequencies without affecting the target audio frequency band.

[0060] For example, the system drives the TV speaker array to generate three independent audio beams, each precisely pointed to the area where three viewers are located. Since viewer C is closer, the system slightly reduces the volume gain of the beam directed at him, ensuring that the perceived loudness is consistent for all three viewers. Simultaneously, the system initiates selective noise reduction for the identified low-frequency noise spectrum from the air conditioner.

[0061] In this embodiment, by performing synergistic optimization of audio directionality, distance measurement, and noise reduction based on a unified environment model, the problem of traditional audio systems being unable to provide users in different locations with a personalized and consistent listening experience and having weak anti-interference capabilities is solved.

[0062] In another feasible implementation, step S40 may include: performing zoned dimming of ambient lighting based on the spectral data and brightness data of the dynamic environment model; adjusting the pixel rendering strategy based on the distance data and position distribution data of the dynamic environment model; and adjusting the image sharpness, super-resolution, and sharpening algorithms based on the distance data of the dynamic environment model to optimize the video output.

[0063] It should be noted that the environmental spectral data in the dynamic environment model refers to the intensity distribution information of ambient light at different wavelengths; the brightness data in the dynamic environment model refers to the brightness information of ambient light; local dimming refers to the technique of dividing the backlight of the display screen into multiple independent control areas and adjusting the brightness of each area as needed; pixel rendering strategy refers to the algorithm rules that determine how to drive pixels to display images, which may include the adjustment of parameters such as contrast, color temperature, and sharpness.

[0064] For example, when the system detects strong light from a desk lamp on the right side of the screen, it dynamically increases the brightness of the local backlight module on the right side of the screen. At the same time, it intelligently optimizes the image super-resolution and sharpening algorithms based on the average viewing distance of the three viewers, and performs additional sharpness compensation on the left edge of the image where viewer C, who is viewing off-axis, is located.

[0065] In this embodiment, by performing zoned dimming, rendering, and adaptive sharpness adjustment on the video based on ambient light sensing and user space information, the problem that traditional video optimization cannot cope with local glare and differences in visual experience from different viewing positions is solved.

[0066] The above are merely feasible implementations of step S40 provided in this embodiment. This embodiment does not specifically limit the specific implementation of step S40.

[0067] This embodiment provides an adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion. It acquires spatial data of the user relative to the terminal device from a terahertz sensor; acquires multimodal environmental data from a camera, microphone array, and ambient light sensor; fuses the spatial data with the multimodal environmental data to generate a dynamic environment model; and optimizes audio and video outputs based on the dynamic environment model. This achieves high-precision, multi-dimensional, and non-contact environmental perception capabilities, accurately identifying user distance, location distribution, and micro-movements. Under the guidance of a unified environment model, it achieves dynamic collaborative control of functions such as precise audio beam pointing, adaptive volume adjustment, intelligent noise reduction, and video local backlighting, resolution optimization, and color temperature adaptation. This effectively solves the technical problems of existing smart terminals relying on traditional sensors, resulting in single-dimensional environmental perception, insufficient accuracy, and a lack of synergy in audio and video optimization.

[0068] For example, to help understand the implementation process of the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion in this embodiment, please refer to... Figure 2 , Figure 2 A simplified flowchart of an adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion is provided, specifically: The sensor layer includes a terahertz sensor, a camera, a microphone array, and an ambient light sensor, with the terahertz sensor serving as the core sensor. Its waveform modulation circuit generates a frequency-modulated continuous wave, which is transmitted through an antenna array, which then receives the echo signal. By analyzing the time-of-flight and Doppler shift of the echo signal, the user's distance data is calculated. Through beam scanning and analysis of the energy distribution of the echo signal in different directions, the user's position distribution data is obtained. The core outputs of this step are distance, position, and gesture data. Synchronized with the terahertz sensor, the camera acquires visual data (for texture and contour extraction), the microphone array acquires sound field data (for background noise and sound field structure analysis), and the ambient light sensor acquires light intensity / color temperature data.

[0069] Terahertz spatial data is input into the processing and fusion layer along with multimodal environmental data from other sensors. First, data preprocessing and alignment (corresponding to timestamp synchronization and coordinate system unification) are performed to obtain spatiotemporally aligned multi-source data. Then, feature extraction and fusion (corresponding to feature-level fusion) are performed, such as complementary integration of precise terahertz coordinates, visual contours, ambient light intensity zones, and acoustic noise source directions. Based on the fused features, environmental model calculations are performed, ultimately generating a unified, digital, dynamic environmental model. This model includes elements such as audience, space, noise, and lighting (e.g., ...). Figure 2 (As shown).

[0070] The generated dynamic environment model enters the execution and optimization layer, where it is fed into the audio optimization engine and video optimization engine, respectively. The audio optimization engine, based on model information (such as viewer position, distance, and noise), executes strategies such as beamforming, adaptive volume, and intelligent noise reduction, ultimately outputting optimized audio through speakers / audio output devices. The video optimization engine, based on model information (such as ambient lighting, user distance, and position), executes strategies such as local dimming, pixel rendering, and sharpness optimization, ultimately displaying optimized video through a display screen / backlight output device.

[0071] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion. Any simple modifications based on this technical concept are within the protection scope of this application.

[0072] This application also provides an adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion. Please refer to [reference needed]. Figure 3 The intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion includes: The first acquisition module 10 is used to acquire spatial data of the user relative to the terminal device collected by the terahertz sensor. The second acquisition module 20 is used to acquire multimodal environmental data collected by the camera, microphone array and ambient light sensor; Modeling module 30 is used to fuse the spatial data with the multimodal environment data to generate a dynamic environment model; The optimization module 40 is used to optimize the audio output and video output based on the dynamic environment model.

[0073] The adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion provided in this application adopts the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion in the above embodiments, and can solve the technical problem of adaptive audio and video optimization for intelligent terminals based on terahertz sensing and multimodal fusion. Compared with the prior art, the beneficial effects of the adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion provided in this application are the same as the beneficial effects of the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion provided in the above embodiments, and other technical features in the adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0074] The first acquisition module 10 is further configured to control the waveform modulation circuit to generate a frequency-modulated continuous wave; control the antenna array to transmit the frequency-modulated continuous wave and receive the echo signal of the frequency-modulated continuous wave; obtain the flight time and Doppler frequency shift of the echo signal based on the echo signal; obtain the distance data between the user and the terminal device based on the flight time and the Doppler frequency shift; control the antenna array to perform beam scanning and analyze the energy distribution of the echo signal in different scanning directions to obtain the user's location distribution data.

[0075] The second acquisition module 20 is also used to obtain texture data and contour data based on image data acquired by the camera; to obtain background noise data and sound field structure data based on sound field data acquired by the microphone array; and to obtain spectral data and brightness data based on optical data acquired by the ambient light sensor.

[0076] The modeling module 30 is also used to synchronize the spatial data and the multimodal environment data with timestamps and coordinate system unification to obtain spatiotemporally aligned data; and to perform feature-level fusion on the spatiotemporally aligned data to generate a dynamic environment model.

[0077] The optimization module 40 is further configured to control the microphone array to perform beamforming based on the position distribution data of the dynamic environment model; control the microphone array to adjust the output volume and sound field width based on the distance data of the dynamic environment model; and use a selective noise reduction algorithm to suppress the background noise data based on the background noise data of the dynamic environment model to complete the optimization of audio output.

[0078] The optimization module 40 is further configured to perform zoned dimming of ambient lighting based on the spectral and brightness data of the dynamic environment model; adjust the pixel rendering strategy based on the distance and position distribution data of the dynamic environment model; and adjust the image sharpness, super-resolution, and sharpening algorithms based on the distance data of the dynamic environment model to optimize the video output.

[0079] The second acquisition module 20 is further configured to extract phase change features or frequency change features caused by user micro-movements based on the echo signal; match the phase change features or frequency change features with a preset gesture pattern to identify the user gesture; and realize contactless operation of the terminal device based on the user gesture.

[0080] This application provides an adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion. The adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion in the above embodiment 1.

[0081] The following is for reference. Figure 4 This document illustrates a structural schematic diagram of an intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion, suitable for implementing embodiments of this application. The intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion in the embodiments of this application can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4The intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0082] like Figure 4 As shown, the terahertz sensing and multimodal fusion-based intelligent terminal adaptive audio and video optimization device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into random access memory (RRAM) 1004. RAM 1004 also stores various programs and data required for the operation of the terahertz sensing and multimodal fusion-based intelligent terminal adaptive audio and video optimization device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the terahertz sensing and multimodal fusion-based intelligent terminal adaptive audio-visual optimization device to wirelessly or wiredly communicate with other devices to exchange data. Although the figure shows an terahertz sensing and multimodal fusion-based intelligent terminal adaptive audio-visual optimization device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0083] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0084] The adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion provided in this application adopts the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion in the above embodiments, and can solve the technical problem of adaptive audio and video optimization for intelligent terminals based on terahertz sensing and multimodal fusion. Compared with the prior art, the beneficial effects of the adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion provided in this application are the same as the beneficial effects of the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion provided in the above embodiments, and other technical features in the adaptive audio and video optimization device for intelligent terminals based on terahertz sensing and multimodal fusion are the same as the features disclosed in the previous embodiment method, and will not be repeated here.

[0085] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0086] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0087] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multimodal fusion in the above embodiments.

[0088] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0089] The aforementioned computer-readable storage medium may be included in an intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion; or it may exist independently and not assembled into the intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion.

[0090] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion, the intelligent terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion performs the following: acquiring spatial data of the user relative to the terminal device collected by the terahertz sensor; acquiring multimodal environmental data collected by the camera, microphone array, and ambient light sensor; fusing the spatial data with the multimodal environmental data to generate a dynamic environmental model; and optimizing the audio and video outputs based on the dynamic environmental model.

[0091] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0093] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0094] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion. This method can solve the technical problem of adaptive audio and video optimization for intelligent terminals based on terahertz sensing and multimodal fusion. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion provided in the above embodiments, and will not be repeated here.

[0095] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described adaptive audio and video optimization method for intelligent terminals based on terahertz sensing and multimodal fusion.

[0096] The computer program product provided in this application can solve the technical problem of adaptive audio and video optimization for smart terminals based on terahertz sensing and multimodal fusion. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion provided in the above embodiments, and will not be repeated here.

[0097] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for adaptive audio and video optimization of intelligent terminals based on terahertz sensing and multimodal fusion, characterized in that, The method includes: Acquire spatial data of the user relative to the terminal device from a terahertz sensor; Acquire multimodal environmental data collected by cameras, microphone arrays, and ambient light sensors; The spatial data is fused with the multimodal environmental data to generate a dynamic environment model; Based on the dynamic environment model, the audio and video outputs are optimized.

2. The method as described in claim 1, characterized in that, The antenna array and waveform modulation circuit of the terahertz sensor are integrated into the frame or back of the terminal device, and the spatial data includes distance data and location distribution data. The acquisition of spatial data of the user relative to the terminal device collected by the terahertz sensor includes: The waveform modulation circuit is controlled to generate a frequency-modulated continuous wave; Control the antenna array to transmit the frequency-modulated continuous wave and receive the echo signal of the frequency-modulated continuous wave; Based on the echo signal, the flight time and Doppler frequency shift of the echo signal are obtained; Based on the flight time and the Doppler frequency shift, the distance data between the user and the terminal device is obtained; The antenna array is controlled to perform beam scanning, and the energy distribution of the echo signal in different scanning directions is analyzed to obtain the user's location distribution data.

3. The method as described in claim 1, characterized in that, The multimodal environmental data includes texture data, contour data, background noise data, sound field structure data, spectral data, and brightness data; The acquisition of multimodal environmental data collected by the camera, microphone array, and ambient light sensor includes: Based on the image data captured by the camera, texture data and contour data are obtained; Based on the sound field data collected by the microphone array, background noise data and sound field structure data are obtained. Based on the optical data collected by the ambient light sensor, spectral data and brightness data are obtained.

4. The method as described in claim 1, characterized in that, The step of fusing the spatial data with the multimodal environment data to generate a dynamic environment model includes: The spatial data and the multimodal environment data are time-stamped and aligned to a coordinate system to obtain spatiotemporally aligned data; Feature-level fusion is performed on the spatiotemporal aligned data to generate a dynamic environment model.

5. The method as described in claim 1, characterized in that, Based on the aforementioned dynamic environment model, the audio output is optimized, including: Based on the location distribution data of the dynamic environment model, the microphone array is controlled to perform beamforming. Based on the distance data from the dynamic environment model, the microphone array is controlled to adjust the output volume and sound field width. Based on the background noise data from the dynamic environment model, a selective noise reduction algorithm is used to suppress the background noise data, thereby optimizing the audio output.

6. The method as described in claim 1, characterized in that, Based on the aforementioned dynamic environment model, the video output is optimized, including: Based on the spectral and brightness data of the dynamic environment model, ambient lighting is dimmed in different zones. Based on the distance and location distribution data of the dynamic environment model, the pixel rendering strategy is adjusted; Based on the distance data from the dynamic environment model, the image sharpness, super-resolution, and sharpening algorithms are adjusted to optimize the video output.

7. The method as described in claim 2, characterized in that, After controlling the antenna array to transmit the frequency-modulated continuous wave and receiving the echo signal of the frequency-modulated continuous wave, the method further includes: Based on the echo signal, extract the phase change features or frequency change features caused by the user's micro-motion; The phase change feature or the frequency change feature is matched with a preset gesture pattern to identify the user's gesture; Based on the user's gestures, contactless operation of the terminal device can be achieved.

8. A smart terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion, characterized in that, The device includes: The first acquisition module is used to acquire spatial data of the user relative to the terminal device collected by the terahertz sensor. The second acquisition module is used to acquire multimodal environmental data collected by the camera, microphone array and ambient light sensor; The modeling module is used to fuse the spatial data with the multimodal environment data to generate a dynamic environment model; The optimization module is used to optimize the audio and video outputs based on the dynamic environment model.

9. A smart terminal adaptive audio and video optimization device based on terahertz sensing and multimodal fusion, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the adaptive audio and video optimization method for smart terminals based on terahertz sensing and multimodal fusion as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the intelligent terminal adaptive audio and video optimization method based on terahertz sensing and multimodal fusion as described in any one of claims 1 to 7.