Multi-scene switching audio loudspeaker tone quality adaptive adjustment system
By integrating multi-dimensional data and adopting a fully closed-loop architecture, the audio speaker sound quality adaptive adjustment system with multi-scene switching solves the problem of one-sided sound quality adjustment in existing technologies, and realizes dynamic sound quality adjustment and listening experience optimization of audio speakers in complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-07
AI Technical Summary
Existing audio speaker sound quality adaptive adjustment technology fails to effectively integrate environmental characteristics, audio content characteristics, biological characteristics, copyright information, and device aging data, resulting in a one-sided adjustment basis that cannot match the diverse needs in complex scenarios.
The audio speaker sound quality adaptive adjustment system adopts multi-scene switching, including a four-dimensional perception layer, a mixed scene-content-user-copyright fusion analysis module, an auditory memory analysis module, a multi-user preference arbitration module, an extreme environment compensation unit, a copyright-sound quality collaborative adjustment module, a device aging adaptation module, a dynamic decision layer, a feedback execution layer, and a listening perception quantitative evaluation engine. It realizes data interaction and adjustment through a fully closed-loop architecture.
It enables dynamic adjustment of sound quality in multiple scenarios, improving the audio speaker's ability to adapt to different environments, users, and device conditions, ensuring consistency and optimization of the listening experience.
Smart Images

Figure CN121815158A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data analysis and artificial intelligence technology, specifically an audio speaker sound quality adaptive adjustment system that switches between multiple scenarios. Background Technology
[0002] An audio speaker is an electroacoustic transducer that converts electrical signals into sound wave signals. It is the core sound-generating unit of an audio playback device. Its core function is to receive electrical signals (analog or digital signals, which need to be decoded and converted into analog electrical signals) from an audio source (such as a mobile phone, player, or computer), and convert the electrical signals into sound waves that can be perceived by the human ear through the mechanical vibration of its internal structure, ultimately realizing the audio content.
[0003] The core working principle of an audio loudspeaker is based on the synergistic effect of electromagnetic induction and mechanical vibration, using the most common moving-coil loudspeaker as an example. The analog electrical signal (current changes with the audio waveform) output from the audio source is transmitted to the "voice coil" (a copper coil wound on a frame that can move freely in a magnetic field) inside the loudspeaker. The voice coil is placed in a fixed magnetic field formed by a "permanent magnet" (such as a neodymium iron boron magnet). According to the "left-hand rule", the voice coil, which carries a changing current, will be subjected to a magnetic force whose direction and magnitude change with the current, thus producing reciprocating mechanical motion. The voice coil is fixedly connected to the loudspeaker's "diaphragm" (usually made of paper, plastic, or metal film, with good elasticity and rigidity). The reciprocating motion of the voice coil drives the diaphragm to vibrate synchronously. The vibration of the diaphragm compresses or stretches the surrounding air, forming mechanical waves (i.e., sound waves) with alternating compression and sparseness. The sound waves travel through the air to the human ear and are ultimately perceived as audio content.
[0004] Existing audio quality adaptive adjustment technologies suffer from insufficient multi-dimensional data perception and fusion: most technologies only collect environmental or user preference data without integrating environmental features, audio content features, biometric features, copyright information, and device aging data for standardized processing. Furthermore, they lack quantitative fusion and analysis of multi-dimensional features, resulting in one-sided adjustment criteria that cannot match the diverse needs of complex scenarios. Therefore, this paper proposes an audio speaker quality adaptive adjustment system with multi-scenario switching to address the above problems. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies and solve at least one of the technical problems mentioned in the background art, this invention proposes an audio speaker sound quality adaptive adjustment system with multi-scene switching.
[0006] The technical solution adopted by the present invention to solve its technical problem is: the audio speaker sound quality adaptive adjustment system with multi-scene switching as described in the present invention includes: a four-dimensional perception layer, a mixed scene-content-user-copyright fusion analysis module, an auditory memory analysis module, a multi-user preference arbitration module, an extreme environment compensation unit, a copyright-sound quality collaborative adjustment module, a device aging adaptation module, a dynamic decision layer, a feedback execution layer, a listening quantification evaluation engine, and a cloud collaborative engine. The four-dimensional perception layer is used to collect environmental features, audio content features, biometric features, copyright information, and device aging data, and output standardized perception data; the hybrid scene-content-user-copyright fusion parsing module receives the data from the four-dimensional perception layer, integrates multi-dimensional features of scene, content, user identity, and copyright, and quantifies and outputs sub-scene weights, user weights, and copyright levels. The auditory memory analysis module dynamically adjusts parameters based on the user's physiological data output by the fusion analysis module and through an auditory fatigue model. The multi-user preference arbitration module resolves multi-user preference conflicts through weight allocation and conflict negotiation based on the user weights output by the fusion parsing module. The extreme environment compensation unit calibrates and enhances the data collected by the four-dimensional perception layer to improve the perception accuracy under extreme environments. The copyright-sound quality collaborative adjustment module matches a differentiated adjustment strategy based on the copyright level output by the fusion analysis module; The device aging adaptation module corrects hardware parameter constraints based on the device aging data output by the four-dimensional perception layer. The dynamic decision-making layer integrates auditory memory, multi-user preferences, extreme environment compensation, copyright constraints, and aging constraints to generate compliant adjustment instructions. The feedback execution layer receives instructions from the dynamic decision layer, performs sound quality adjustment, and implements multi-user audio partitioning. The auditory perception quantification evaluation engine collects audio data output from the feedback execution layer and evaluates the auditory perception effect through objective indicators; The cloud-based collaborative engine synchronizes the quantitative evaluation results of auditory perception, multi-user preferences, and device aging data to the cloud database, providing global data support for the dynamic decision-making layer.
[0007] Preferably, the auditory memory analysis module performs the following steps: S1. Using the modified ISO 1999 auditory fatigue model, a mapping relationship between "listening duration and auditory sensitivity" is established, with the following formula: Where: S(t,f) is the auditory sensitivity at frequency f at time t, S0(f) is the initial sensitivity (collected at t=0), k=0.002 (fatigue coefficient, calibrated through 1000 sets of user listening data), t is the continuous listening duration (unit: minutes), β=0.3 (frequency influence coefficient, β increases to 0.4 when f≥8kHz); S2. EEG brainwave sensor data is collected every 5 minutes, and the proportion of alpha waves is extracted: If the alpha wave accounts for ≤40%, the sensitivity change is calculated using the formula S(t,f), and the EQ parameters are adjusted accordingly. For example, when t=60 minutes and f=8kHz, S(t,f)=S0(f)×(1-0.002×60×8000^0.3). If the sensitivity decreases by 12%, the gain of the 8kHz band will be automatically increased by 1.5dB. If the alpha wave accounts for more than 40%, the "mild fatigue mode" is triggered: the volume is reduced by 3dB, and the gain of high frequencies above 10kHz is reduced by 1dB to avoid hearing damage. S3. Transmit the corrected EQ parameters to the dynamic decision-making layer as the basis for generating adjustment instructions.
[0008] Preferably, the multi-user preference arbitration module performs the following steps: S1. User Weight Allocation: User IDs are obtained via a Bluetooth BLE module (with an accuracy rate ≥ 99%), and combined with user data synchronized from the cloud-based collaborative engine, user weight W_u is calculated. Identity weights: Head of household W_identity = 0.3, Family member W_identity = 0.2, Visitor W_identity = 0.1; Usage frequency weight: For high-frequency users who have used the service ≥ 10 times in the past 30 days, W_frequency = W_identity × 1.2; Real-time engagement weight: For users who actively fine-tune parameters through the APP, W_engagement = W_frequency × 1.3; End-user weight W_u = W_participation; S2. Preference Fusion Calculation: For the same adjustment parameter (e.g., 100Hz gain), collect the preference parameters P_u of all access users (supporting ≥5 people), and calculate the final parameter using the formula: For example: User 1 (W=0.39, P=+5dB), User 2 (W=0.26, P=-3dB), User 3 (W=0.13, P=+2dB), then P_final=(0.39×5+0.26×(-3)+0.13×2) / (0.39+0.26+0.13)=+1.4dB; S3. Conflict Negotiation: If the difference in P_u between any two users is greater than 5dB, a pop-up negotiation window will be triggered in the mobile app. Displays the user preference parameters and weights, and the default strategy is "majority user preference + weight compensation" - for example, 3 people support +3dB, 2 people support -2dB, and after weight calculation, P_final=+1.8dB; The negotiation results are synchronized to the cloud-based collaborative engine, updating the multi-user preference database to ensure consistency in subsequent adjustments; S4. Transmit P_final to the dynamic decision layer, with a single user's listening satisfaction ≥85%.
[0009] The multi-user preference arbitration module identifies users via Bluetooth and calculates user weights by combining identity, usage frequency, and real-time participation. It uses a weighted average to fuse the preference values of multiple users for the same parameter. If there are preference conflicts, the module negotiates through an app pop-up and synchronizes the results to the cloud. Finally, it outputs a unified parameter that takes into account the needs of multiple users, ensuring the listening satisfaction of individual users.
[0010] Preferably, the extreme environment compensation unit includes a biosensor calibration module, a radar signal enhancement module, and an electromagnetic interference suppression module, which respectively perform the following steps: S1. Biosensor Calibration (High Temperature > 45℃): Collect data from a temperature and humidity sensor (model SHT35, range -40℃-125℃) to obtain the ambient temperature T; Temperature compensation is applied to the raw signal PPG_raw from the PPG heart rate sensor (model MAX30102) using the following formula: The accuracy error of the calibrated physiological data is ≤5%; S2. Radar signal enhancement (high humidity > 85%RH scenario): For the spatial data acquired by the millimeter-wave radar (model AWR1642), an adaptive Kalman filter algorithm is used to remove water vapor interference; The 3dB attenuation gain is recovered through an antenna gain compensation algorithm, ensuring that the spatial measurement error is ≤±0.2m; S3. Electromagnetic Interference Suppression (EMC Level ≥ EN55032 Class B Scenarios): The Bluetooth module (version 5.3) uses frequency hopping communication with a hopping rate of ≥1600 times / second to avoid interference frequency bands; The audio semantic acquisition submodule employs signal shielding algorithms (such as wavelet threshold denoising) to reduce the impact of electromagnetic interference on audio parsing, achieving a semantic parsing misjudgment rate of ≤5%. S4. Transmit the calibrated sensing data to the four-dimensional sensing layer to ensure the reliability of the standardized data.
[0011] Preferably, the copyright-sound quality collaborative adjustment module performs the following steps: S1. Copyright Restriction Resolution: Read the copyright information from the audio file's metadata (ID3 tag) or streaming media protocol (HLS) and establish a "Copyright Level - Adjustment Permission" mapping table: Copyright Level Allowed EQ Gain Range Prohibited Sound Quality Optimization Direction Lossless (FLAC) ±6dB None 20Hz-20kHz Full-Band Equalization Enhancement Standard (MP3) ±4dB Prohibited 100-200Hz Low-Frequency Over-Enhancement (Single Gain ≤2dB) 1-3kHz Vocal Band Optimization Decryption (DRM) ±2dB Prohibited High-Frequency Modification Above 8kHz, Prohibited Oversampling Processing Maintain 200-5kHz Basic Listening Clarity S2. Dynamic permission adaptation: If the resolution is encrypted copyright, the adjustment function of the frequency band above 8kHz is locked through the hardware interface, and the EQ gain adjustment range is limited to ±2dB to avoid triggering the DRM copyright protection mechanism and causing playback interruption. If the resolution is deemed lossless, unlock full-band adjustment permissions, allowing ±6dB gain adjustment; S3. Unlocking Sound Quality Potential (For Lossless Audio Only): An oversampling algorithm is used to increase the audio sampling rate from 44.1kHz to 192kHz (achieved through the DSP chip TITMS320C5535); combined with the content characteristics output by the fusion analysis module (such as the need to enhance low frequencies in symphonic music), EQ parameters are optimized within ±6dB, achieving a sound quality potential utilization rate of ≥90%. S4. Transmit the adapted adjustment permissions and optimization parameters to the dynamic decision-making layer to ensure a balance between compliance and sound quality.
[0012] Preferably, the device aging adaptation module performs the following steps: S1. Aging Level Assessment: The speaker's operating current is collected using a current sensor (model ACS712), and the current fluctuation amplitude ΔI (%) is calculated as follows: (Maximum Current - Minimum Current) / Rated Current × 100%; Read the cumulative speaker usage time T (in hours) stored in the EEPROM, and evaluate the aging level L using the formula: Where L∈[1,5] (Level 1 is brand new, Level 5 is severely aged), for example, T=1500 hours, ΔI=18%, then L=1+(1500 / 1000)+(18 / 15)=3.2, which is rounded to Level 3; S2. Hardware parameter correction: The speaker frequency response range and distortion threshold are adjusted according to the aging level L: L=1 level: Frequency response 20Hz-20kHz, distortion threshold ≤3%; L=3 level: Frequency response 50Hz-18kHz, distortion threshold ≤4%; L=5 level: Frequency response 80Hz-16kHz, distortion threshold ≤5%; S3. Aging Compensation Adjustment: Automatically increases the gain of the corresponding frequency band to address frequency band attenuation caused by aging. L=Level 2: 50-100Hz gain +1dB; L=Level 3: 50-100Hz gain +2dB, 1-3kHz vocal band gain +0.5dB; L=4 levels: 50-100Hz gain +3dB, 1-3kHz gain +1dB; S4. Transmit the corrected hardware parameters and compensation gain to the dynamic decision layer to ensure that the adjustment parameters match the hardware capabilities after aging, with a distortion of ≤3%.
[0013] Preferably, the auditory quantification evaluation engine performs the following steps: S1. Objective index calculation: Speech intelligibility index (STI): The transmission index of the signal in the 125Hz-8kHz frequency band is calculated by collecting the speech signal output from the speaker (such as the standard ITU-T P.501 test signal). The range is 0-1, and ≥0.7 is judged as clear. Auditory Comfort Index (ACI): Combining the proportion of alpha waves in EEG (≤40% is comfortable) and the volume level (≤85dB is comfortable), the formula is ACI=(1-alpha wave proportion / 100)×(1-(volume-60) / 25)×100, with a range of 0-100, and ≥80 is considered comfortable; Total Harmonic Distortion (THD): Acquire a 1kHz sinusoidal test signal, calculate the power ratio of harmonic components to fundamental components, and a value ≤3% is considered acceptable. S2. Evaluation-adjustment closed loop: If STI < 0.7: Automatically increase the gain of the 1-3kHz vocal band by 0.5-1dB until STI ≥ 0.7; If ACI < 80: Reduce the gain of high frequencies above 8kHz by 0.5dB, and at the same time lower the volume by 2-3dB until ACI ≥ 80; If THD > 3%: Reduce the gain of the current frequency band by 0.3-0.5dB (prioritize adjusting the low frequency 100-200Hz) until THD ≤ 3%; S3. User feedback calibration: After every 10 closed-loop adjustments, the user's subjective evaluation (clear / average / fuzzy) is obtained through an APP pop-up window. If the deviation between subjective evaluation and objective indicators is greater than 10% (e.g., STI=0.75 but user feedback is ambiguous), adjust the indicator weights—for example, increase the proportion of ACI in the evaluation, so that the deviation between quantitative evaluation and subjective feelings is ≤3%; S4. Synchronize the evaluation results and the revised indicator weights to the cloud-based collaborative engine to optimize subsequent adjustment strategies.
[0014] Preferably, the four-dimensional perception layer includes the following sub-modules and parameters: S1. Auditory fatigue perception sub-module: integrates an EEG brainwave sensor, sampling frequency ≥256Hz, data transmission delay ≤10ms, used to collect alpha wave signals; S2. Multi-user identity perception unit: adopts a Bluetooth BLE 5.3 module, supports simultaneous access of ≥5 user devices, identity recognition accuracy ≥99%, response time ≤50ms; S3. Device aging monitoring unit: integrates a current sensor, measurement range 0-5A, accuracy ±3%, used to monitor speaker operating current fluctuations; EEPROM storage chip (capacity ≥4KB), records the speaker's cumulative usage time (accuracy ±1 hour); S4. Copyright information acquisition sub-module: reads data via USB / Bluetooth interface. Audio file metadata is retrieved with ID3 tag parsing latency ≤20ms; copyright information extraction for HLS / DASH streaming media protocols is supported with an accuracy rate ≥98%; S5. Extreme Environment Monitoring Submodule: Temperature and humidity sensor, range -40℃-125℃ (temperature), 0%-100%RH (humidity), accuracy ±0.3℃ / ±2%RH; Electromagnetic interference sensor, measurement range -60dBm-0dBm, used to monitor electromagnetic interference intensity; S6. Environment-Content Awareness Submodule: Noise sensor, range 30-120dB, frequency response 20Hz-20kHz; Millimeter-wave radar, spatial measurement range 0.2-10m, accuracy ±0.1m; Audio semantic acquisition submodule, used to parse audio content type.
[0015] Preferably, the dynamic decision-making layer adopts a five-level architecture of "rule engine + reinforcement learning + physiological constraints + copyright constraints + aging constraints" and performs the following steps: S1. Constraint priority sorting: highest priority: physiological constraints (EEG alpha wave proportion > 40% triggers fatigue warning), copyright constraints (encrypted audio is prohibited from being modified above 8kHz); Second highest priority: Aging constraint (L=3 level time frequency response limit is 50Hz-18kHz); Basic priorities: rule engine (e.g., default bass +2dB in home scenarios), reinforcement learning (optimizing parameters based on historical adjustment data); S2. Adjustment instruction generation: Receive parameters output from the multi-dimensional adaptation layer (auditory memory, multi-user arbitration, etc.), first check whether they meet the highest priority constraints - for example, if the encrypted audio parameter contains 8kHz+1dB, it will be directly rejected and corrected to 8kHz±0dB; If all constraints are met, the parameters are optimized using a reinforcement learning model (based on the Q-Learning algorithm, with ≥1000 iterations), and the instruction generation frequency is ≥20Hz. If the second-highest priority constraint is not met (e.g., parameters exceed the frequency response range after aging), the gain of the frequency band exceeding the limit will be automatically clipped (e.g., the gain below 50Hz will be set to 0dB). S3. Response lag control: Adopting a pipelined processing architecture, the perception data acquisition, constraint verification, and instruction generation are executed in parallel, with a dynamic scene response lag of ≤20ms; S4. Permission Verification Output: After the instruction is generated, it is verified again by the permission verification unit (storing the copyright-aging constraint table) to ensure that there are no illegal parameters (such as lossless audio gain ≤6dB). After the verification is passed, it is output to the feedback execution layer.
[0016] Preferably, the feedback execution layer includes the following sub-modules and steps: S1. Multi-user audio partition unit: Based on Bluetooth BLE grouping technology, access users are divided into different audio zones (such as the living room sofa area and the study area). For each zone, output the corresponding preference parameters—for example, the sofa zone (elderly, 100Hz+2dB) and the study zone (children, 1kHz+1dB), with a zone switching response time of ≤100ms; S2. Aging Protection Module: Real-time monitoring of the speaker's operating power (calculated via a current sensor: power = voltage × current) and comparison with the rated power after aging (e.g., the original 20W is reduced to 15W after aging); If the actual power is ≥90% of the rated power (e.g., 15W × 90% = 13.5W), the volume gain will be automatically limited to ≤10dB to avoid hardware overload. If the power continues for 10 seconds or more at or above the rated power, an audible and visual warning will be triggered (LED red light flashing + 1kHz warning sound). S3. Adaptive frequency division unit: Analyze audio content features (such as vocals / instruments / sound effects) and dynamically allocate frequency bands: Human voice content (news / conference): The 1-3kHz frequency band is processed separately to improve separation; Symphony content: The 20-200Hz low frequency and 2-8kHz high frequency are divided into separate frequencies to enhance the sense of layering; After crossover, the output is through a Class D amplifier (model TITPA3116D2), improving vocal separation by 20%. S4. Adjustment effect feedback: The audio signal output from the power amplifier is collected and transmitted to the listening quantification evaluation engine to provide data support for closed-loop optimization.
[0017] The beneficial effects of this invention are: This invention provides a multi-scene switching audio speaker sound quality adaptive adjustment system. The system adopts a fully closed-loop architecture of "perception-analysis-processing-decision-execution-evaluation-coordination." Each module interacts with the others via hardware interfaces (such as SPI, I2C, and Bluetooth 5.3) and software protocols (such as MQTT). The four-dimensional perception layer collects and standardizes environmental, content, biological, copyright, and aging data; the hybrid fusion analysis module extracts and quantifies multi-dimensional features; the auditory memory, multi-user arbitration, extreme environment compensation, copyright-sound quality coordination, and device aging adaptation modules process data from the dimensions of fatigue, preference, environment, copyright, and hardware, respectively; the dynamic decision-making layer merges the outputs of multiple modules according to constraint priorities to generate compliant adjustment instructions; the feedback execution layer executes the adjustment and implements zoning; the auditory quantification evaluation engine optimizes the adjustment effect through objective indicators and user feedback; and the cloud collaboration engine synchronizes global data to support decision-making, forming a fully closed-loop adaptive adjustment system of "perception-processing-decision-execution-evaluation-coordination."
[0018] This invention provides an audio speaker sound quality adaptive adjustment system that allows for multi-scene switching. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention.
[0020] In the attached diagram: Figure 1 This is a schematic diagram of the system flow of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Specific implementation examples are given below.
[0023] Please see Figure 1 This invention provides an audio speaker sound quality adaptive adjustment system with multi-scene switching, including: a four-dimensional perception layer, a mixed scene-content-user-copyright fusion analysis module, an auditory memory analysis module, a multi-user preference arbitration module, an extreme environment compensation unit, a copyright-sound quality collaborative adjustment module, a device aging adaptation module, a dynamic decision layer, a feedback execution layer, a listening perception quantitative evaluation engine, and a cloud collaborative engine. The four-dimensional perception layer is used to collect environmental features, audio content features, biometric features, copyright information, and device aging data, and output standardized perception data; the hybrid scene-content-user-copyright fusion parsing module receives the data from the four-dimensional perception layer, integrates multi-dimensional features of scene, content, user identity, and copyright, and quantifies and outputs sub-scene weights, user weights, and copyright levels. The auditory memory analysis module dynamically adjusts parameters based on the user's physiological data output by the fusion analysis module and through an auditory fatigue model. The multi-user preference arbitration module resolves multi-user preference conflicts through weight allocation and conflict negotiation based on the user weights output by the fusion parsing module. The extreme environment compensation unit calibrates and enhances the data collected by the four-dimensional perception layer to improve the perception accuracy under extreme environments. The copyright-sound quality collaborative adjustment module matches a differentiated adjustment strategy based on the copyright level output by the fusion analysis module; The device aging adaptation module corrects hardware parameter constraints based on the device aging data output by the four-dimensional perception layer. The dynamic decision-making layer integrates auditory memory, multi-user preferences, extreme environment compensation, copyright constraints, and aging constraints to generate compliant adjustment instructions. The feedback execution layer receives instructions from the dynamic decision layer, performs sound quality adjustment, and implements multi-user audio partitioning. The auditory perception quantification evaluation engine collects audio data output from the feedback execution layer and evaluates the auditory perception effect through objective indicators; The cloud-based collaborative engine synchronizes the quantitative evaluation results of auditory perception, multi-user preferences, and device aging data to the cloud database, providing global data support for the dynamic decision-making layer.
[0024] This system adopts a fully closed-loop architecture of "perception-analysis-processing-decision-execution-evaluation-collaboration". Each module interacts with the other through hardware interfaces (such as SPI, I2C, Bluetooth 5.3) and software protocols (such as MQTT). The four-dimensional perception layer collects and standardizes environmental, content, biological, copyright, and aging data; the hybrid fusion analysis module extracts and quantifies multi-dimensional features; the auditory memory, multi-user arbitration, extreme environment compensation, copyright-sound quality collaboration, and device aging adaptation modules process data from the dimensions of fatigue, preference, environment, copyright, and hardware, respectively; the dynamic decision-making layer merges the outputs of multiple modules according to constraint priority to generate compliant adjustment instructions; the feedback execution layer executes the adjustment and implements zoning; the auditory quantification evaluation engine optimizes the adjustment effect through objective indicators and user feedback; and the cloud collaboration engine synchronizes global data to support decision-making, forming a fully closed-loop adaptive adjustment of "perception-processing-decision-execution-evaluation-collaboration".
[0025] Furthermore, such as Figure 1As shown, the auditory memory analysis module performs the following steps: S1. Using the modified ISO 1999 auditory fatigue model, a mapping relationship between "listening duration and auditory sensitivity" is established, with the following formula: Where: S(t,f) is the auditory sensitivity at frequency f at time t, S0(f) is the initial sensitivity (collected at t=0), k=0.002 (fatigue coefficient, calibrated through 1000 sets of user listening data), t is the continuous listening duration (unit: minutes), β=0.3 (frequency influence coefficient, β increases to 0.4 when f≥8kHz); S2. EEG brainwave sensor data is collected every 5 minutes, and the proportion of alpha waves is extracted: If the alpha wave accounts for ≤40%, the sensitivity change is calculated using the formula S(t,f), and the EQ parameters are adjusted accordingly. For example, when t=60 minutes and f=8kHz, S(t,f)=S0(f)×(1-0.002×60×8000^0.3). If the sensitivity decreases by 12%, the gain of the 8kHz band will be automatically increased by 1.5dB. If the alpha wave accounts for more than 40%, the "mild fatigue mode" is triggered: the volume is reduced by 3dB, and the gain of high frequencies above 10kHz is reduced by 1dB to avoid hearing damage. S3. Transmit the corrected EQ parameters to the dynamic decision-making layer as the basis for generating adjustment instructions.
[0026] The auditory memory analysis module is based on the modified ISO1999 model. It uses EEG sensors to periodically collect the proportion of alpha waves to determine the user's auditory fatigue state. When not fatigued, it corrects EQ parameters based on the mapping relationship of "listening duration-frequency-sensitivity" to compensate for the decrease in sensitivity. When fatigued, it triggers the protection mode to adjust the volume and high-frequency gain. Finally, the adaptation parameters are transmitted to the dynamic decision layer to achieve dynamic matching of "physiological state-sound quality parameters".
[0027] Furthermore, such as Figure 1 As shown, the multi-user preference arbitration module performs the following steps: S1. User Weight Allocation: User IDs are obtained via a Bluetooth BLE module (with an accuracy rate ≥ 99%), and combined with user data synchronized from the cloud-based collaborative engine, user weight W_u is calculated. Identity weights: Head of household W_identity = 0.3, Family member W_identity = 0.2, Visitor W_identity = 0.1; Usage frequency weight: For high-frequency users who have used the service ≥ 10 times in the past 30 days, W_frequency = W_identity × 1.2; Real-time engagement weight: For users who actively fine-tune parameters through the APP, W_engagement = W_frequency × 1.3; End-user weight W_u = W_participation; S2. Preference Fusion Calculation: For the same adjustment parameter (e.g., 100Hz gain), collect the preference parameters P_u of all access users (supporting ≥5 people), and calculate the final parameter using the formula: For example: User 1 (W=0.39, P=+5dB), User 2 (W=0.26, P=-3dB), User 3 (W=0.13, P=+2dB), then P_final=(0.39×5+0.26×(-3)+0.13×2) / (0.39+0.26+0.13)=+1.4dB; S3. Conflict Negotiation: If the difference in P_u between any two users is greater than 5dB, a pop-up negotiation window will be triggered in the mobile app. Displays the user preference parameters and weights, and the default strategy is "majority user preference + weight compensation" - for example, 3 people support +3dB, 2 people support -2dB, and after weight calculation, P_final=+1.8dB; The negotiation results are synchronized to the cloud-based collaborative engine, updating the multi-user preference database to ensure consistency in subsequent adjustments; S4. Transmit P_final to the dynamic decision layer, with a single user's listening satisfaction ≥85%.
[0028] The multi-user preference arbitration module identifies users via Bluetooth and calculates user weights by combining identity, usage frequency, and real-time participation. It uses a weighted average to fuse the preference values of multiple users for the same parameter. If there are preference conflicts, the module negotiates through an app pop-up and synchronizes the results to the cloud. Finally, it outputs a unified parameter that takes into account the needs of multiple users, ensuring the listening satisfaction of individual users.
[0029] Furthermore, such as Figure 1 As shown, the extreme environment compensation unit includes a biosensor calibration module, a radar signal enhancement module, and an electromagnetic interference suppression module, which respectively perform the following steps: S1. Biosensor Calibration (High Temperature > 45℃): Collect data from a temperature and humidity sensor (model SHT35, range -40℃-125℃) to obtain the ambient temperature T; Temperature compensation is applied to the raw signal PPG_raw from the PPG heart rate sensor (model MAX30102) using the following formula: The accuracy error of the calibrated physiological data is ≤5%; S2. Radar signal enhancement (high humidity > 85%RH scenario): For the spatial data acquired by the millimeter-wave radar (model AWR1642), an adaptive Kalman filter algorithm is used to remove water vapor interference; The 3dB attenuation gain is recovered through an antenna gain compensation algorithm, ensuring that the spatial measurement error is ≤±0.2m; S3. Electromagnetic Interference Suppression (EMC Level ≥ EN55032 Class B Scenarios): The Bluetooth module (version 5.3) uses frequency hopping communication with a hopping rate of ≥1600 times / second to avoid interference frequency bands; The audio semantic acquisition submodule employs signal shielding algorithms (such as wavelet threshold denoising) to reduce the impact of electromagnetic interference on audio parsing, achieving a semantic parsing misjudgment rate of ≤5%. S4. Transmit the calibrated sensing data to the four-dimensional sensing layer to ensure the reliability of the standardized data.
[0030] The extreme environment compensation unit employs differentiated processing for different extreme scenarios: at high temperatures, it calibrates the PPG sensor signal using a linear formula to compensate for the impact of temperature on biological data; at high humidity, it removes radar moisture interference and compensates for antenna gain using adaptive Kalman filtering; and in the event of electromagnetic interference, it suppresses interference through Bluetooth frequency hopping and wavelet denoising. Finally, it outputs high-precision sensing data to ensure the accuracy of system regulation under extreme environments.
[0031] Furthermore, such as Figure 1 As shown, the copyright-sound quality collaborative adjustment module performs the following steps: S1. Copyright Restriction Resolution: Read the copyright information from the audio file's metadata (ID3 tag) or streaming media protocol (HLS) and establish a "Copyright Level - Adjustment Permission" mapping table: Copyright Level Allowed EQ Gain Range Prohibited Sound Quality Optimization Direction Lossless (FLAC) ±6dB None 20Hz-20kHz Full-Band Equalization Enhancement Standard (MP3) ±4dB Prohibited 100-200Hz Low-Frequency Over-Enhancement (Single Gain ≤2dB) 1-3kHz Vocal Band Optimization Decryption (DRM) ±2dB Prohibited High-Frequency Modification Above 8kHz, Prohibited Oversampling Processing Maintain 200-5kHz Basic Listening Clarity S2. Dynamic permission adaptation: If the resolution is encrypted copyright, the adjustment function of the frequency band above 8kHz is locked through the hardware interface, and the EQ gain adjustment range is limited to ±2dB to avoid triggering the DRM copyright protection mechanism and causing playback interruption. If the resolution is deemed lossless, unlock full-band adjustment permissions, allowing ±6dB gain adjustment; S3. Unlocking Sound Quality Potential (For Lossless Audio Only): An oversampling algorithm is used to increase the audio sampling rate from 44.1kHz to 192kHz (achieved through the DSP chip TITMS320C5535); combined with the content characteristics output by the fusion analysis module (such as the need to enhance low frequencies in symphonic music), EQ parameters are optimized within ±6dB, achieving a sound quality potential utilization rate of ≥90%. S4. Transmit the adapted adjustment permissions and optimization parameters to the dynamic decision-making layer to ensure a balance between compliance and sound quality.
[0032] The copyright-sound quality collaborative adjustment module obtains the copyright level by parsing audio file metadata or streaming media protocols, and matches adjustment permissions based on a preset mapping table; encrypted copyright locks high-frequency adjustment and oversampling, standard copyright restricts low-frequency gain, and lossless copyright unlocks full permissions; for lossless audio, it optimizes sound quality through DSP oversampling and content adaptation; finally, it outputs compliant and optimized adjustment parameters, balancing copyright protection and listening experience.
[0033] Furthermore, such as Figure 1 As shown, the device aging adaptation module performs the following steps: S1. Aging Level Assessment: The speaker's operating current is collected using a current sensor (model ACS712), and the current fluctuation amplitude ΔI (%) is calculated as follows: (Maximum Current - Minimum Current) / Rated Current × 100%; Read the cumulative speaker usage time T (in hours) stored in the EEPROM, and evaluate the aging level L using the formula: Where L∈[1,5] (Level 1 is brand new, Level 5 is severely aged), for example, T=1500 hours, ΔI=18%, then L=1+(1500 / 1000)+(18 / 15)=3.2, which is rounded to Level 3; S2. Hardware parameter correction: The speaker frequency response range and distortion threshold are adjusted according to the aging level L: L=1 level: Frequency response 20Hz-20kHz, distortion threshold ≤3%; L=3 level: Frequency response 50Hz-18kHz, distortion threshold ≤4%; L=5 level: Frequency response 80Hz-16kHz, distortion threshold ≤5%; S3. Aging Compensation Adjustment: Automatically increases the gain of the corresponding frequency band to address frequency band attenuation caused by aging. L=Level 2: 50-100Hz gain +1dB; L=Level 3: 50-100Hz gain +2dB, 1-3kHz vocal band gain +0.5dB; L=4 levels: 50-100Hz gain +3dB, 1-3kHz gain +1dB; S4. Transmit the corrected hardware parameters and compensation gain to the dynamic decision layer to ensure that the adjustment parameters match the hardware capabilities after aging, with a distortion of ≤3%.
[0034] The device aging adaptation module collects the speaker current fluctuation amplitude through a current sensor, combines it with the cumulative usage time stored in the EEPROM, and evaluates the aging level using a formula; it corrects the speaker frequency response range and distortion threshold according to the aging level; it automatically compensates for the corresponding frequency band gain for different levels of frequency band attenuation; and finally outputs the parameters to adapt to the aging hardware, ensuring that the sound quality adjustment is within the hardware's capabilities and avoiding distortion.
[0035] Furthermore, such as Figure 1 As shown, the auditory quantification evaluation engine performs the following steps: S1. Objective index calculation: Speech intelligibility index (STI): The transmission index of the signal in the 125Hz-8kHz frequency band is calculated by collecting the speech signal output from the speaker (such as the standard ITU-T P.501 test signal). The range is 0-1, and ≥0.7 is judged as clear. Auditory Comfort Index (ACI): Combining the proportion of alpha waves in EEG (≤40% is comfortable) and the volume level (≤85dB is comfortable), the formula is ACI=(1-alpha wave proportion / 100)×(1-(volume-60) / 25)×100, with a range of 0-100, and ≥80 is considered comfortable; Total Harmonic Distortion (THD): Acquire a 1kHz sinusoidal test signal, calculate the power ratio of harmonic components to fundamental components, and a value ≤3% is considered acceptable. S2. Evaluation-adjustment closed loop: If STI < 0.7: Automatically increase the gain of the 1-3kHz vocal band by 0.5-1dB until STI ≥ 0.7; If ACI < 80: Reduce the gain of high frequencies above 8kHz by 0.5dB, and at the same time lower the volume by 2-3dB until ACI ≥ 80; If THD > 3%: Reduce the gain of the current frequency band by 0.3-0.5dB (prioritize adjusting the low frequency 100-200Hz) until THD ≤ 3%; S3. User feedback calibration: After every 10 closed-loop adjustments, the user's subjective evaluation (clear / average / fuzzy) is obtained through an APP pop-up window. If the deviation between subjective evaluation and objective indicators is greater than 10% (e.g., STI=0.75 but user feedback is ambiguous), adjust the indicator weights—for example, increase the proportion of ACI in the evaluation, so that the deviation between quantitative evaluation and subjective feelings is ≤3%; S4. Synchronize the evaluation results and the revised indicator weights to the cloud-based collaborative engine to optimize subsequent adjustment strategies.
[0036] The auditory quantification assessment engine calculates three objective indicators—STI (Sharpness Intensity), ACI (Adequacy Intensity), and THD (Total Distortion)—by collecting test signals. It then performs closed-loop adjustments on indicators that fail to meet the standards (boosting vocal frequencies, reducing high frequencies and volume, and optimizing low-frequency gain). The engine also incorporates user feedback to calibrate indicator weights, reducing the discrepancy between quantification and subjective perception. Finally, the assessment results and weights are synchronized to the cloud, providing a basis for system adjustments.
[0037] Furthermore, such as Figure 1 As shown, the four-dimensional perception layer includes the following sub-modules and parameters: S1. Auditory fatigue perception sub-module: integrates an EEG brainwave sensor, sampling frequency ≥256Hz, data transmission delay ≤10ms, used to collect alpha wave signals; S2. Multi-user identity perception unit: adopts a Bluetooth BLE 5.3 module, supports simultaneous access of ≥5 user devices, identity recognition accuracy ≥99%, response time ≤50ms; S3. Device aging monitoring unit: integrates a current sensor, measurement range 0-5A, accuracy ±3%, used to monitor speaker operating current fluctuations; EEPROM storage chip (capacity ≥4KB), records the speaker's cumulative usage time (accuracy ±1 hour); S4. Copyright information acquisition sub-module: reads data via USB / Bluetooth interface. Audio file metadata is retrieved with ID3 tag parsing latency ≤20ms; copyright information extraction for HLS / DASH streaming media protocols is supported with an accuracy rate ≥98%; S5. Extreme Environment Monitoring Submodule: Temperature and humidity sensor, range -40℃-125℃ (temperature), 0%-100%RH (humidity), accuracy ±0.3℃ / ±2%RH; Electromagnetic interference sensor, measurement range -60dBm-0dBm, used to monitor electromagnetic interference intensity; S6. Environment-Content Awareness Submodule: Noise sensor, range 30-120dB, frequency response 20Hz-20kHz; Millimeter-wave radar, spatial measurement range 0.2-10m, accuracy ±0.1m; Audio semantic acquisition submodule, used to parse audio content type.
[0038] The four-dimensional perception layer achieves multi-dimensional data acquisition through six sub-modules: the auditory fatigue perception sub-module acquires EEG alpha waves, the multi-user identity perception unit identifies user IDs and roles, the device aging monitoring unit acquires current fluctuations and usage time, the copyright information acquisition sub-module parses copyright levels, the extreme environment monitoring sub-module acquires temperature, humidity, and electromagnetic interference, and the environment-content perception sub-module acquires noise, spatial location, and content type. All data is standardized, encapsulated, and transmitted to the fusion and parsing module, providing basic data support for subsequent system processing.
[0039] Furthermore, such as Figure 1As shown, the dynamic decision-making layer adopts a five-level architecture of "rule engine + reinforcement learning + physiological constraints + copyright constraints + aging constraints" and executes the following steps: S1. Constraint priority sorting: highest priority: physiological constraints (EEG alpha wave proportion > 40% triggers fatigue warning), copyright constraints (encrypted audio is prohibited from being modified above 8kHz); Second highest priority: Aging constraint (L=3 level time frequency response limit is 50Hz-18kHz); Basic priorities: rule engine (e.g., default bass +2dB in home scenarios), reinforcement learning (optimizing parameters based on historical adjustment data); S2. Adjustment instruction generation: Receive parameters output from the multi-dimensional adaptation layer (auditory memory, multi-user arbitration, etc.), first check whether they meet the highest priority constraints - for example, if the encrypted audio parameter contains 8kHz+1dB, it will be directly rejected and corrected to 8kHz±0dB; If all constraints are met, the parameters are optimized using a reinforcement learning model (based on the Q-Learning algorithm, with ≥1000 iterations), and the instruction generation frequency is ≥20Hz. If the second-highest priority constraint is not met (e.g., parameters exceed the frequency response range after aging), the gain of the frequency band exceeding the limit will be automatically clipped (e.g., the gain below 50Hz will be set to 0dB). S3. Response lag control: Adopting a pipelined processing architecture, the perception data acquisition, constraint verification, and instruction generation are executed in parallel, with a dynamic scene response lag of ≤20ms; S4. Permission Verification Output: After the instruction is generated, it is verified again by the permission verification unit (storing the copyright-aging constraint table) to ensure that there are no illegal parameters (such as lossless audio gain ≤6dB). After the verification is passed, it is output to the feedback execution layer.
[0040] The dynamic decision-making layer first clarifies the constraint priorities (physiological and copyright constraints are highest, followed by aging, and then rules and reinforcement learning fundamentals). After receiving data from multiple modules, it verifies the constraints sequentially according to priority, rejecting or correcting non-compliant parameters. It then overlays scenario-adaptive parameters through a rule engine and optimizes parameters through reinforcement learning. A pipelined architecture is adopted to reduce response lag, and secondary verification ensures instruction compliance. Finally, it outputs precise and compliant adjustment instructions to coordinate the functions of various modules in the system. Furthermore, such as Figure 1 As shown, the feedback execution layer includes the following sub-modules and steps: S1. Multi-user audio partition unit: Based on Bluetooth BLE grouping technology, access users are divided into different audio zones (such as the living room sofa area and the study area). For each zone, output the corresponding preference parameters—for example, the sofa zone (elderly, 100Hz+2dB) and the study zone (children, 1kHz+1dB), with a zone switching response time of ≤100ms; S2. Aging Protection Module: Real-time monitoring of the speaker's operating power (calculated via a current sensor: power = voltage × current) and comparison with the rated power after aging (e.g., the original 20W is reduced to 15W after aging); If the actual power is ≥90% of the rated power (e.g., 15W × 90% = 13.5W), the volume gain will be automatically limited to ≤10dB to avoid hardware overload. If the power continues for 10 seconds or more at or above the rated power, an audible and visual warning will be triggered (LED red light flashing + 1kHz warning sound). S3. Adaptive frequency division unit: Analyze audio content features (such as vocals / instruments / sound effects) and dynamically allocate frequency bands: Human voice content (news / conference): The 1-3kHz frequency band is processed separately to improve separation; Symphony content: The 20-200Hz low frequency and 2-8kHz high frequency are divided into separate frequencies to enhance the sense of layering; After crossover, the output is through a Class D amplifier (model TITPA3116D2), improving vocal separation by 20%. S4. Adjustment effect feedback: The audio signal output from the power amplifier is collected and transmitted to the listening quantification evaluation engine to provide data support for closed-loop optimization.
[0041] The feedback execution layer enables multi-user audio partitioning through Bluetooth grouping and radar positioning, outputting parameters that match user preferences for different partitions; the aging protection module monitors speaker power in real time, triggering warnings and gain limiting to prevent hardware overload; the adaptive crossover unit dynamically allocates frequency bands according to the audio content, improving sound quality; finally, the collected output signal is fed back to the evaluation engine, forming a closed loop of adjustment effect to ensure user listening experience and hardware safety.
[0042] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A multi-scene switching audio speaker sound quality adaptive adjustment system, characterized in that: include: The system includes a four-dimensional perception layer, a mixed scene-content-user-copyright fusion analysis module, an auditory memory analysis module, a multi-user preference arbitration module, an extreme environment compensation unit, a copyright-sound quality collaborative adjustment module, a device aging adaptation module, a dynamic decision-making layer, a feedback execution layer, an auditory quantitative evaluation engine, and a cloud-based collaborative engine. The four-dimensional perception layer is used to collect environmental features, audio content features, biometric features, copyright information, and device aging data, and output standardized perception data; the hybrid scene-content-user-copyright fusion parsing module receives the data from the four-dimensional perception layer, integrates multi-dimensional features of scene, content, user identity, and copyright, and quantifies and outputs sub-scene weights, user weights, and copyright levels. The auditory memory analysis module dynamically adjusts parameters based on the user's physiological data output by the fusion analysis module and through an auditory fatigue model. The multi-user preference arbitration module resolves multi-user preference conflicts through weight allocation and conflict negotiation based on the user weights output by the fusion parsing module. The extreme environment compensation unit calibrates and enhances the data collected by the four-dimensional perception layer to improve the perception accuracy under extreme environments. The copyright-sound quality collaborative adjustment module matches a differentiated adjustment strategy based on the copyright level output by the fusion analysis module; The device aging adaptation module corrects hardware parameter constraints based on the device aging data output by the four-dimensional perception layer. The dynamic decision-making layer integrates auditory memory, multi-user preferences, extreme environment compensation, copyright constraints, and aging constraints to generate compliant adjustment instructions. The feedback execution layer receives instructions from the dynamic decision layer, performs sound quality adjustment, and implements multi-user audio partitioning. The auditory perception quantification evaluation engine collects audio data output from the feedback execution layer and evaluates the auditory perception effect through objective indicators; The cloud-based collaborative engine synchronizes the quantitative evaluation results of auditory perception, multi-user preferences, and device aging data to the cloud database, providing global data support for the dynamic decision-making layer.
2. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The auditory memory analysis module performs the following steps: S1. Using the modified ISO 1999 auditory fatigue model, a mapping relationship between "listening duration and auditory sensitivity" is established, with the following formula: Where: S(t,f) is the auditory sensitivity at frequency f at time t, S0(f) is the initial sensitivity (collected at t=0), k=0.002 (fatigue coefficient, calibrated through 1000 sets of user listening data), t is the continuous listening duration (unit: minutes), β=0.3 (frequency influence coefficient, β increases to 0.4 when f≥8kHz); S2. EEG brainwave sensor data is collected every 5 minutes, and the proportion of alpha waves is extracted: If the alpha wave accounts for ≤40%, the sensitivity change is calculated using the formula S(t,f), and the EQ parameters are adjusted accordingly. For example, when t=60 minutes and f=8kHz, S(t,f)=S0(f)×(1-0.002×60×8000^0.3). If the sensitivity decreases by 12%, the gain of the 8kHz band will be automatically increased by 1.5dB. If the alpha wave accounts for more than 40%, the "mild fatigue mode" is triggered: the volume is reduced by 3dB, and the gain of high frequencies above 10kHz is reduced by 1dB to avoid hearing damage. S3. Transmit the corrected EQ parameters to the dynamic decision-making layer as the basis for generating adjustment instructions.
3. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The multi-user preference arbitration module performs the following steps: S1. User Weight Allocation: User IDs are obtained via a Bluetooth BLE module (with an accuracy rate ≥ 99%), and combined with user data synchronized from the cloud-based collaborative engine, user weight W_u is calculated. Identity weights: Head of household W_identity = 0.3, Family member W_identity = 0.2, Visitor W_identity = 0.1; Usage frequency weight: For high-frequency users who have used the service ≥ 10 times in the past 30 days, W_frequency = W_identity × 1.2; Real-time engagement weight: For users who actively fine-tune parameters through the APP, W_engagement = W_frequency × 1.3; End-user weight W_u = W_participation; S2. Preference Fusion Calculation: For the same adjustment parameter (e.g., 100Hz gain), collect the preference parameters P_u of all access users (supporting ≥5 people), and calculate the final parameter using the formula: ; For example: User 1 (W=0.39, P=+5dB), User 2 (W=0.26, P=-3dB), User 3 (W=0.13, P=+2dB), then P_final=(0.39×5+0.26×(-3)+0.13×2) / (0.39+0.26+0.13)=+1.4dB; S3. Conflict Negotiation: If the difference in P_u between any two users is greater than 5dB, a pop-up negotiation window will be triggered in the mobile app. Displays the preference parameters and weights of each user. The default strategy is "majority preference + weight compensation" - for example, if 3 people support it, the weight is +3dB and 2 people support it, the weighted P_final = +1.8dB. The negotiation results are synchronized to the cloud-based collaborative engine, updating the multi-user preference database to ensure consistency in subsequent adjustments; S4. Transmit P_final to the dynamic decision layer, with a single user's listening satisfaction ≥85%.
4. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The extreme environment compensation unit includes a biosensor calibration module, a radar signal enhancement module, and an electromagnetic interference suppression module, which respectively perform the following steps: S1. Biosensor Calibration (High Temperature > 45℃): Collect data from a temperature and humidity sensor (model SHT35, range -40℃-125℃) to obtain the ambient temperature T; Temperature compensation is applied to the raw signal PPG_raw from the PPG heart rate sensor (model MAX30102) using the following formula: ; The accuracy error of the calibrated physiological data is ≤5%; S2. Radar signal enhancement (high humidity > 85%RH scenario): For the spatial data acquired by the millimeter-wave radar (model AWR1642), an adaptive Kalman filter algorithm is used to remove water vapor interference; The attenuated 3dB gain is recovered through an antenna gain compensation algorithm, ensuring that the spatial measurement error is ≤±0.2m; S3. Electromagnetic Interference Suppression (EMC Level ≥ EN55032 Class B Scenarios): The Bluetooth module (version 5.3) uses frequency hopping communication with a hopping rate of ≥1600 times / second to avoid interference frequency bands; The audio semantic acquisition submodule employs signal shielding algorithms (such as wavelet threshold denoising) to reduce the impact of electromagnetic interference on audio parsing, achieving a semantic parsing misjudgment rate of ≤5%. S4. Transmit the calibrated sensing data to the four-dimensional sensing layer to ensure the reliability of the standardized data.
5. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The copyright-sound quality collaborative adjustment module performs the following steps: S1. Copyright Restriction Resolution: Read the copyright information from the audio file's metadata (ID3 tag) or streaming media protocol (HLS) and establish a "Copyright Level - Adjustment Permission" mapping table: Copyright Level Allowed EQ Gain Range Prohibited Sound Quality Optimization Direction Lossless (FLAC) ±6dB None 20Hz-20kHz Full-Band Equalization Enhancement Standard (MP3) ±4dB Prohibited 100-200Hz Low-Frequency Over-Enhancement (Single Gain ≤2dB) 1-3kHz Vocal Band Optimization Decryption (DRM) ±2dB Prohibited High-Frequency Modification Above 8kHz, Prohibited Oversampling Processing Maintain 200-5kHz Basic Listening Clarity S2. Dynamic permission adaptation: If the resolution is encrypted copyright, the adjustment function of the frequency band above 8kHz is locked through the hardware interface, and the EQ gain adjustment range is limited to ±2dB to avoid triggering the DRM copyright protection mechanism and causing playback interruption. If the resolution is deemed lossless, unlock full-band adjustment permissions, allowing ±6dB gain adjustment; S3. Unlocking Sound Quality Potential (For Lossless Audio Only): An oversampling algorithm is used to increase the audio sampling rate from 44.1kHz to 192kHz (achieved through the DSP chip TITMS320C5535); combined with the content characteristics output by the fusion analysis module (such as the need to enhance low frequencies in symphonic music), EQ parameters are optimized within ±6dB, achieving a sound quality potential utilization rate of ≥90%. S4. Transmit the adapted adjustment permissions and optimization parameters to the dynamic decision-making layer to ensure a balance between compliance and sound quality.
6. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The device aging adaptation module performs the following steps: S1. Aging Level Assessment: The speaker's operating current is collected using a current sensor (model ACS712), and the current fluctuation amplitude ΔI (%) is calculated as follows: (Maximum Current - Minimum Current) / Rated Current × 100%; Read the cumulative speaker usage time T (in hours) stored in the EEPROM, and evaluate the aging level L using the formula: ; Where L∈[1,5] (Level 1 is brand new, Level 5 is severely aged), for example, T=1500 hours, ΔI=18%, then L=1+(1500 / 1000)+(18 / 15)=3.2, which is rounded to Level 3; S2. Hardware parameter correction: The speaker frequency response range and distortion threshold are adjusted according to the aging level L: L=1 level: Frequency response 20Hz-20kHz, distortion threshold ≤3%; L=3 level: Frequency response 50Hz-18kHz, distortion threshold ≤4%; L=5 level: Frequency response 80Hz-16kHz, distortion threshold ≤5%; S3. Aging Compensation Adjustment: Automatically increases the gain of the corresponding frequency band to address frequency band attenuation caused by aging. L=Level 2: 50-100Hz gain +1dB; L=Level 3: 50-100Hz gain +2dB, 1-3kHz vocal band gain +0.5dB; L=4 levels: 50-100Hz gain +3dB, 1-3kHz gain +1dB; S4. Transmit the corrected hardware parameters and compensation gain to the dynamic decision layer to ensure that the adjustment parameters match the hardware capabilities after aging, with a distortion of ≤3%.
7. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The auditory quantification evaluation engine performs the following steps: S1. Objective index calculation: Speech intelligibility index (STI): The transmission index of the signal in the 125Hz-8kHz frequency band is calculated by collecting the speech signal output from the speaker (such as the standard ITU-T P.501 test signal). The range is 0-1, and ≥0.7 is judged as clear. Auditory Comfort Index (ACI): Combining the proportion of alpha waves in EEG (≤40% is comfortable) and the volume level (≤85dB is comfortable), the formula is ACI=(1-alpha wave proportion / 100)×(1-(volume-60) / 25)×100, with a range of 0-100, and ≥80 is considered comfortable; Total Harmonic Distortion (THD): Acquire a 1kHz sinusoidal test signal, calculate the power ratio of harmonic components to fundamental components, and a value ≤3% is considered acceptable. S2. Evaluation-Adjustment Closed Loop: If STI < 0.7: Automatically increase the gain of the 1-3kHz vocal band by 0.5-1dB until STI ≥ 0.7; If ACI < 80: Reduce the gain of high frequencies above 8kHz by 0.5dB, and at the same time lower the volume by 2-3dB until ACI ≥ 80; If THD > 3%: Reduce the gain of the current frequency band by 0.3-0.5dB (prioritize adjusting the low frequency 100-200Hz) until THD ≤ 3%; S3. User feedback calibration: After every 10 closed-loop adjustments, the user's subjective evaluation (clear / average / fuzzy) is obtained through an APP pop-up window. If the deviation between subjective evaluation and objective indicators is greater than 10% (e.g., STI=0.75 but user feedback is ambiguous), adjust the indicator weights—for example, increase the proportion of ACI in the evaluation, so that the deviation between quantitative evaluation and subjective feelings is ≤3%; S4. Synchronize the evaluation results and the revised indicator weights to the cloud-based collaborative engine to optimize subsequent adjustment strategies.
8. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The four-dimensional perception layer includes the following sub-modules and parameters: S1. Auditory fatigue perception sub-module: integrates an EEG brainwave sensor, sampling frequency ≥256Hz, data transmission delay ≤10ms, used to collect alpha wave signals; S2. Multi-user identity perception unit: adopts a Bluetooth BLE 5.3 module, supports simultaneous access of ≥5 user devices, identity recognition accuracy ≥99%, response time ≤50ms; S3. Device aging monitoring unit: integrates a current sensor, measurement range 0-5A, accuracy ±3%, used to monitor speaker operating current fluctuations; EEPROM storage chip (capacity ≥4KB), records the speaker's cumulative usage time (accuracy ±1 hour); S4. Copyright information acquisition sub-module: reads audio data via USB / Bluetooth interface. The video file metadata has an ID3 tag parsing latency of ≤20ms; it supports HLS / DASH streaming media protocol copyright information extraction with an accuracy of ≥98%; S5. Extreme Environment Monitoring Submodule: Temperature and humidity sensor, range -40℃-125℃ (temperature), 0%-100%RH (humidity), accuracy ±0.3℃ / ±2%RH; Electromagnetic interference sensor, measurement range -60dBm-0dBm, used to monitor electromagnetic interference intensity; S6. Environment-Content Awareness Submodule: Noise sensor, range 30-120dB, frequency response 20Hz-20kHz; Millimeter-wave radar, spatial measurement range 0.2-10m, accuracy ±0.1m; Audio semantic acquisition submodule, used to parse audio content type.
9. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The dynamic decision-making layer adopts a five-level architecture of "rule engine + reinforcement learning + physiological constraints + copyright constraints + aging constraints" and executes the following steps: S1. Constraint priority sorting: highest priority: physiological constraints (EEG alpha wave proportion > 40% triggers fatigue warning), copyright constraints (encrypted audio is prohibited from being modified above 8kHz); Second highest priority: Aging constraint (L=3 level time frequency response limit is 50Hz-18kHz); Basic priorities: rule engine (e.g., default bass +2dB in home scenarios), reinforcement learning (optimizing parameters based on historical adjustment data); S2. Adjustment instruction generation: Receive parameters output from the multi-dimensional adaptation layer (auditory memory, multi-user arbitration, etc.), first check whether they meet the highest priority constraints - for example, if the encrypted audio parameter contains 8kHz+1dB, it will be directly rejected and corrected to 8kHz±0dB; If all constraints are met, the parameters are optimized using a reinforcement learning model (based on the Q-Learning algorithm, with ≥1000 iterations), and the instruction generation frequency is ≥20Hz. If the second-highest priority constraint is not met (e.g., the parameter exceeds the frequency response range after aging), the gain of the frequency band exceeding the limit will be automatically clipped (e.g., the gain below 50Hz will be set to 0dB). S3. Response lag control: Adopting a pipelined processing architecture, the perception data acquisition, constraint verification, and instruction generation are executed in parallel, and the response lag in dynamic scenarios is ≤20ms; S4. Permission Verification Output: After the instruction is generated, it is verified again by the permission verification unit (storing the copyright-aging constraint table) to ensure that there are no illegal parameters (such as lossless audio gain ≤6dB). After the verification is passed, it is output to the feedback execution layer.
10. The audio speaker sound quality adaptive adjustment system with multi-scene switching as described in claim 1, characterized in that: The feedback execution layer includes the following sub-modules and steps: S1. Multi-user audio partition unit: Based on Bluetooth BLE packet technology, the access users are divided into different audio partitions (such as the living room sofa area and the study area). For each zone, output the corresponding preference parameters—for example, sofa zone (elderly, 100Hz+2dB) and study zone (children, 1kHz+1dB), with a zone switching response time of ≤100ms. S2. Aging Protection Module: Real-time monitoring of speaker operating power (calculated via current sensor: power = voltage × current), compared with the rated power after aging (e.g., the original 20W is reduced to 15W after aging); If the actual power is ≥90% of the rated power (e.g., 15W × 90% = 13.5W), the volume gain will be automatically limited to ≤10dB to avoid hardware overload. If the power continues for 10 seconds or more at or above the rated power, an audible and visual warning will be triggered (LED red light flashing + 1kHz warning sound). S3. Adaptive Crossover Unit: Analyzes audio content characteristics (such as vocals / instruments / sound effects) and dynamically allocates frequency bands. Human voice content (news / conference): The 1-3kHz frequency band is processed separately to improve separation; Symphony content: The 20-200Hz low frequency and 2-8kHz high frequency are divided into separate frequencies to enhance the sense of layering; After crossover, the output is through a Class D amplifier (model TITPA3116D2), improving vocal separation by 20%. S4. Adjustment effect feedback: Collect the audio signal output by the power amplifier and transmit it to the listening quantification evaluation engine to provide data support for closed-loop optimization.