Acoustic light adjusting method based on artificial intelligence
Through sensors, audio, video and depth data are collected, combined with technologies such as multi-objective knowledge distillation and grayscale release, the problem of inaccurate audio lighting adjustment in complex environments is solved, accurate understanding and intelligent adjustment of performance scenes are achieved, and response efficiency and system robustness are improved.
Patent Information
- Application Number
- CN202510535488.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the prior art to accurately grasp multi-dimensional information such as crowd emotions, spatial atmosphere and stage effects in complex public environments, resulting in insufficient sound lighting adjustment.
Audio, video and depth data are collected through sensors, multi-modal data is generated after preprocessing, and multi-objective knowledge distillation, importance pruning and low-bit quantization technology compressed models. Combined with grayscale release, real-time delay monitoring and dual-model parallel inference, the most reliable results are selected and the audio and lighting is intelligently adjusted.
It realizes an accurate understanding of the performance atmosphere, character status and spatial layout, improves the response efficiency and accuracy of audio and lighting adjustment, reduces manual operations, and improves the robustness and real-time system.
Smart Images

Figure CN120456381A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based sound and lighting adjustment method. Background Art
[0002] In modern intelligent environmental control, artificial intelligence (AI) technology is widely used to automate the adjustment of audio and lighting systems, creating a more comfortable and efficient user experience. By integrating sensors, machine learning algorithms, and user preference data, AI can intelligently adjust parameters such as audio volume and quality, as well as lighting brightness and color.
[0003] The patent publication number CN119668177A states in its specification that "the present invention provides a method and system for linking hotel central control audio and lighting scenes, which is applied to the field of central control data processing; the present invention can accurately identify user needs through the scene content preset by the hotel terminal, and determine the corresponding operation type based on this, so that the system can adjust the audio and lighting settings according to the user's specific needs, ensuring that each scene can achieve personalized equipment adjustment. At the same time, the user can transmit customized linkage data through the mobile terminal, including audio playback content, lighting color, etc. After synchronizing the content to be operated to the hotel terminal, the room equipment can be accurately controlled based on user needs, and ensure that the equipment operation meets the user's personalized needs." Although the above technology synchronizes the content to be operated of the room facilities through the mobile terminal and receives the customized linkage data sent back by the user, so as to achieve the advantage of users being able to customize the linkage effect of audio and lighting and meet the needs of diverse scenes, it is difficult to accurately grasp multi-dimensional information such as crowd emotions, spatial atmosphere, and stage effects in a complex public environment.
[0004] To sum up, developing an artificial intelligence-based audio and lighting adjustment method is still a key issue that needs to be urgently solved in the field of artificial intelligence technology. Summary of the Invention
[0005] The purpose of the present invention is to solve the problem in the prior art that although the above-mentioned technology synchronizes the pending operation content of room facilities through the mobile terminal and receives customized linkage data sent back by the user, it can achieve the advantage of users customizing the linkage effect of sound and lighting and meet the needs of diverse scenarios. However, it is difficult to accurately grasp the multi-dimensional information such as crowd emotions, spatial atmosphere, stage effects, etc. in a complex public environment.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] The present invention provides an artificial intelligence-based sound and lighting adjustment method, comprising the following steps: S1, collecting raw data through a sensor, performing preprocessing, and outputting multimodal data;
[0008] S2. compressing the full model into multiple lightweight versions with a trade-off between accuracy and volume, storing them hierarchically according to performance indicators, and then inputting the multimodal data into the lightweight model;
[0009] S3: Use grayscale release combined with real-time latency and accuracy monitoring to launch new models through gradual volume expansion and threshold determination.
[0010] S4: Use dual-model parallel reasoning and output entropy evaluation to select the most reliable results in real time. Combined with the sliding window win rate, statistics are collected for automatic switching and fallback to invalid versions.
[0011] S5. According to the statistical results of the rollback failure version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models.
[0012] Furthermore, in step S1, the raw data collected by the sensor is preprocessed and the method of outputting multimodal data is as follows:
[0013] The sensors include audio, video and depth sensors, and the raw data collected include audio data q au , video data q vi and depth data q de The raw data is processed using STFT (STFT stands for Short-Time Fourier Transform, a time-frequency analysis technique used to analyze the local frequency and phase characteristics of non-stationary signals), multi-scale convolution, and voxelization to extract audio, video, and depth features. The expression formula is:
[0014] where q au_cv (t) is the audio signal convolution output, q vi_cv (t) is the video signal convolution output, Q au (f,t) is the audio signal, q vi (t) is the video signal, Conv s is the sth convolutional layer, S is the number of convolutional layers, is the summation operation, q de_vo (i, j, k) is the depth map data after voxelization, q de (x i ,y j ,z k ) is the original depth map data, Voxelize(·) is the voxelization operation, (i, j, k) is the position index in the voxel grid, (x i ,y j ,z k) is the three-dimensional spatial coordinate in the original depth map. After the feature extraction of audio, video and depth data, weighted fusion technology is used to combine them into a unified feature vector, output multimodal data, and combine hash fingerprint technology to generate a batch ID.
[0015] Furthermore, in step S2, the full model is compressed into multiple lightweight versions with a trade-off between precision and volume, and stored hierarchically according to performance indicators. The multimodal data is then input into the lightweight model in the following manner:
[0016] The lightweight model adopts multi-objective knowledge distillation, importance pruning and low-bit quantization technology to compress the full version model. Let the full version model be W full Containing multiple levels and parameters, the knowledge learned in the full version of the multi-objective knowledge distillation model is transferred to a smaller lightweight model, and the smaller lightweight model is set to W ql , expression formula:
[0017] where Y ql , Y wz are the outputs of the lightweight model and the full version of the model, I ql , I wz is the predicted probability distribution of the lightweight model and the full version of the model, U KL represents the Kullback-Leibler divergence, E dis Represents the total loss of the model when performing knowledge distillation, is the squared error of the L2 norm, R[·] is the expected value, To balance the two losses controlled by the weight coefficient, the importance pruning is performed by deleting neurons and connections that have little impact on model performance, as expressed in the formula: Where O(P i ) is the pruning importance of the i-th layer weight, is the loss function with respect to weight P i The gradient of , ||·||1 represents the L1 norm measuring the absolute importance of the weight, and the low-bit quantization further compresses the model by reducing the number of bits for each parameter.
[0018] Furthermore, in step S2, the full model is compressed into multiple lightweight versions with a trade-off between precision and volume, and stored hierarchically according to performance indicators. The multimodal data is then input into the lightweight model in the following manner:
[0019] By combining multi-objective knowledge distillation, importance pruning, and low-bit quantization, the full model is compressed into multiple lightweight models with a precision-volume trade-off. The volume and precision of each lightweight model are weighed based on task requirements, device resources, and performance indicators, as expressed in the following formula: in represents the kth lightweight version, is the compression ratio and accuracy index set for this version. The multimodal data is input into the compressed lightweight model for inference, expressed as: Among them D out is the output of the model, X inp is the input feature of multimodal data.
[0020] Furthermore, in step S3, a grayscale release is adopted in combination with real-time delay and accuracy monitoring. Through gradual volume expansion and threshold determination, the method for launching the new model is as follows:
[0021] The grayscale release is a gradual online strategy. Initially, the new model is deployed on a small number of nodes. As the number of nodes gradually increases, it is finally fully online. The real-time delay monitoring monitors the delay of the model in real time, setting the average delay of the new model in the kth stage to Δt k , then the real-time delay monitoring model is expressed as: where N k is the number of nodes tested in the kth phase, Δt k,i is the response delay of the i-th node, Δt k is the average delay time of the kth cycle, It is from the i=1 request to the N k The "sum symbol" of the requests is displayed. At the same time as the grayscale is released, the accuracy of the new model is monitored in real time, and the accuracy of the new model in the kth stage is set to Acc k , then the accuracy monitoring formula is: where N k is the number of nodes tested in the kth phase, is the precision of the ith node, It is from the i=1 request to the N k The "sum symbol" of the request, Acc k is the average precision of the k-th node.
[0022] Furthermore, in step S3, a grayscale release is adopted in combination with real-time delay and accuracy monitoring. Through gradual volume expansion and threshold determination, the method for launching the new model is as follows:
[0023] The threshold judgment includes the triggering conditions of the fallback mechanism. If the average delay in the g stage exceeds 50ms, it will immediately fall back to the stable version. If the average accuracy in the g stage is lower than 85%, it will immediately fall back to the stable version. The expression formula is: where R k=1 indicates that the current release needs to be rolled back to the stable version, otherwise the release process will continue to be gradually promoted. The gradual release means that the new model will update the release progress under the condition that the latency and accuracy indicators of the current stage meet the expectations. The expression formula is: UP: ρ k+1 =min(ρ k +Δρ,1), where Δρ is the increase in the release ratio in each stage, ρ k+1 is the “online ratio” after the k+1th update, ρ k is the proportion of online users at the time of the kth update, min(·) is the “minimum value” of the expression in the brackets, 1 means “all online”, until it is released to all nodes (i.e. ρ k =1), the fallback mechanism in a certain stage k is triggered (i.e. R k =1), the release progress stops and rolls back to the stable version. The release progress will be reset and re-released from the previous stage. When all stages are completed and the rollback mechanism is not triggered, the final model will be fully launched, and the latency and accuracy performance data of each stage will be recorded. The expression formula is:
[0024] DeployFully:if max(Δt k )≤Δt max andmin(Acc k )≥Acc min , where DeployFully means "fully online", max(Δt k ) is the maximum response delay among all nodes, Δt max The maximum allowed response delay threshold is 50ms, min(Acc k ) is the minimum accuracy value among all nodes, Acc min The minimum allowed accuracy threshold is 85%, at which point the new model is considered successfully launched and can run safely on all nodes.
[0025] Furthermore, in step S4, the method of using dual-model parallel reasoning and output entropy evaluation to select the most reliable result in real time, and combining the sliding window win rate to calculate the automatic switching and fallback failure versions is as follows:
[0026] The dual-model parallel reasoning and output entropy evaluation are used to select the most reliable result in real time. Two models and M are used for reasoning, and multimodal data is used as input data, so that the outputs of the two models are M1 and M2 respectively. The output entropy evaluation selects the most reliable reasoning result and performs entropy evaluation on the outputs of the two models. The level of entropy indicates the confidence level of the model in the output result. The expression formula is: where p(h i ) is the probability distribution of the model output h, L is the total number of categories, G(h) represents the information entropy value of the predicted output h, Is a summation symbol, p(h i ) represents the predicted probability of the i-th category in the model prediction results, logp(h i ) is p(h i ), the lower the entropy value, the more certain the model output is, and the higher the entropy value, the greater the uncertainty of the model. According to the entropy value, the model with the lower entropy value is selected as the inference result at the current moment. The expression formula is: where h sel is the predicted value finally selected as the output result, h1 represents the predicted output of model M1, h2 represents the predicted output of model M2, and G(·) represents the entropy function.
[0027] Furthermore, in step S4, the method of using dual-model parallel reasoning and output entropy evaluation to select the most reliable result in real time, and combining the sliding window win rate to calculate the automatic switching and fallback failure versions is as follows:
[0028] The winning rates of the sliding windows are counted separately. In a sliding window with a window size of Z, the performance of the two models is evaluated by their winning rates. Each time the window slides, the winning rates of the two models in the window are counted. The expression formula is: in is the predicted value of the i-th inference model M1 and M2, V(·) is the indicator function, H 1,k It represents the accuracy of model M1 in the latest Z samples in the sliding window at time k, where Z represents the width of the sliding window. It is the sum of the inference results of all these moments from the k-Z+1th moment to the current kth moment. is the predicted output of model M1 in the i-th inference, h (i) is the true label of the i-th sample, is the predicted output of model M2 in the i-th inference, H 2,k It represents the accuracy of model M2 in the latest Z samples within the sliding window at time k. It is 1 when the prediction is correct, otherwise it is 0. The long-term performance of the two models is judged based on the winning rate of the sliding window statistics. The model with a higher winning rate will be given priority. According to the real-time winning rate statistics, a threshold B is set. When it is detected that the winning rate of any model of the two models is continuously lower than the threshold B, the automatic switching mechanism is triggered to switch out the invalid model and roll back. After several consecutive rollbacks, it is determined that the model is no longer applicable.
[0029] Furthermore, in step S5, based on the statistical results of the rollback failure version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models:
[0030] The incremental retraining is triggered after multiple backoffs, and the predicted value of each inference is recorded at each backoff. and the true label c (i) , and regard it as a set of failed samples After the number of rollbacks exceeds the threshold B, incremental retraining is triggered, and the incremental retraining sets the current version to The incremental retraining uses a set of failed samples Update model parameters and express the formula: Where Γ is the loss function, incr represents the total loss value of incremental training, represents the sum of all failed samples i′, Represents the predicted value of sample i and the true label h (i) The incremental retraining introduces a distribution alignment discriminator, which forces the newly generated samples to be aligned with the historical failure samples in the feature space through the Wasserstein distance constraint, setting the two distributions and The Wasserstein distance is defined as: in It's all from and The joint distribution of samples in , ||x′-y′|| is the distance between samples x′ and y′ (usually using Euclidean distance), Represents the new sample distribution and the distribution of failed samples The Wasserstein distance between Indicates that in all joint distributions α (satisfying the marginal distribution is and ) minimizes the expected value, represents sampling a pair x′, y′ under the joint distribution α.
[0031] Furthermore, in step S5, based on the statistical results of the rollback failure version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models:
[0032] The incremental retraining introduces entropy-weighted sample selection, and the entropy of the prediction is calculated for each sample i. When the entropy value is high, the sample has greater potential to improve the model, so it is given a higher sampling weight. The weighting coefficient of the sample is The expression formula is: in The sum of the predicted entropy values of all n samples is used as the denominator of normalization, is the weighting coefficient of sample i, is the predicted entropy value of the i-th sample, and the weighted loss function is: where Γ wei represents the weighted loss function, Indicates the sum of all failed samples. Represents the model's predicted value for the i-th sample and the true label h (i) The incremental retraining eliminates the error between versions through version entropy decay, regularly evaluates the inference entropy of each candidate model, calculates the average inference entropy of all candidate models, and eliminates the 50% versions with the fastest entropy growth, and regularly deletes old candidate models, keeping only Time interval candidate models as the next generation candidate models.
[0033] Beneficial effects
[0034] Compared with the known public technology, the technical solution provided by the present invention has the following advantages:
[0035] Beneficial effects:
[0036] When in use, the present invention achieves an accurate understanding of the atmosphere of the performance, the status of the characters, and the spatial layout through the joint modeling of multi-source data of audio, video, and depth, which facilitates the accurate capture of multi-dimensional information of crowd emotions, spatial atmosphere, and stage effects in complex public environments, and then intelligently adjusts the sound and lighting based on artificial intelligence, which is conducive to improving response efficiency and accuracy, reducing manual operations, and greatly improving the robustness and real-time performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 The present invention is a flowchart of an artificial intelligence-based audio and lighting adjustment method. DETAILED DESCRIPTION
[0038] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0039] It should be noted that the terms "first," "second," and the like in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.
[0040] The present invention is described in further detail below with reference to the accompanying drawings:
[0041] Example:
[0042] like Figure 1 As shown, the present invention provides an artificial intelligence-based sound and lighting adjustment method, comprising the following steps: S1, collecting raw data through a sensor, preprocessing it, and outputting multimodal data;
[0043] Furthermore, in step S1, the raw data collected by the sensor is preprocessed and the method of outputting multimodal data is as follows:
[0044] The sensors include audio, video and depth sensors, and the raw data collected include audio data q au , video data q vi and depth data q de The raw data is processed using STFT (STFT stands for Short-Time Fourier Transform, a time-frequency analysis technique used to analyze the local frequency and phase characteristics of non-stationary signals), multi-scale convolution, and voxelization to extract audio, video, and depth features. The expression formula is:
[0045] where q au_cv (t) is the audio signal convolution output, q vi_cv (t) is the video signal convolution output, Q au (f,t) is the audio signal, q vi (t) is the video signal, Conv s is the sth convolutional layer, S is the number of convolutional layers, is the summation operation, q de_vo (i, j, k) is the depth map data after voxelization, q de (x i ,y j ,z k) is the original depth map data, Voxelize(·) is the voxelization operation, (i, j, k) is the position index in the voxel grid, (x i ,y j ,z k ) is the three-dimensional spatial coordinate in the original depth map. After the feature extraction of audio, video and depth data, weighted fusion technology is used to combine them into a unified feature vector, output multimodal data, and combine hash fingerprint technology to generate a batch ID.
[0046] In this embodiment, the method extracts multidimensional information from the time domain, frequency domain, and spatial structure through STFT, multi-scale convolution, and voxelization operations to improve perception accuracy. The weighted fusion strategy fully utilizes the complementarity of each modal data to enhance the method's ability to understand complex scenes. The hash algorithm is then used to generate a unique ID for each batch of multimodal data, so that each training and inference is based on clean and traceable input.
[0047] S2. compressing the full model into multiple lightweight versions with a trade-off between accuracy and volume, storing them hierarchically according to performance indicators, and then inputting the multimodal data into the lightweight model;
[0048] Furthermore, in step S2, the full model is compressed into multiple lightweight versions with a trade-off between precision and volume, and stored hierarchically according to performance indicators. The multimodal data is then input into the lightweight model in the following manner:
[0049] The lightweight model adopts multi-objective knowledge distillation, importance pruning and low-bit quantization technology to compress the full version model. Let the full version model be W full Containing multiple levels and parameters, the knowledge learned in the full version of the multi-objective knowledge distillation model is transferred to a smaller lightweight model, and the smaller lightweight model is set to W ql , expression formula:
[0050] where Y ql , Y wz are the outputs of the lightweight model and the full version of the model, I ql , I wz is the predicted probability distribution of the lightweight model and the full version of the model, U KL represents the Kullback-Leibler divergence, E dis Represents the total loss of the model when performing knowledge distillation, is the squared error of the L2 norm, R[·] is the expected value, To balance the two losses controlled by the weight coefficient, the importance pruning is performed by deleting neurons and connections that have little impact on model performance, as expressed in the formula: Where O(P i) is the pruning importance of the i-th layer weight, is the loss function with respect to weight P i The gradient of , ||·||1 represents the L1 norm measuring the absolute importance of the weight, and the low-bit quantization further compresses the model by reducing the number of bits for each parameter.
[0051] Furthermore, in step S2, the full model is compressed into multiple lightweight versions with a trade-off between precision and volume, and stored hierarchically according to performance indicators. The multimodal data is then input into the lightweight model in the following manner:
[0052] By combining multi-objective knowledge distillation, importance pruning, and low-bit quantization, the full model is compressed into multiple lightweight models with a precision-volume trade-off. The volume and precision of each lightweight model are weighed based on task requirements, device resources, and performance indicators, as expressed in the following formula: in represents the kth lightweight version, is the compression ratio and accuracy index set for this version. The multimodal data is input into the compressed lightweight model for inference, expressed as: Among them D out is the output of the model, X inp is the input feature of multimodal data.
[0053] In this embodiment, this method combines knowledge distillation with importance pruning, supplemented by low-bit quantization, to refine multiple lightweight models that balance accuracy and volume and store them in hierarchical levels by performance, thereby meeting the flexible deployment requirements from high-end GPU servers to resource-constrained edge devices, which is conducive to significantly reducing model deployment costs and dynamically matching the optimal model according to the scenario.
[0054] S3: Use grayscale release combined with real-time latency and accuracy monitoring to launch new models through gradual volume expansion and threshold determination.
[0055] Furthermore, in step S3, a grayscale release is adopted in combination with real-time delay and accuracy monitoring. Through gradual volume expansion and threshold determination, the method for launching the new model is as follows:
[0056] The grayscale release is a gradual online strategy. Initially, the new model is deployed on a small number of nodes. As the number of nodes gradually increases, it is finally fully online. The real-time delay monitoring monitors the delay of the model in real time, setting the average delay of the new model in the kth stage to Δt k , then the real-time delay monitoring model is expressed as: where N k is the number of nodes tested in the kth phase, Δt k,i is the response delay of the i-th node, Δt kis the average delay time of the kth cycle, It is from the i=1 request to the N k The "sum symbol" of the requests is displayed. At the same time as the grayscale is released, the accuracy of the new model is monitored in real time, and the accuracy of the new model in the kth stage is set to Acc k , then the accuracy monitoring formula is: where N k is the number of nodes tested in the kth phase, is the precision of the ith node, It is from the i=1 request to the N k The "sum symbol" of the request, Acc k is the average precision of the k-th node.
[0057] Furthermore, in step S3, a grayscale release is adopted in combination with real-time delay and accuracy monitoring. Through gradual volume expansion and threshold determination, the method for launching the new model is as follows:
[0058] The threshold judgment includes the triggering conditions of the fallback mechanism. If the average delay in the g stage exceeds 50ms, it will immediately fall back to the stable version. If the average accuracy in the g stage is lower than 85%, it will immediately fall back to the stable version. The expression formula is: where R k =1 indicates that the current release needs to be rolled back to the stable version, otherwise the release process will continue to be gradually promoted. The gradual release means that the new model will update the release progress under the condition that the latency and accuracy indicators of the current stage meet the expectations. The expression formula is: UP: ρ k+1 =min(ρ k +Δρ,1), where Δρ is the increase in the release ratio in each stage, ρ k+1 is the “online ratio” after the k+1th update, ρ k is the proportion of online users at the time of the kth update, min(·) is the “minimum value” of the expression in the brackets, 1 means “all online”, until it is released to all nodes (i.e. ρ k =1), the fallback mechanism in a certain stage k is triggered (i.e. R k =1), the release progress stops and rolls back to the stable version. The release progress will be reset and re-released from the previous stage. When all stages are completed and the rollback mechanism is not triggered, the final model will be fully launched, and the latency and accuracy performance data of each stage will be recorded. The expression formula is:
[0059] DeployFully:if max(Δt k )≤Δt max andmin(Acc k )≥Acc min, where DeployFully means "fully online", max(Δt k ) is the maximum response delay among all nodes, Δt max The maximum allowed response delay threshold is 50ms, min(Acc k ) is the minimum accuracy value among all nodes, Acc min The minimum allowed accuracy threshold is 85%, at which point the new model is considered successfully launched and can run safely on all nodes.
[0060] In this embodiment, when going online, this method relies on the grayscale expansion strategy, combined with a 50ms delay threshold and an 85% accuracy threshold, to first verify the effect of the new model on a small number of nodes. Once performance degradation is detected, it will immediately and seamlessly roll back to the stable version to avoid the spread of potential risks. Finally, in the online inference stage, the stable version and the candidate version of the model run in parallel, and the most reliable result is dynamically selected by outputting entropy difference and sliding win rate. At the same time, all rollback events and abnormal samples are automatically fed back to the preprocessing and offline training links, triggering data re-sampling and incremental retraining, forming an end-to-end adaptive closed loop, and continuously optimizing multimodal perception and scene understanding capabilities.
[0061] S4: Use dual-model parallel reasoning and output entropy evaluation to select the most reliable results in real time. Combined with the sliding window win rate, statistics are collected for automatic switching and fallback to invalid versions.
[0062] Furthermore, in step S4, the method of using dual-model parallel reasoning and output entropy evaluation to select the most reliable result in real time, and combining the sliding window win rate to calculate the automatic switching and fallback failure versions is as follows:
[0063] The dual-model parallel reasoning and output entropy evaluation are used to select the most reliable result in real time. Two models and M are used for reasoning, and multimodal data is used as input data, so that the outputs of the two models are M1 and M2 respectively. The output entropy evaluation selects the most reliable reasoning result and performs entropy evaluation on the outputs of the two models. The level of entropy indicates the confidence level of the model in the output result. The expression formula is: where p(h i ) is the probability distribution of the model output h, L is the total number of categories, G(h) represents the information entropy value of the predicted output h, Is a summation symbol, p(h i ) represents the predicted probability of the i-th category in the model prediction results, logp(h i ) is p(h i ), the lower the entropy value, the more certain the model output is, and the higher the entropy value, the greater the uncertainty of the model. According to the entropy value, the model with the lower entropy value is selected as the inference result at the current moment. The expression formula is: where h selis the predicted value finally selected as the output result, h1 represents the predicted output of model M1, h2 represents the predicted output of model M2, and G(·) represents the entropy function.
[0064] Furthermore, in step S4, the method of using dual-model parallel reasoning and output entropy evaluation to select the most reliable result in real time, and combining the sliding window win rate to calculate the automatic switching and fallback failure versions is as follows:
[0065] The winning rates of the sliding windows are counted separately. In a sliding window with a window size of Z, the performance of the two models is evaluated by their winning rates. Each time the window slides, the winning rates of the two models in the window are counted. The expression formula is: in is the predicted value of the i-th inference model M1 and M2, V(·) is the indicator function, H 1,k It represents the accuracy of model M1 in the latest Z samples in the sliding window at time k, where Z represents the width of the sliding window. It is the sum of the inference results of all these moments from the k-Z+1th moment to the current kth moment. is the predicted output of model M1 in the i-th inference, h (i) is the true label of the i-th sample, is the predicted output of model M2 in the i-th inference, H 2,k It represents the accuracy of model M2 in the latest Z samples within the sliding window at time k. It is 1 when the prediction is correct, otherwise it is 0. The long-term performance of the two models is judged based on the winning rate of the sliding window statistics. The model with a higher winning rate will be given priority. According to the real-time winning rate statistics, a threshold B is set. When it is detected that the winning rate of any model of the two models is continuously lower than the threshold B, the automatic switching mechanism is triggered to switch out the invalid model and roll back. After several consecutive rollbacks, it is determined that the model is no longer applicable.
[0066] In this embodiment, the method evaluates the output quality of the two models in real time, so that each inference result comes from a "more credible" model, which is convenient for enhancing the stability and reliability of inference. Even if the new model has performance fluctuations, the old model can be run in parallel as "insurance" to fall back at any time, reducing deployment risks and improving the security of new model deployment. When the scene suddenly changes, such as light changes or noise interference, causing the performance of one of the models to degrade, the method automatically detects and adjusts the usage strategy, which is conducive to dynamic adaptation to changes in different scenes.
[0067] S5. Based on the statistical results of the failed rollback version, trigger incremental retraining after multiple rollbacks to generate a new generation of candidate models;
[0068] Furthermore, in step S5, based on the statistical results of the rollback failure version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models:
[0069] The incremental retraining is triggered after multiple backoffs, and the predicted value of each inference is recorded at each backoff. and the true label c (i) , and regard it as a set of failed samples After the number of rollbacks exceeds the threshold B, incremental retraining is triggered, and the incremental retraining sets the current version to The incremental retraining uses a set of failed samples Update model parameters and express the formula: Where Γ is the loss function, incr represents the total loss value of incremental training, represents the sum of all failed samples i′, Represents the predicted value of sample i and the true label h (i) The incremental retraining introduces a distribution alignment discriminator, which forces the newly generated samples to be aligned with the historical failure samples in the feature space through the Wasserstein distance constraint, setting the two distributions and The Wasserstein distance is defined as: in It's all from and The joint distribution of samples in , ||x′-y′|| is the distance between samples x′ and y′ (usually using Euclidean distance), Represents the new sample distribution and the distribution of failed samples The Wasserstein distance between Indicates that in all joint distributions α (satisfying the marginal distribution is and ) minimizes the expected value, represents sampling a pair x′, y′ under the joint distribution α.
[0070] Furthermore, in step S5, based on the statistical results of the rollback failure version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models:
[0071] The incremental retraining introduces entropy-weighted sample selection, and the entropy of the prediction is calculated for each sample i. When the entropy value is high, the sample has greater potential to improve the model, so it is given a higher sampling weight. The weighting coefficient of the sample is The expression formula is: in The sum of the predicted entropy values of all n samples is used as the denominator of normalization, is the weighting coefficient of sample i, is the predicted entropy value of the i-th sample, and the weighted loss function is: where Γ wei represents the weighted loss function, Indicates the sum of all failed samples. Represents the model's predicted value for the i-th sample and the true label h (i) The incremental retraining eliminates the error between versions through version entropy decay, regularly evaluates the inference entropy of each candidate model, calculates the average inference entropy of all candidate models, and eliminates the 50% versions with the fastest entropy growth, and regularly deletes old candidate models, keeping only Time interval candidate models as the next generation candidate models.
[0072] In this embodiment, this method uses failure cases exposed in actual operation as "reverse teaching materials" to avoid the recurrence of the same problem. The introduction of distribution alignment and entropy weighting mechanisms is conducive to improving the robustness of the model to edge samples and complex scenarios. At the same time, the entire process is automatically triggered, trained and evaluated by the system, without the need for manual labeling intervention, which is conducive to reducing manual operation and maintenance costs. Through the entropy trend elimination mechanism, inferior candidate models generated by training are avoided from being mixed into the official version, thereby ensuring model quality. Incremental retraining is faster and more efficient than full training, and can respond to on-site changes more quickly.
[0073] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An artificial intelligence-based sound and lighting adjustment method, characterized in that: The following steps are involved: S1, collects raw data through sensors, performs preprocessing, and outputs multimodal data; S2. compressing the full model into multiple lightweight versions with a trade-off between accuracy and volume, storing them hierarchically according to performance indicators, and then inputting the multimodal data into the lightweight model; S3: Use grayscale release combined with real-time latency and accuracy monitoring to launch new models through gradual volume expansion and threshold determination. S4: Use dual-model parallel reasoning and output entropy evaluation to select the most reliable results in real time. Combined with the sliding window win rate, statistics are collected for automatic switching and fallback to invalid versions. S5. According to the statistical results of the failed rollback version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models.
2. The method for adjusting sound and lighting based on artificial intelligence according to claim 1, characterized in that: In step S1, the raw data collected by the sensor is preprocessed and the method of outputting multimodal data is as follows: The sensors include audio, video and depth sensors, and the raw data collected include audio data q au , video data q vi and depth data q de The raw data is processed using STFT (STFT stands for Short-Time Fourier Transform, a time-frequency analysis technique used to analyze the local frequency and phase characteristics of non-stationary signals), multi-scale convolution, and voxelization to extract audio, video, and depth features. The expression formula is: where q au_cv (t) is the audio signal convolution output, q vi_cv (t) is the video signal convolution output, Q au (f,t) is the audio signal, q vi (t) is the video signal, Conv s is the sth convolutional layer, S is the number of convolutional layers, is the summation operation, q de_vo (i, j, k) is the depth map data after voxelization, q de (x i ,y j ,z k ) is the original depth map data, Voxelize(·) is the voxelization operation, (i, j, k) is the position index in the voxel grid, (x i ,y j ,z k ) is the three-dimensional spatial coordinate in the original depth map. After the feature extraction of audio, video and depth data, weighted fusion technology is used to combine them into a unified feature vector, output multimodal data, and combine hash fingerprint technology to generate a batch ID.
3. The method for adjusting sound and lighting based on artificial intelligence according to claim 2, characterized in that: In step S2, the full model is compressed into multiple lightweight versions with a trade-off between precision and volume, and stored hierarchically according to performance indicators. The multimodal data is then input into the lightweight model as follows: The lightweight model adopts multi-objective knowledge distillation, importance pruning and low-bit quantization technology to compress the full version model. Let the full version model be W full Containing multiple levels and parameters, the knowledge learned in the full version of the multi-objective knowledge distillation model is transferred to a smaller lightweight model, and the smaller lightweight model is set to W ql , expression formula: where Y ql , Y wz are the outputs of the lightweight model and the full version of the model, I ql , I wz is the predicted probability distribution of the lightweight model and the full version of the model, U KL represents the Kullback-Leibler divergence, E dis Represents the total loss of the model when performing knowledge distillation, is the squared error of the L2 norm, R[·] is the expected value, To balance the two losses controlled by the weight coefficient, the importance pruning is performed by deleting neurons and connections that have little impact on model performance, as expressed in the formula: Where O(P i ) is the pruning importance of the i-th layer weight, is the loss function with respect to weight P i The gradient of , ||·||1 represents the L1 norm measuring the absolute importance of the weight, and the low-bit quantization further compresses the model by reducing the number of bits for each parameter.
4. The method for adjusting sound and lighting based on artificial intelligence according to claim 3, characterized in that: In step S2, the full model is compressed into multiple lightweight versions with a trade-off between precision and volume, and stored hierarchically according to performance indicators. The multimodal data is then input into the lightweight model as follows: By combining multi-objective knowledge distillation, importance pruning, and low-bit quantization, the full model is compressed into multiple lightweight models with a precision-volume trade-off. The volume and precision of each lightweight model are weighed based on task requirements, device resources, and performance indicators, as expressed in the following formula: in represents the kth lightweight version, is the compression ratio and accuracy index set for this version. The multimodal data is input into the compressed lightweight model for inference, expressed as: Among them D out is the output of the model, X inp is the input feature of multimodal data.
5. The method for adjusting sound and lighting based on artificial intelligence according to claim 4, characterized in that: In step S3, a grayscale release is used in combination with real-time delay and accuracy monitoring. The new model is launched through gradual volume increase and threshold determination. The grayscale release is a gradual online strategy. Initially, the new model is deployed on a small number of nodes. As the number of nodes gradually increases, it is finally fully online. The real-time delay monitoring monitors the delay of the model in real time, setting the average delay of the new model in the kth stage to Δt k , then the real-time delay monitoring model is expressed as: where N k is the number of nodes tested in the kth phase, Δt k,i is the response delay of the i-th node, Δt k is the average delay time of the kth cycle, It is from the i=1 request to the N k The "sum symbol" of the requests is displayed. At the same time as the grayscale is released, the accuracy of the new model is monitored in real time, and the accuracy of the new model in the kth stage is set to Acc k , then the accuracy monitoring formula is: where N k is the number of nodes tested in the kth phase, is the precision of the ith node, It is from the i=1 request to the N k The "sum symbol" of the request, Acc k is the average precision of the k-th node.
6. The method for adjusting sound and lighting based on artificial intelligence according to claim 5, characterized in that: In step S3, a grayscale release is used in combination with real-time delay and accuracy monitoring. The new model is launched through gradual volume increase and threshold determination. The threshold judgment includes the triggering conditions of the fallback mechanism. If the average delay in the g stage exceeds 50ms, it will immediately fall back to the stable version. If the average accuracy in the g stage is lower than 85%, it will immediately fall back to the stable version. The expression formula is: where R k =1 indicates that the current release needs to be rolled back to the stable version, otherwise the release process will continue to be gradually promoted. The gradual release means that the new model will update the release progress under the condition that the latency and accuracy indicators of the current stage meet the expectations. The expression formula is: UP: ρ k+1 =min(ρ k +Δρ,1), where Δρ is the increase in the release ratio in each stage, ρ k+1 is the "online ratio" after the k+1th update, ρ k is the proportion of online users at the time of the kth update, min(·) is the "minimum value" of the expression in the brackets, 1 means "full online", until it is released to all nodes (i.e. ρ k =1), the fallback mechanism in a certain stage k is triggered (i.e. R k =1), the release progress stops and rolls back to the stable version. The release progress will be reset and re-released from the previous stage. When all stages are completed and the rollback mechanism is not triggered, the final model will be fully launched, and the latency and accuracy performance data of each stage will be recorded. The expression formula is: DeployFully:if max(Δt k )≤Δt max andmin(Acc k )≥Acc min , where DeployFully means "fully online", max(Δt k ) is the maximum response delay among all nodes, Δt max The maximum response delay threshold allowed is 50ms, min(Acc k ) is the minimum accuracy value among all nodes, Acc min The minimum allowed accuracy threshold is 85%, at which point the new model is considered successfully launched and can run safely on all nodes.
7. The method for adjusting sound and lighting based on artificial intelligence according to claim 6, characterized in that: In step S4, the method of using dual-model parallel reasoning and output entropy evaluation to select the most reliable result in real time, and combining the sliding window win rate to calculate the automatic switching and fallback failure versions is as follows: The dual-model parallel reasoning and output entropy evaluation are used to select the most reliable result in real time. Two models and M are used for reasoning, and multimodal data is used as input data, so that the outputs of the two models are M1 and M2 respectively. The output entropy evaluation selects the most reliable reasoning result and performs entropy evaluation on the outputs of the two models. The level of entropy indicates the confidence level of the model in the output result. The expression formula is: where p(h i ) is the probability distribution of the model output h, L is the total number of categories, G(h) represents the information entropy value of the predicted output h, Is a summation symbol, p(h i ) represents the predicted probability of the i-th category in the model prediction results, logp(h i ) is p(h i ), the lower the entropy value, the more certain the model output is, and the higher the entropy value, the greater the uncertainty of the model. According to the entropy value, the model with the lower entropy value is selected as the inference result at the current moment. The expression formula is: where h sel is the predicted value finally selected as the output result, h1 represents the predicted output of model M1, h2 represents the predicted output of model M2, and G(·) represents the entropy function.
8. The method for adjusting sound and lighting based on artificial intelligence according to claim 7, characterized in that: In step S4, the method of using dual-model parallel reasoning and output entropy evaluation to select the most reliable result in real time, and combining the sliding window win rate to calculate the automatic switching and fallback failure versions is as follows: The winning rates of the sliding windows are counted separately. In a sliding window with a window size of Z, the performance of the two models is evaluated by their winning rates. Each time the window slides, the winning rates of the two models in the window are counted. The expression formula is: in is the predicted value of the i-th inference model M1 and M2, V(·) is the indicator function, H 1,k It represents the accuracy of model M1 in the latest Z samples in the sliding window at time k, where Z represents the width of the sliding window. It is the sum of the inference results of all these moments from the k-Z+1th moment to the current kth moment. is the predicted output of model M1 in the i-th inference, h (i) is the true label of the i-th sample, is the predicted output of model M2 in the i-th inference, H 2,k It represents the accuracy of model M2 in the latest Z samples within the sliding window at time k. It is 1 when the prediction is correct, otherwise it is 0. The long-term performance of the two models is judged based on the winning rate of the sliding window statistics. The model with a higher winning rate will be given priority. According to the real-time winning rate statistics, a threshold B is set. When it is detected that the winning rate of any model of the two models is continuously lower than the threshold B, the automatic switching mechanism is triggered to switch out the invalid model and roll back. After several consecutive rollbacks, it is determined that the model is no longer applicable.
9. The method for adjusting sound and lighting based on artificial intelligence according to claim 8, characterized in that: In step S5, based on the statistical results of the failed rollback version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models: The incremental retraining is triggered after multiple backoffs, and the predicted value of each inference is recorded at each backoff. and the true label c (i) , and regard it as a set of failed samples After the number of rollbacks exceeds the threshold B, incremental retraining is triggered, and the incremental retraining sets the current version to The incremental retraining uses a set of failed samples Update model parameters and express the formula: Where Γ is the loss function, incr represents the total loss value of incremental training, represents the sum of all failed samples i′, Represents the predicted value of sample i and the true label h (i) The incremental retraining introduces a distribution alignment discriminator, which forces the newly generated samples to be aligned with the historical failure samples in the feature space through the Wasserstein distance constraint, setting the two distributions and The Wasserstein distance is defined as: in It's all from and The joint distribution of samples in , ||x′-y′|| is the distance between samples x′ and y′ (usually using Euclidean distance), Represents the new sample distribution and the distribution of failed samples The Wasserstein distance between Indicates that in all joint distributions α (satisfying the marginal distribution is and ) minimizes the expected value, represents sampling a pair x′, y′ under the joint distribution α.
10. The method for adjusting sound and lighting based on artificial intelligence according to claim 8, characterized in that: In step S5, based on the statistical results of the failed rollback version, incremental retraining is triggered after multiple rollbacks to generate a new generation of candidate models: The incremental retraining introduces entropy-weighted sample selection, and the entropy of the prediction is calculated for each sample i. When the entropy value is high, the sample has greater potential to improve the model, so it is given a higher sampling weight. The weighting coefficient of the sample is The expression formula is: in The sum of the predicted entropy values of all n samples is used as the denominator of normalization, is the weighting coefficient of sample i, is the predicted entropy value of the i-th sample, and the weighted loss function is: where Γ wei represents the weighted loss function, Indicates the sum of all failed samples. Represents the model's predicted value for the i-th sample and the true label h (i) The incremental retraining eliminates the error between versions through version entropy decay, regularly evaluates the inference entropy of each candidate model, calculates the average inference entropy of all candidate models, and eliminates the 50% versions with the fastest entropy growth, and regularly deletes old candidate models, keeping only Time interval candidate models as the next generation candidate models.
Citation Information
Patent Citations
Linkage method and system of hotel central control sound equipment and light scene
CN119668177A
Cited By
Lamplight control method and system for sound equipment based on multi-modal data
CN121397808A