An Adaptive Audio Data Processing System and Method for Audio Devices

By analyzing video and audio data, intelligently adjusting the audio volume, the problem of user manual adjustment of volume is solved, and the user experience and family entertainment quality is improved.

CN119645340BActive Publication Date: 2025-07-22VEST AUDIO TECH (GUANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411782266.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-07-22
Estimated Expiration
2044-12-05

AI Technical Summary

Technical Problem

Existing smart speakers need to manually adjust the volume when a user calls or needs to talk to avoid affecting the call and causing a poor user experience.

Method used

By obtaining the video data of the monitoring device and the audio data of the audio device, analyzing the feature area and target local area, intercepting the pre-segment segment, obtaining similarity values and environmental spectrograms, adjusting the volume model in combination with historical data, and intelligently adjusting the audio volume.

Benefits of technology

Automatically adjust the audio volume to avoid excessive sound affecting normal calls and improve user experience and family entertainment quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119645340B_ABST
    Figure CN119645340B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive audio data processing system and method for an audio device, which relates to the technical field of device data processing and includes: obtaining historical video data of a monitoring device and historical audio data of an audio device to obtain a feature region and a target local region, and intercepting a preamble segment in the video data; obtaining the video data at the current moment to obtain a similarity value corresponding to the current video data; collecting the audio data at the current moment to obtain an environmental spectrogram; determining a volume adjustment model for the device volume, and intelligently adjusting the current volume of the audio device in combination with the similarity value. By using the historical video data and audio data, the present invention obtains the device volume that should be set at the current moment, avoids the excessive sound of the speaker affecting normal calls, brings convenience and comfort to users, makes the household audio device more in line with the actual needs of users, and improves the quality of home entertainment and life.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of device data processing, and specifically to an adaptive audio data processing system and method for audio devices. Background Art

[0002] With the progress of technology, smart homes can achieve more intelligent decision-making and automated control, bringing significant development uses. They not only greatly enrich people's entertainment methods, bringing endless happy times to families, but also play an important role in creating a warm and harmonious family atmosphere; as an indispensable audio device for home use, speakers play a crucial role in people's daily lives. Smart speakers are generally equipped with microphones and can listen to users' instructions at any time. However, in actual use, it often happens that if the user receives a call or needs to talk, the speaker volume needs to be manually adjusted so as not to affect the conversation, which is not only troublesome but also the frequent adjustment of the volume affects the user experience. Therefore, according to the audio data in the environment and the home surveillance camera, intelligently adjusting the volume of the speaker device to avoid the speaker sound being too loud and affecting normal calls can bring convenience and comfort to users, making the home audio device more in line with the actual needs of users, and further improving the quality of home entertainment and life. Summary of the Invention

[0003] The purpose of the present invention is to provide an adaptive audio data processing system and method for audio devices to solve the problems raised in the prior art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] An adaptive audio data processing method for audio devices, comprising the following steps:

[0006] Step S100: Obtain the historical video data of the monitoring device and the historical audio data of the audio device, and according to the amplitude change in the audio data, obtain the feature area and the target local area in the monitoring screen; establish a plane coordinate system for the monitoring screen, and intercept the preamble segment in the video data according to the coordinate position change of the feature area and the target local area;

[0007] Step S200: Extract the feature actions in the preamble segment, obtain the video data at the current moment, compare and analyze the current video data with the feature actions in the preamble segment, and obtain the similarity value corresponding to the current video data;

[0008] Step S300: Collect the audio data at the current moment, and convert both the audio input information and the audio output information in the audio data into spectrograms, and then obtain the environmental spectrogram;

[0009] A spectrogram is a tool that visualizes the variation of the frequency and amplitude of a sound signal over time, with frequency on the x-axis and amplitude on the y-axis. There is a corresponding spectrogram for each moment. In the spectrogram of the audio output information, there is only the audio information of the speaker, while in the spectrogram of the audio input information, it also includes the audio information of the human voice. Therefore, by combining and subtracting the two, the ambient spectrogram can be obtained;

[0010] Step S400: Pre-acquire the historical ambient spectrogram of the audio device and the change in the device volume, determine the volume adjustment model of the device volume, and intelligently adjust the current volume of the audio device in combination with the similarity value.

[0011] Further, step S100 includes:

[0012] Step S110: Obtain the audio data of the audio device in the recent M days. The audio data includes each moment, as well as the audio input information, audio output information, and device volume at each moment; obtain the amplitude of the audio output information at each moment. Take the device volume at a certain moment T0 as V0. If the device volume within a time period S1 before a certain moment T0 is all V0, and the audio volume V1 at the next moment T1 of a certain moment T0 satisfies V0 > V1, and the audio volumes at subsequent moments decrease in turn until when a certain moment T n , and a certain moment T n , the next moment T n+1 , the corresponding device volume is the same, obtain the average amplitude F1 within the time period S1, and the average amplitude F2 from a certain moment T0 to a certain moment T n ;

[0013] Step S120: If F2 < F1, obtain the monitoring screen of the monitoring device at a certain moment T0, establish a plane coordinate system for the monitoring screen, and through a neural network algorithm, obtain the characteristic area of the characteristic item in the monitoring screen and the local parts of the moving object, and the corresponding local areas in the monitoring screen; if the intersection area of the characteristic area and a certain local area in the monitoring screen is greater than 0, then take a certain local area as the target local area;

[0014] In this solution, the audio device is a smart speaker. The audio input information is the audio data received by the microphone of the smart speaker, the audio output information is the audio data output by the speaker, and the device volume is the volume of the smart speaker. The reason for judging F2 < F1 here is as follows: In some cases, turning down the volume of the speaker may not only be because someone is speaking, but also because the sound played by the speaker suddenly increases, and the reason for the increase is generally related to the suddenly switched song or some suddenly appearing background music. Therefore, it is necessary to make a judgment to ensure that it is indeed caused by external factors. Both the audio device and the monitoring device are deployed in the living room. The characteristic item is the remote control for controlling the volume of the audio device, the moving object is a person, and the local parts include the hand and the mouth. In this step, it is the hand, specifically: If people pick up the remote control to control the volume of the speaker and then the volume of the speaker becomes smaller, then the hand of people picking up the remote control is used as the target local area; generally, only when situations such as conversation or making a phone call occur will the volume of the speaker be adjusted;

[0015] Step S130: Obtain the central coordinate P of the target local area at a certain moment T in the time period S1, obtain the central coordinates of the target local area at each moment within the previous time period S2 starting from the moment T, and calculate the average coordinate P^, and take the central coordinate of the characteristic area at the moment T as P T ; Taking the coordinate P as the starting point and P T as the end point, obtain the vector V0. Take the time period from the moment T until the moment when the intersection area between the target local area and the characteristic area is not 0 as S3. Taking the central coordinates of the target local area at each moment in the time period S3 as the starting point and P T as the end point, obtain all vectors;

[0016] Step S140: If 1 / H < L1 / L2 < H and the angle between each vector and the vector V0 is less than the angle threshold, where H is the distance coefficient, L1 is the distance between the coordinate P T and the coordinate P, and L2 is the distance between the coordinate P T and the coordinate P^, then take the time period S2 as the preamble segment, and thus obtain all the preamble segments in the video data.

[0017] A certain moment T0 is the moment when people are just about to switch the volume of the speaker, the moment T is the moment when people are just about to approach the remote control, the time period S2 is the time period before the moment T when people are answering the phone, and the specific length is determined according to the actual situation; the time period S3 is the moment when people are just about to approach the remote control, that is, the moment T, and the moment when they are about to touch or have touched the remote control. In this solution, there may be multiple preamble segments, and by comprehensively analyzing multiple preamble segments, the data will be more reliable.

[0018] Further, step S200 includes:

[0019] Step S210: Obtain all local regions in a certain pre - segment, and obtain the first local region and the second local region therein. Obtain the grayscale histogram of the first local region at each moment in the pre - segment. According to the average grayscale in each grayscale histogram, calculate the grayscale variance of the grayscale histograms of any two adjacent moments T a-1 and T a . If the grayscale variance is greater than the variance threshold, mark the moment T a . If the total duration D1 obtained by adding the marked moments is greater than the first duration threshold, take the grayscale change of the first local region as the first characteristic action;

[0020] Step S220: Obtain the region R of the target device in the monitoring screen d . If the total duration D2 of the intersection area of the second local region, which is the region R2 in the monitoring screen, and the region R d is greater than the area threshold and greater than the second duration threshold, take the change in the intersection area of the second local region as the second characteristic action. Furthermore, according to the total duration corresponding to all pre - segments, obtain the first total duration average value d1 and the second total duration average value d2, as well as the weight value W1 of the first characteristic action: W1 = d2 / (d1 + d2), and the weight value W2 of the second characteristic action: W2 = d1 / (d1 + d2);

[0021] Step S230: Obtain the video data at the current moment. If the total duration corresponding to the first local region of a certain moving object therein and the total duration corresponding to the second local region both meet the judgment of step S200, obtain the first similarity and the second similarity . Furthermore, obtain the similarity value of the current video data and the pre - segment: Q = W1*Q1+W1*Q2.

[0022] The first characteristic action is speaking, and the first local region is the mouth. When a person is speaking, the grayscale value of the mouth will change; the second characteristic action is a person holding a mobile phone in hand, and the second local region is the hand. The target device is a mobile phone. When a person is making a call, the mobile phone will be in the hand. Therefore, according to the position overlap situation between the hand and the mobile phone part, it can be judged whether there may be an action of making a call. By comparing the current video data with each characteristic action, the similarity value can be obtained, and based on the similarity value, it can be judged whether the volume should be adjusted and what the adjusted volume value should be.

[0023] Further, step S300 includes: obtaining the audio input information and audio output information corresponding to each moment in all the previous segments, and converting both the audio input information and the audio output information into spectrograms; if the amplitude difference of a certain frequency in the two spectrograms corresponding to a certain moment is greater than the amplitude threshold, then mark the certain frequency at the certain moment, and if the number of marked occurrences of the same certain frequency in each moment is greater than the marking number threshold, then regard the certain frequency as a changing frequency, and take the average value of the amplitude differences as the change amplitude value of the changing frequency; and then establish an environmental spectrogram according to the changing frequencies and the corresponding change amplitude values.

[0024] Further, step S400 includes:

[0025] Step S410: Obtain the device volume V0 of the audio device at a certain moment T0 in step S100, and the device volume V n at a certain moment T n , and obtain the volume change value Y = V0 - V n of the audio data; add the amplitudes of each frequency in the spectrogram at a certain moment in a certain previous segment, and perform normalization to obtain the total amplitude corresponding to each moment, and obtain the amplitude average value D corresponding to the certain previous segment; mark the moments with the total amplitude greater than the amplitude threshold, take the ratio of the sum of the marked moments to the total duration of the certain previous segment as C, and obtain the environmental audio coefficient X = C * D, and determine the volume adjustment model of the device volume as Y = k * X + b according to the least squares method, where k is the slope and b is the intercept;

[0026] Generally speaking, when people are on the phone, if there is sound at each moment, it means that the hands-free function is turned on. When the hands-free function is turned on, the speaking voice is relatively loud, and there is no need to deliberately reduce the volume of the speaker. Therefore, generally in this case, the volume change of the speaker should be small. If there is not sound at each moment, and if the volume is set small, some information may be missed. In order to be able to hear the sound normally, the volume change should be large at this time;

[0027] Step S420: If the total duration corresponding to the current moment and satisfy and , then obtain as the monitoring period, and obtain the environmental audio coefficient X corresponding to the monitoring period according to the corresponding calculation process of a certain previous segment in step S410 ^ , and then obtain the current volume adjustment value Z = G * Q * (k * X ^ + b) of the audio device according to the slope k, the intercept b, and the similarity value Q, where G is the adjustment coefficient. If Z < Z0, let Z = Z0, and Z0 is the adjustment threshold.

[0028] An adaptive audio data processing system for an audio device, comprising a module for intercepting a front - end segment, a module for obtaining a similarity value, a module for obtaining an environmental spectrogram, and a module for adjusting the volume of the audio device;

[0029] The module for intercepting a front - end segment: used to obtain the historical video data of the monitoring device and the historical audio data of the audio device, and according to the amplitude change in the audio data, obtain the feature area and the target local area in the monitoring screen; establish a plane coordinate system of the monitoring screen, and intercept the front - end segment in the video data according to the coordinate position change of the feature area and the target local area;

[0030] The module for obtaining a similarity value: used to extract the feature actions in the front - end segment, obtain the video data at the current moment, compare and analyze the current video data with the feature actions in the front - end segment, and obtain the similarity value corresponding to the current video data;

[0031] The module for obtaining an environmental spectrogram: used to collect the audio data at the current moment, and convert both the audio input information and the audio output information in the audio data into spectrograms, thereby obtaining an environmental spectrogram;

[0032] The module for adjusting the volume of the audio device: used to pre - obtain the historical environmental spectrogram of the audio device and the device volume change situation, determine the volume adjustment model of the device volume, and intelligently adjust the current volume of the audio device in combination with the similarity value.

[0033] Furthermore, the module for intercepting a front - end segment includes an amplitude analysis unit, a unit for obtaining the target local area, a vector obtaining unit, and a front - end segment intercepting unit;

[0034] The amplitude analysis unit: used to obtain the audio data of the audio device, where the audio data includes each moment, as well as the audio input information, audio output information, and device volume at each moment; obtain the amplitude of the audio output information at each moment, and obtain the average amplitude of the set time period;

[0035] The unit for obtaining the target local area: used to obtain the monitoring screen, establish a plane coordinate system of the monitoring screen, obtain the feature area of the feature item in the monitoring screen, as well as the respective local areas corresponding to each local part of the moving object in the monitoring screen, and the target local area;

[0036] The vector obtaining unit: used to obtain the center coordinates of the target local area, and obtain all vectors according to the center coordinates of the feature area and the change situation of each coordinate;

[0037] The front - end segment intercepting unit: used to make a judgment to obtain all the front - end segments in the video data.

[0038] Further, the audio device volume adjustment module includes a volume adjustment model obtaining unit and a volume change value calculation unit;

[0039] The volume adjustment model obtaining unit: It is used to obtain the volume change value corresponding to the audio data and the environmental audio coefficient, and determine the volume adjustment model of the device volume according to the least square method;

[0040] The volume change value calculation unit: It is used to obtain the monitoring period, obtain the environmental audio coefficient corresponding to the monitoring period, and then obtain the current volume adjustment value of the audio device.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides an adaptive audio device audio data processing system and method, including: obtaining the historical video data of the monitoring device and the historical audio data of the audio device, obtaining the feature area and the target local area, and intercepting the preamble segment in the video data; obtaining the video data at the current moment, and obtaining the similarity value corresponding to the current video data; collecting the audio data at the current moment, and then obtaining the environmental spectrogram; determining the volume adjustment model of the device volume, and combining the similarity value to intelligently adjust the current volume of the audio device. By using the historical video data and audio data, the present invention obtains the device volume that should be set at the current moment, avoids the excessive sound of the speaker affecting normal calls, brings convenience and comfort to users, makes the home audio device more in line with the actual needs of users, and improves the quality of home entertainment and life. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic flowchart of an adaptive audio device audio data processing method of the present invention;

[0043] Figure 2 It is a structural diagram of an adaptive audio device audio data processing system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Embodiment: As Figure 1 shown, the present invention provides a technical solution for an adaptive audio device audio data processing system and method, including the following steps:

[0046] Step S100: Obtain the historical video data of the monitoring device and the historical audio data of the audio device. According to the amplitude changes in the audio data, obtain the feature area and the target local area in the monitoring screen; establish a plane coordinate system for the monitoring screen, and intercept the preposed segment in the video data according to the coordinate position changes of the feature area and the target local area.

[0047] Step S110: Obtain the audio data of the audio device in the recent M days. The audio data includes each moment, as well as the audio input information, audio output information, and device volume at each moment; obtain the amplitude of the audio output information at each moment. Take the device volume at a certain moment T0 as V0. If the device volume within the time period S1 before a certain moment T0 is all V0, and the audio volume V1 at the next moment T1 of a certain moment T0 satisfies V0 > V1, and the audio volumes at subsequent moments decrease in sequence until when a certain moment T n , is the same as the device volume corresponding to the next moment T n of a certain moment T n+1 , obtain the average amplitude F1 within the time period S1, and the average amplitude F2 from a certain moment T0 to a certain moment T n .

[0048] Step S120: If F2 < F1, obtain the monitoring screen of the monitoring device at a certain moment T0, establish a plane coordinate system for the monitoring screen, and through the neural network algorithm, obtain the feature area of the feature item in the monitoring screen and each local part of the moving object, and the corresponding local areas in the monitoring screen; if the intersection area of the feature area and a certain local area in the monitoring screen is greater than 0, then take a certain local area as the target local area.

[0049] In this solution, the audio device is a smart speaker. The audio input information is the audio data received by the microphone of the smart speaker, the audio output information is the audio data output by the speaker, and the device volume is the volume of the smart speaker. The reason for judging F2 < F1 here is that in some cases, turning down the volume of the speaker may not only be because someone is talking, but also because the sound played by the speaker suddenly increases, and the reason for the increase is generally related to the suddenly switched song or some suddenly appearing background music. Therefore, it is necessary to make a judgment to ensure that it is indeed caused by external factors. Both the audio device and the monitoring device are deployed in the living room. The feature item is the remote control for controlling the volume of the audio device, the moving object is a person, and the local parts include the hand and the mouth. In this step, it is the hand. Specifically, if people pick up the remote control to control the volume of the speaker and then the volume of the speaker will become smaller, then the hand of people picking up the remote control is taken as the target local area; and generally, only when situations such as conversation or making a phone call occur will the volume of the speaker be adjusted.

[0050] Step S130: Obtain the central coordinate P of the target local area at a certain moment T in the time period S1. Obtain the central coordinates of the target local area at each moment within the time period S2 starting from moment T and moving forward in time, and calculate the average coordinate P^. Use the central coordinate of the feature area at moment T as P T ; Using the coordinate P as the starting point and P T as the ending point, obtain the vector V0. Use the time period starting from moment T until the moment when the intersection area between the target local area and the feature area is non-zero as S3. Using the central coordinates of the target local area at each moment within the time period S3 as the starting point and P T as the ending point, obtain all vectors.

[0051] Step S140: If 1 / H < L1 / L2 < H and the angle between each vector and the vector V0 is less than the angle threshold, where H is the distance coefficient, L1 is the distance between the coordinate P T and the coordinate P, and L2 is the distance between the coordinate P T and the coordinate P^, then use the time period S2 as the preamble segment, and thus obtain all the preamble segments in the video data.

[0052] A certain moment T0 is the moment when people are just about to switch the volume of the speaker, moment T is the moment when people are just about to approach the remote control, and the time period S2 is the time period before moment T when people are answering the phone, and the specific length is determined according to the actual situation; the time period S3 is the moment when people are just about to approach the remote control, that is, moment T, to the moment when they are about to touch or have touched the remote control. In this solution, there may be multiple preamble segments, and by comprehensively analyzing multiple preamble segments, the data will be more reliable. The distance coefficient H is adjusted according to the actual situation and is 1.2 in this embodiment.

[0053] Step S200: Extract the characteristic actions in the preamble segment, obtain the video data at the current moment, compare and analyze the current video data with the characteristic actions in the preamble segment, and obtain the similarity value corresponding to the current video data.

[0054] Step S210: Obtain all the local areas in a certain preamble segment, and obtain the first local area and the second local area among them. Obtain the gray-level histogram of the first local area at each moment in the preamble segment. According to the average gray level in each gray-level histogram, calculate the gray-level variance of the gray-level histograms at any two adjacent moments T a-1 and T a . If the gray-level variance is greater than the variance threshold, then mark the moment T a . If the total duration D1 obtained by adding the marked moments is greater than the first duration threshold, then use the gray-level change of the first local area as the first characteristic action.

[0055] Step S220: Obtain the area R of the target device in the monitoring screen d , if the area R2 of the second partial area of a moving object in the monitoring screen intersects with the area R d , and the duration D2 during which the intersection area is greater than the area threshold is greater than the second duration threshold, take the total duration D2 as the second time period, and take the change in the intersection area of the second partial area as the second characteristic action; obtain the weight W1 of the first characteristic action = D2 / (D1 + D2), and the weight W2 of the second characteristic action = D1 / (D1 + D2).

[0056] Step S230: Obtain the video data at the current moment. If the gray value change of the first partial area of the target moving object and the change in the intersection area of the second partial area in the video data satisfy the judgment in step S200, obtain the first time period TF1 and the second time period TF2 corresponding to the target moving object, obtain the first similarity Q1 = min(D1, TF1) / max(D1, TF1), obtain the first similarity Q2 = min(D2, TF2) / max(D2, TF2), and further obtain the similarity value between the current video data and the previous segment: Q = W1 * Q1 + W1 * Q2.

[0057] The first characteristic action is speaking, and the first partial area is the mouth. When a person is speaking, the gray value of the mouth will change; the second characteristic action is that a person holds a mobile phone in their hand, and the second partial area is the hand. The target device is a mobile phone. When a person is making a call, the mobile phone will be in their hand. Therefore, according to the position overlap between the hand and the mobile phone part, it is judged that there may be a behavior of making a call; comparing the current video data with each characteristic action, a similarity value can be obtained, and based on the similarity value, it is judged whether the volume should be adjusted and what the adjusted volume value should be.

[0058] Step S300: Collect the audio data at the current moment, and convert both the audio input information and the audio output information in the audio data into spectrograms, and then obtain the environmental spectrogram.

[0059] Obtain the audio input information and the audio output information corresponding to each moment in all the previous segments, and convert both the audio input information and the audio output information into spectrograms; if the amplitude difference of a certain frequency in the two spectrograms corresponding to a certain moment is greater than the amplitude threshold, then mark a certain frequency in a certain moment. If the number of marked times of the same certain frequency in each moment is greater than the marking number threshold, then take a certain frequency as the changing frequency, and take the average value of the amplitude differences as the changing amplitude value of the changing frequency; then, based on each changing frequency and the corresponding changing amplitude value, establish the environmental spectrogram.

[0060] The spectrogram is a tool that visualizes the variation of the frequency and amplitude of a sound signal over time, with frequency on the horizontal axis and amplitude on the vertical axis. There is a corresponding spectrogram at each moment. In the spectrogram of the audio output information, there is only the audio information of the speaker, while in the spectrogram of the audio input information, it also includes the audio information of the human voice. Therefore, by combining and subtracting the two, the environmental spectrogram can be obtained.

[0061] Step S400: Pre-obtain the historical environmental spectrogram of the audio device and the change in device volume, determine the volume adjustment model of the device volume, and intelligently adjust the current volume of the audio device in combination with the similarity value.

[0062] Step S410: Obtain the device volume V0 of the audio device at a certain moment T0 and the device volume V n at a certain moment T n in step S100, and obtain the volume change value Y = V0 - V n corresponding to the audio data; add the amplitudes of each frequency in the spectrogram at a certain moment in a certain preamble segment, and perform normalization to obtain the total amplitude corresponding to each moment, and obtain the amplitude average value D corresponding to a certain preamble segment; mark the moments when the total amplitude is greater than the amplitude threshold, and take the ratio of the sum of the marked moments to the total duration of a certain preamble segment as C, and obtain the environmental audio coefficient X = C * D. According to the least squares method, determine the volume adjustment model of the device volume as Y = k * X + b, where k is the slope and b is the intercept.

[0063] Generally speaking, when people are on the phone, if there is sound at every moment, it means the speakerphone is turned on. When the speakerphone is turned on, the speaking voice is relatively loud, and there is no need to deliberately reduce the speaker volume. Therefore, generally in this case, the volume change of the speaker should be small. If there is not sound at every moment, if the volume is set too small, some information may be missed. In order to be able to hear the sound normally, the volume change at this time should be large.

[0064] Step S420: If the total duration and corresponding to the current moment satisfy and , then obtain the time period as the monitoring period, and obtain the environmental audio coefficient X ^ corresponding to the monitoring period according to the corresponding calculation process of a certain preamble segment in step S410. Then, according to the slope k, the intercept b, and the similarity value Q, obtain the current volume adjustment value Z = G * Q * (k * X ^ + b) of the audio device, where G is the adjustment coefficient. If Z < Z0, let Z = Z0, and Z0 is the adjustment threshold.

[0065] The environmental audio coefficient X ^The calculation is the same as that of step S410, which is to first calculate the amplitude average value, then calculate the ratio, and then multiply to obtain the environmental audio coefficient X. ^ In this embodiment, the adjustment threshold Z0 is the current audio volume multiplied by 0.2.

[0066] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. An adaptive audio data processing method for an audio device, characterized in that The following steps are involved: Step S100: acquiring the video data of the monitoring device history and the audio data of the audio device history, and obtaining the characteristic area and the target local area in the monitoring screen according to the amplitude change in the audio data; establishing a plane coordinate system of the monitoring screen, and intercepting the front segment in the video data according to the coordinate position change of the characteristic area and the target local area; Step S200: extracting characteristic actions in the preceding segment, obtaining video data at the current moment, comparing and analyzing the current video data with the characteristic actions in the preceding segment, and obtaining a similarity value corresponding to the current video data; Step S300: collecting audio data at the current moment, and converting audio input information and audio output information in the audio data into a spectrum graph, thereby obtaining an environment spectrum graph; Step S400: pre-acquire the historical environmental spectrum of the audio device and the volume change of the device, determine the volume adjustment model of the device volume, and intelligently adjust the current volume of the audio device in combination with the similarity value; Step S400 includes: Step S410: Obtain, in step S100, the device volume V0 of the audio device at a certain moment T0, and the device volume V n at a certain moment T n , and obtain the volume change value Y = V0 - V corresponding to the audio data n ; Add the amplitudes of each frequency in the spectrogram at a certain moment in a certain preamble segment, and perform normalization to obtain the total amplitude corresponding to each moment, and obtain the amplitude average value D corresponding to a certain preamble segment; Mark the moments when the total amplitude is greater than the amplitude threshold, and use the ratio of the sum of the marked moments to the total duration of the certain preamble segment as C, and obtain the environmental audio coefficient X = C * D. According to the least squares method, determine the volume adjustment model of the device volume as Y = k * X + b, where k is the slope and b is the intercept; Step S420: If the total duration corresponding to the current moment and meet and then obtain of the time period as the monitoring period. According to the corresponding calculation process of a certain pre - segment in step S410, obtain the environmental audio coefficient X^ corresponding to the monitoring period. Furthermore, according to the slope k, the intercept b, and the similarity value Q, obtain the current volume adjustment value Z of the audio device as Z = G * Q * (k * X^ + b), where G is the adjustment coefficient. If Z < Z0, let Z = Z0, and Z0 is the adjustment threshold.

2. The audio data processing method of an adaptive audio device according to claim 1, wherein, Step S100 includes: Step S110: Obtain the audio data of the audio device in the recent M days. The audio data includes each moment, as well as the audio input information, audio output information, and device volume at each moment; obtain the amplitude of the audio output information at each moment. Take the device volume at a certain moment T0 as V0. If the device volume in the time period S1 before a certain moment T0 is all V0, and the audio volume V1 at the next moment T1 of a certain moment T0 satisfies V0 > V1, and the audio volume at subsequent moments decreases in turn until when a certain moment T n , is the same as a certain moment T n , and the device volume corresponding to the next moment T n+1 of it, obtain the average amplitude F1 in the time period S1, and the average amplitude F2 from a certain moment T0 to a certain moment T n ; Step S120: if F2<F1, obtain the monitoring screen of the monitoring device at a certain time T0, establish a plane coordinate system of the monitoring screen, and obtain the characteristic area of the characteristic object in the monitoring screen and the local areas corresponding to the local parts of the moving object in the monitoring screen through a neural network algorithm; if the area of intersection between the characteristic area and a certain local area in the monitoring screen is greater than 0, take the certain local area as the target local area; Step S130: Obtain the center coordinate P of the target local region at a certain moment T in the time period S1. Obtain the center coordinates of the target local region at each moment within the previous time period S2 starting from moment T, and calculate the average coordinate \(\hat{P}\). Take the center coordinate of the feature region at moment T as P T ; Using the coordinate P as the starting point and P T as the ending point, obtain the vector V0. Take the time period from moment T until the moment when the intersection area between the target local region and the feature region is non-zero as S3. Using the center coordinates of the target local region at each moment within the time period S3 as the starting point and P T as the ending point, obtain all vectors; Step S140: If 1 / H < L1 / L2 < H and the angle between each vector and vector V0 is less than the angle threshold, where H is the distance coefficient, L1 is the distance between coordinate P T and coordinate P, and L2 is the distance between coordinate P T and coordinate P^, then take the time period S2 as the preamble segment, and thus obtain all the preamble segments in the video data.

3. An audio data processing method for an adaptive audio device according to claim 2, characterized in that Step S200 includes: Step S210: Obtain all local regions in a certain pre - segment, and obtain the first local region and the second local region therein. Obtain the grayscale histogram of the first local region at each moment in the pre - segment. According to the average grayscale in each grayscale histogram, calculate the grayscale variance of the grayscale histograms at any two adjacent moments T a-1 and T a . If the grayscale variance is greater than the variance threshold, then mark the moment T a . If the total duration D1 obtained by adding the marked moments is greater than the first duration threshold, then take the grayscale change of the first local region as the first characteristic action; Step S220: Obtain the area R of the target device in the monitoring screen d , if the total duration D2 of the area R2 of the second partial area in the monitoring screen intersecting with the area R d is greater than the second duration threshold, take the change in the intersection area of the second partial area as the second characteristic action; furthermore, according to the total duration corresponding to all the previous segments, obtain the first total duration average value d1 and the second total duration average value d2, as well as the weight W1 of the first characteristic action = d2 / (d1 + d2), and the weight W2 of the second characteristic action = d1 / (d1 + d2); Step S230: Obtain the video data at the current moment. If the total duration corresponding to the first local region of a certain moving object and the total duration corresponding to the second local region both satisfy the judgment in step S200, then obtain the first similarity The second similarity Furthermore, obtain the similarity value between the current video data and the previous segment: Q = W1 * Q1 + W2 * Q2.

4. The audio data processing method of an adaptive audio device according to claim 3, characterized in that Step S300 includes: obtaining the audio input information and audio output information corresponding to each moment in all the preceding segments, and converting the audio input information and the audio output information into a spectrum diagram; if the difference in the amplitude of a certain frequency in the two spectrum diagrams corresponding to a certain moment is greater than the amplitude threshold, then the certain frequency at a certain moment is marked, and if the number of marks of the same certain frequency at each moment is greater than the marking number threshold, then the certain frequency is used as the changing frequency, and the average value of the difference between the amplitudes is used as the changing amplitude value of the changing frequency; and then an environmental spectrum diagram is established according to each changing frequency and the corresponding changing amplitude value.

5. An audio data processing system for an audio device, which is used to execute an adaptive audio data processing method for an audio device described in any one of claims 1-4, characterized in that, The system includes a module for intercepting a pre-segment, a module for obtaining a similarity value, a module for obtaining an environmental spectrum diagram, and a module for adjusting the volume of an audio device; The module for capturing the front segment is used to obtain the historical video data of the monitoring device and the historical audio data of the audio device, and obtain the characteristic area and the target local area in the monitoring screen according to the amplitude change in the audio data; establish the plane coordinate system of the monitoring screen, and capture the front segment in the video data according to the coordinate position change of the characteristic area and the target local area; A module for obtaining a similarity value is used to extract the characteristic action in the preceding segment, obtain the video data at the current moment, compare and analyze the current video data with the characteristic action in the preceding segment, and obtain the similarity value corresponding to the current video data; Obtaining an environmental spectrum graph module: used to collect audio data at the current moment, and convert the audio input information and audio output information in the audio data into a spectrum graph, thereby obtaining an environmental spectrum graph; The audio device volume adjustment module is used to obtain the historical environmental spectrum of the audio device and the device volume changes in advance, determine the volume adjustment model of the device volume, and intelligently adjust the current volume of the audio device in combination with the similarity value.

6. The audio data processing system of an audio device according to claim 5, characterized in that, The module for intercepting the preceding segment includes an amplitude analysis unit, a unit for obtaining a target local area, a unit for obtaining a vector, and a unit for intercepting the preceding segment; Amplitude analysis unit: used to obtain audio data of the audio device, including each moment, as well as audio input information, audio output information and device volume at each moment; obtain the amplitude of the audio output information at each moment, and obtain the average amplitude of the set time period; Obtaining target local area unit: used to obtain the monitoring screen, establish the plane coordinate system of the monitoring screen, obtain the characteristic area of the characteristic object in the monitoring screen, and each local part of the moving object, the corresponding local areas in the monitoring screen, and the target local area; Vector obtaining unit: used to obtain the center coordinates of the local target area, and obtain all vectors according to the center coordinates of the feature area and the changes of each coordinate; The pre-fragment extraction unit is used to make a judgment and obtain all pre-fragments in the video data.

7. An audio data processing system for an audio device according to claim 6, characterized in that, The module for adjusting the volume of the audio device includes a volume adjustment model obtaining unit and a volume change value calculating unit; Obtaining a volume adjustment model unit: used to obtain a volume change value corresponding to the audio data and an environmental audio coefficient, and determine a volume adjustment model of the device volume according to a least squares method; Volume change value calculation unit: used to obtain the monitoring period and the environmental audio coefficient corresponding to the monitoring period, and then obtain the current volume adjustment value of the audio device.

Citation Information

Patent Citations

  • Audio volume intelligent adjustment method and device, electronic equipment and storage medium

    CN114489561A

  • Volume adjusting method, device and equipment

    CN118016091A