Smart device-based system and method for environment scenario identification and decision-making
Through the environmental scene recognition and decision-making system of the smart device, the audio configuration parameters are automatically adjusted, which solves the problem of poor user experience in different scenarios of the smart device and achieves a better user experience.
Patent Information
- Application Number
- PCT/CN2024/071382
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-17
AI Technical Summary
In different usage scenarios, smart devices lack the ability to dynamically adjust audio configuration, resulting in poor user experience.
The environment scene recognition and decision-making system based on intelligent devices is adopted, including the acquisition module, the data processing module, the environment scene recognition module and the decision-making module. By collecting environmental data, processing and identifying environmental scene characteristics, the audio configuration parameters are automatically adjusted.
It effectively improves the user experience of smart devices in different environmental scenarios, and automatically adapts audio configuration parameters.
Smart Images

Figure CN2024071382_17072025_PF_FP_ABST
Abstract
Description
Environmental scene recognition and decision-making system and method based on intelligent equipment Technical Field
[0001] The present application relates to the technical field of smart devices, and in particular to a system and method for environmental scene recognition and decision-making based on smart devices. Background Art
[0002] Smart devices such as smartphones, smart headphones, and AR / VR devices are becoming increasingly common in our daily lives and work. Users are placing increasing demands on the functionality of these devices, even hoping they can compensate for the performance losses caused by varying environmental characteristics. Common usage scenarios include adjusting the phone's external speaker's timbre in noisy environments to enhance playback clarity by increasing mid- and high-frequency sounds, or switching smart headphones from noise-canceling mode to transparency mode while someone is conversing with the user. However, in different usage scenarios, the default parameters or configurations of smart devices are no longer suitable and require manual adjustments by the user. The device lacks the ability to dynamically adjust to the environmental scenario, resulting in a poor user experience. Technical issues
[0003] The present application provides an environmental scene recognition and decision-making system and method based on smart devices, which can recognize different environmental scenes and adapt the audio configuration parameters of corresponding smart devices in different environmental scenes, which can effectively improve the user experience. Technical Solutions
[0004] To solve the above technical problems, this application adopts a technical solution: to provide an environment scene recognition and decision-making system based on smart devices, including:
[0005] A collection module, provided in the smart device, for collecting environmental data;
[0006] A data processing module, connected to the acquisition module, for processing the environmental data to obtain environmental scene features;
[0007] An environmental scene recognition module, connected to the data processing module, is used to recognize the environmental scene features and obtain the corresponding environmental scene;
[0008] The decision module is connected to the environmental scene recognition module and is used to call a preset configuration strategy according to the environmental scene to automatically adjust the audio configuration parameters of the smart device.
[0009] According to one embodiment of the present application, the environmental scene recognition module includes a finite state machine model or a classification network model.
[0010] According to one embodiment of the present application, when the environmental scene recognition module is a finite state machine model, the data processing module includes: one or more of a human voice detection unit, a quiet / noisy detection unit, a long-term noise detection unit, an outdoor / indoor detection unit, a traffic mode detection unit, and a motion / stationary detection unit, and the environmental scene feature is the environmental scene type output by each detection unit.
[0011] According to one embodiment of the present application, when the environmental scene recognition module is a classification network model, the data processing module includes: a feature extraction unit or a feature extraction network, and the environmental scene feature is a feature used to identify the type of the environmental scene.
[0012] According to an embodiment of the present application, the environmental scene characteristics include one or more of human voice characteristics, noise sound pressure level characteristics, noise spectrum characteristics, synthetic acceleration characteristics, angular acceleration characteristics, and carrier-to-noise ratio (CNR) of visible satellites.
[0013] According to one embodiment of the present application, the acquisition module includes one or more of an accelerometer, a gyroscope, a magnetometer, a GPS / GNSS receiver, a light sensor, a proximity sensor, and a microphone.
[0014] According to one embodiment of the present application, the environmental scenes include quiet indoor scenes, quiet outdoor scenes, long-term continuous noise scenes, quiet indoor scenes with people talking, outdoor sports scenes, public transportation scenes, noisy outdoor scenes, and noisy indoor scenes.
[0015] According to one embodiment of the present application, the configuration strategy includes a first audio configuration strategy corresponding to the quiet indoor scene, a second audio configuration strategy corresponding to the quiet outdoor scene, a third audio configuration strategy corresponding to the long-lasting noise scene, a fourth audio configuration strategy corresponding to the quiet indoor scene with someone talking, a fifth audio configuration strategy corresponding to the outdoor sports scene, a sixth audio configuration strategy corresponding to the public transportation scene, a seventh audio configuration strategy corresponding to the noisy outdoor scene, and an eighth audio configuration strategy corresponding to the noisy indoor scene;
[0016] The first audio configuration strategy is that the media playback volume is medium and the tone is balanced;
[0017] The second audio configuration strategy is to play the media at a medium volume and increase the low frequency of the sound effects;
[0018] The third audio configuration strategy is to increase the volume of media playback, reduce low frequencies and increase mid- and high-frequency sound effects;
[0019] The fourth audio configuration strategy is to reduce the media playback volume;
[0020] The fifth audio configuration strategy is to enhance low frequencies during Bluetooth playback;
[0021] The sixth audio configuration strategy is to reduce the media playback volume;
[0022] The seventh audio configuration strategy is to increase the volume of media playback and reduce the low frequency of sound effects;
[0023] The eighth audio configuration strategy is to increase the volume of media playback and reduce the low frequency of sound effects.
[0024] According to one embodiment of the present application, the audio configuration parameters include volume and / or sound effects.
[0025] To solve the above technical problems, another technical solution adopted by this application is to provide an environment scene recognition and decision-making method based on an intelligent device, which is applied to the environment scene recognition and decision-making system. The environment scene recognition and decision-making method includes:
[0026] Collecting environmental data based on the collection module in the smart device;
[0027] Processing the environmental data to obtain environmental scene features;
[0028] Using a finite state machine model or a classification network model to identify the environmental scene features to obtain a corresponding environmental scene;
[0029] A preset configuration strategy is called according to the environmental scenario to automatically adjust the audio configuration parameters of the smart device. Beneficial effects
[0030] The beneficial effects of the present application are: the environmental scene recognition and decision-making system based on smart devices includes an acquisition module, which is provided in the smart device and is used to collect environmental data; a data processing module, which is connected to the acquisition module and is used to process the environmental data and obtain environmental scene characteristics; an environmental scene recognition module, which is connected to the data processing module and is used to recognize the environmental scene characteristics and obtain the corresponding environmental scene; a decision module, which is connected to the environmental scene recognition module and is used to call a preset configuration strategy according to the environmental scene to automatically adjust the audio configuration parameters of the smart device, which can recognize different environmental scenes and adapt the audio configuration parameters of the corresponding smart device in different environmental scenes, and can effectively improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIG1 is a schematic diagram of the architecture of an environment scene recognition and decision-making system based on a smart device according to an embodiment of the present application;
[0032] FIG2 is a schematic diagram of the architecture of an environment scene recognition and decision-making system based on a smart device according to another embodiment of the present application;
[0033] FIG3 is a schematic diagram of states and transition conditions of a finite state machine model according to an embodiment of the present application;
[0034] FIG4 is a schematic diagram of the structure of a neural network model according to an embodiment of the present application;
[0035] FIG5 is a schematic diagram of the architecture of an environment scene recognition and decision-making system based on a smart device according to another embodiment of the present application;
[0036] FIG6 is a schematic diagram of the architecture of an environment scene recognition and decision-making system based on a smart device according to another embodiment of the present application;
[0037] FIG7 is a flow chart of an environment scene recognition and decision-making method based on a smart device according to an embodiment of the present application. Modes for Carrying Out the Invention
[0038] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0039] The terms "first," "second," and "third" in this application are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features identified. Therefore, features identified as "first," "second," or "third" may explicitly or implicitly include at least one of such features. In the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined. All directional designations in the embodiments of this application (such as up, down, left, right, front, back, etc.) are intended only to illustrate the relative positional relationships and movement of components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional designations will also change accordingly. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to such process, method, product, or apparatus.
[0040] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0041] Figure 1 is a schematic diagram of the architecture of an environmental scene recognition and decision-making system based on a smart device according to one embodiment of the present application. As shown in Figure 1, the environmental scene recognition and decision-making system 100 includes: an acquisition module 10, a data processing module 20, an environmental scene recognition module 30, and a decision-making module 40. The data processing module 20 is connected to the acquisition module 10, the environmental scene recognition module 30 is connected to the data processing module 20, and the decision-making module 40 is connected to the environmental scene recognition module 30. The acquisition module 10 is located in the smart device and is used to collect environmental data. The acquisition module 10 may include one or more of an accelerometer, a gyroscope, a magnetometer, a GPS / GNSS receiver, a light sensor, a proximity sensor, and a microphone. The data processing module 20 is used to process the environmental data and obtain environmental scene characteristics; the environmental scene recognition module 30 is used to identify the environmental scene characteristics and obtain the corresponding environmental scene; and the decision-making module 40 is used to invoke a preset configuration policy based on the environmental scene to automatically adjust the audio configuration parameters of the smart device. This environmental scene recognition and decision-making system 100 can recognize different environmental scenes and adapt the corresponding smart device audio configuration parameters to each scene, effectively improving the user experience.
[0042] In one possible implementation, the smart device includes, but is not limited to, a smartphone, a smart headset, etc. The audio configuration parameters include volume and / or sound effects. Exemplarily, the audio configuration parameters include volume. Exemplarily, the audio configuration parameters include sound effects. Exemplarily, the audio configuration parameters include volume and sound effects.
[0043] In one achievable embodiment, the environmental scene recognition module 30 includes a finite state machine model 31 or a classification network model 32. The finite state machine model 31 can solve the problem of transitioning between a finite number of interrelated states. The classification network model 32 can be one or a combination of a feedforward neural network, a convolutional neural network, or a recurrent neural network. The classification network can also be other machine learning algorithms.
[0044] In a feasible implementation, the environmental scenes include a quiet indoor scene, a quiet outdoor scene, a long-lasting noise scene, a quiet indoor scene with people talking, an outdoor sports scene, a public transportation scene, a noisy outdoor scene, and a noisy indoor scene.
[0045] In an achievable embodiment, the configuration strategy includes a first audio configuration strategy corresponding to a quiet indoor scene, a second audio configuration strategy corresponding to a quiet outdoor scene, a third audio configuration strategy corresponding to a long-lasting noise scene, a fourth audio configuration strategy corresponding to a quiet indoor scene with people talking, a fifth audio configuration strategy corresponding to an outdoor sports scene, a sixth audio configuration strategy corresponding to a public transportation scene, a seventh audio configuration strategy corresponding to a noisy outdoor scene, and an eighth audio configuration strategy corresponding to a noisy indoor scene.
[0046] The first audio configuration strategy is to set the media playback volume to a medium level and the timbre to be balanced.
[0047] The second audio configuration strategy is to set the media playback volume to medium and increase the low frequency of the sound effects.
[0048] The third audio configuration strategy is to increase the volume of media playback, reduce low frequencies in the sound effects, and increase mid- and high frequencies.
[0049] The fourth audio configuration strategy is to reduce the media playback volume.
[0050] The fifth audio configuration strategy is to enhance the low frequencies of Bluetooth playback.
[0051] The sixth audio configuration strategy is to reduce the media playback volume.
[0052] The seventh audio configuration strategy is to increase the volume of media playback and reduce the low frequency of sound effects.
[0053] The eighth audio configuration strategy is to increase the volume of media playback and reduce the low frequency of sound effects.
[0054] In one possible embodiment, referring to FIG2 , when the environmental scene recognition module 30 is a finite state machine model 31, the data processing module 20 includes one or more of: a human voice detection unit 21, a quiet / noisy detection unit 22, a prolonged noise detection unit 23, an outdoor / indoor detection unit 24, a traffic pattern detection unit 25, and a motion / stationary detection unit 26. The environmental scene characteristics are the environmental scene types output by each detection unit. For example, environmental scene types include quiet indoor scenes, quiet outdoor scenes, prolonged continuous noise scenes, quiet indoor scenes with people talking, outdoor sports scenes, public transportation scenes, noisy outdoor scenes, and noisy indoor scenes. The finite state machine model 31 switches environmental scene types based on transition conditions. As shown in FIG3 , the initial state is quiet indoors. When the carrier-to-noise ratio (CNR) is detected to be increasing outdoors, the state switches to quiet outdoors; when noise is detected to be increasing, the state switches to noisy indoors; when human voices are detected to be increasing, the state switches to quiet indoor scenes with people talking, and so on. In this embodiment, the acquisition module 10 collects environmental data, each detection unit processes the environmental data to obtain the environmental scene type, the finite state machine model 31 switches the environmental scene type through conversion conditions, and outputs the final environmental scene. The decision module 40 calls the preset configuration strategy according to the environmental scene to automatically adjust the audio configuration parameters of the smart device.
[0055] For example, the quiet / noisy detection unit 22 first frames the audio signal captured by the microphone. Each frame is then multiplied by a preset gain value. The A-weighted equivalent continuous sound pressure level ((LAeq, T), Leq for short) of each frame processed in the previous step is calculated. Finally, Leq is smoothed to eliminate jumps, resulting in Leq_smooth. The smoothing method can be a sliding average or a low-pass filter. Leq_smooth is compared with a first preset threshold. If Leq_smooth is greater than the first preset threshold, a detection result Res_noise = 1 (indicating a noisy environment) is output; otherwise, a detection result Res_noise = 0 (indicating a quiet environment) is output. The first preset threshold can be adjusted based on the actual application scenario, for example, 60. Ultimately, the sound pressure level calculated by the quiet / noisy detection unit needs to be calibrated against an audio calibration system. During calibration, the preset gain value is adjusted to ensure that the sound pressure level output by the quiet / noisy detection unit is consistent with the sound pressure level of the calibration system.
[0056] For example, the long-term noise detection unit 23 caches the Res_noise value from the quiet / noisy detection unit at the end of the longterm_buffer. The length of the longterm_buffer corresponds to the duration of the noise. When the noise intensity exceeds a threshold and persists for a period of time, it is defined as long-term noise. In this embodiment, the number of "1"s in the longterm_buffer is counted, which is the length of the longterm_buffer. When the number of "1"s in the buffer reaches a second preset threshold, the detection result Res_long_term_noise = 1 (indicating long-term noise) is output; otherwise, the detection result Res_long_term_noise = 0 (indicating non-long-term noise) is output. In this embodiment, the second preset threshold can be adjusted according to the actual application scenario.
[0057] For example, in the human voice detection unit 21, whether there is human voice is detected from the audio signal collected by the microphone. Generally, human voice segments and non-human voice segments appear alternately. If human voice is detected, if the time interval for the next appearance of human voice is less than the third preset threshold, the audio signal during this period is also considered to be human voice. For example, in the music playing scene, an echo cancellation algorithm is required to eliminate the audio signal played by the speaker, and a noise suppression module is used to reduce noise to obtain a pure human voice. Whether there is human voice is detected by the human voice detection unit and the stored human voice detection algorithm. There are silent segments interspersed in the human voice segment. Detecting human voice does not give the endpoint of the human voice, but continuously outputs the detection result: whether someone is currently speaking. For example: feature extraction processing is performed on the audio signal collected by the microphone, and whether there is human voice is determined based on the feature extraction result. If there is a human voice, the intermediate result noise_count = 0 is output, otherwise noise_count = noise_count + 1. Compare noise_count with a fourth preset threshold. If noise_count is less than the fourth preset threshold, output the intermediate result Res_voice_meddle = 1 (there is a voice) to voice_buffer; otherwise, output the intermediate result Res_voice_meddle = 0 (there is no voice) to voice_buffer. Count whether the number of 1s in voice_buffer is greater than a fifth preset threshold. If so, output the detection result Res_voice = 1 (there is a voice); otherwise, output the detection result Res_voice = 0 (there is no voice).
[0058] For example, in the outdoor / indoor detection unit 24, one method is to use a combination of sensors for detection, such as accelerometers, gyroscopes, and magnetometers for data fusion. Another method is to use GNSS (Global Navigation Satellite System) data for detection. GNSS data refers to the CNR (Carrier to Noise Ratio) data of all visible satellites in the same satellite navigation system, such as the GPS (Global Positioning System) constellation. GPS signals are obstructed by buildings, causing signal attenuation and a lower signal-to-noise ratio. Therefore, the signal strength outdoors is higher than indoors. CNR data represents the signal strength received by a smart device from a visible satellite. In certain scenarios, multiple satellites may be visible at the same time, so CNR data is a multi-valued array. In this embodiment, when using CNR data for detection, each CNR data array is sorted from largest to smallest, and the top N largest values are selected to form a feature vector. If the number of visible satellites at a given moment is less than N, the feature vector is padded with zeros to ensure that the length of the feature vector is fixed at N. In one embodiment, after collecting a large amount of CNR data for indoor / outdoor scenes, machine learning or deep learning algorithms can be used to perform indoor / outdoor scene recognition. Labeled CNR data is used as a training set to learn a model or mapping that can distinguish indoor / outdoor scenes, and the learned model is used for indoor / outdoor scene detection. In another embodiment, the average or median of the N-dimensional CNR feature vector is calculated and compared with a sixth preset threshold. If the calculated result is greater than the sixth preset threshold, the detection result Res_outdoor = 1 (outdoor) is output; otherwise, the detection result Res_outdoor = 0 (indoor) is output.
[0059] For example, in the motion / stationary detection unit 26, when it is detected that a user is walking / running with a smart device such as a smartphone, it is considered to be in motion (category = 1). When the movement stops, it is considered to be in a stationary state (category = 0). Platforms such as Android provide software programs that implement pedometer functions, which can be used to detect changes in the number of steps of a user. If the change in the number of steps over a period of time is greater than the seventh preset threshold, it is judged to be in motion, otherwise it is judged to be in a stationary state. The pedometer function implemented by this software program uses sensors such as the accelerometer / gyroscope of the smartphone at the bottom layer, detects gait through signal processing or pattern recognition methods, and then realizes step counting.
[0060] For example, in the traffic mode detection unit 25, public transportation refers to taking public transportation such as the subway and bus, while non-public transportation includes walking, stopping, running, and driving. This embodiment uses 3-axis data from a smartphone's accelerometer. Through an end-to-end deep learning approach, feature representations are learned from the raw 3-axis data, and the temporal dynamics of the acceleration time series data are modeled to accurately identify various modes of transportation. Compared to existing detection methods, this embodiment uses raw 3-axis acceleration data as input, eliminating the need for pre-processing the acceleration data, such as removing gravity acceleration, smoothing, or calculating the root mean square value (amplitude) of the 3-axis acceleration. This simplifies the data processing process and improves processing efficiency. The specific detection process is as follows: First, acceleration data is collected and stored, and the data is labeled. The acceleration data includes acceleration data collected by different collectors using different phones at different time periods, including subway, bus, car, walking, running, bicycle, and different grip styles. The acceleration data sampling frequency is 50 Hz. Second, a neural network model is built, as shown in Figure 4. The neural network model inputs the 3-axis acceleration data and outputs category labels. The neural network model consists of a first convolutional layer, a second convolutional layer, a third convolutional layer, a bidirectional long short-term memory network, a fully connected layer, and a normalization layer. The first, second, and third convolutional layers are used to automatically extract features from acceleration data. Both layers utilize 2D convolutional neural networks, Reluctant Unit (ReLU) activation functions, batch normalization, and max pooling. The bidirectional long short-term memory network is a two-layer structure used to model temporal dynamics. When identifying public transportation, subways and buses are classified into one category (labeled 1), while all other data are classified into another category (labeled 0). Third, model training begins with an initial learning rate of 0.001, and the Adam algorithm is used for optimization. The batch size is 50. Regularization is used to prevent overfitting and improve generalization. L2 regularization with a size of 0.001 is used. The convolutional layer activation function uses the Reluctant Unit (ReLU) activation function, and a learning adjustment strategy with piecewise constant decay is used, with the learning rate adjusted to 0.9 of the original value within each segment. When making the training set, the data set is appropriately trimmed to reduce redundant data and increase data diversity, so that the accuracy of model prediction can reach 90%.
[0061] In one achievable embodiment, when the environmental scene recognition module 30 is a classification network model 32, the data processing module 20 includes a feature extraction unit 27 or a feature extraction network 28. Environmental scene features are used to identify the type of environmental scene. Environmental scene features include one or more of human voice features, noise sound pressure level features, noise spectrum features, synthetic acceleration features, angular acceleration features, and the carrier-to-noise ratio of visible satellites. In this embodiment, the environmental scene features output by the data processing module 20 serve as input to the classification network model 32. The output of the classification network model 32 is the environmental scene, and the decision module 40 then adjusts the sound effect parameters and volume based on the scene type.
[0062] Referring to FIG5 , when the data processing module 20 is a feature extraction unit 27, the feature extraction unit 27 may include a human voice feature extraction subunit 271, a noise sound pressure level feature extraction subunit 272, a noise spectrum feature extraction subunit 273, a synthetic acceleration feature extraction subunit 274, an angular acceleration feature extraction subunit 275, and a visible satellite carrier-to-noise ratio extraction subunit 276. In this embodiment, the acquisition module 10 collects environmental data and transmits it to the feature extraction unit 27. The feature extraction unit 27 invokes the corresponding feature extraction subunit based on the environmental data, causing the feature extraction subunit to extract features from the environmental data and transmit the feature extraction results to the classification network model 32. The classification network model 32 identifies the environmental scene based on the feature extraction results and outputs them. The decision module 40 is configured to invoke a preset configuration policy based on the environmental scene to automatically adjust the sound effect parameters and volume of the smart device.
[0063] Referring to Figure 6 , when data processing module 20 is a feature extraction network 28, feature extraction network 28 can be a neural network capable of automatically learning feature representations, such as a convolutional neural network. In this embodiment, acquisition module 10 collects environmental data and transmits it to feature extraction network 28. Feature extraction network 28 automatically learns feature representations based on the environmental data and transmits them to classification network model 32. Classification network model 32 identifies and outputs the environmental scene based on the feature extraction results. Decision module 40 is used to invoke a preset configuration policy based on the environmental scene to automatically adjust the sound parameters and volume of the smart device.
[0064] FIG7 is a flow chart of an environment scene recognition and decision-making method based on a smart device according to an embodiment of the present application. It should be noted that the method of the present application is not limited to the flow sequence shown in FIG7 if substantially the same results are achieved. As shown in FIG7 , the method includes the following steps:
[0065] Step S10: Collect environmental data based on the collection module in the smart device.
[0066] In step S10, the acquisition module may include one or more of an accelerometer, a gyroscope, a magnetometer, a GPS / GNSS receiver, a light sensor, a proximity sensor, and a microphone. The environmental data corresponds to the type of acquisition module. For example, the environmental data acquired by the microphone is an audio signal.
[0067] Step S20: Process the environmental data to obtain environmental scene features.
[0068] In step S20, the environmental data is processed using a data processing module to obtain environmental scene features. In one achievable implementation method, the data processing module includes one or more of a human voice detection unit, a quiet / noisy detection unit, a long-term noise detection unit, an outdoor / indoor detection unit, a traffic mode detection unit, and a motion / stationary detection unit. The environmental scene features are the environmental scene types output by each detection unit. In another achievable implementation method, the data processing module includes a feature extraction unit or a feature extraction network. The environmental scene features are features used to identify the environmental scene type, such as one or more of human voice features, noise sound pressure level features, noise spectrum features, synthetic acceleration features, angular acceleration features, and the carrier-to-noise ratio of visible satellites.
[0069] Step S30: using a finite state machine model or a classification network model to identify environmental scene features and obtain a corresponding environmental scene.
[0070] In step S30, the finite state machine model can solve the problem of transitioning between a finite number of interrelated states. The classification network model can be one or a combination of a feedforward neural network, a convolutional neural network, or a recurrent neural network. The classification network can also be another machine learning algorithm. Environmental scenarios include quiet indoor scenes, quiet outdoor scenes, scenes with long-lasting noise, quiet indoor scenes with people talking, outdoor sports scenes, public transportation scenes, noisy outdoor scenes, and noisy indoor scenes.
[0071] Step S40: calling a preset configuration strategy according to the environmental scenario to automatically adjust the audio configuration parameters of the smart device.
[0072] In step S40, the audio configuration parameters include volume and / or sound effects. In one achievable embodiment, the configuration strategies include a first audio configuration strategy for quiet indoor scenes, a second audio configuration strategy for quiet outdoor scenes, a third audio configuration strategy for long-lasting noise scenes, a fourth audio configuration strategy for quiet indoor scenes with people talking, a fifth audio configuration strategy for outdoor sports scenes, a sixth audio configuration strategy for public transportation scenes, a seventh audio configuration strategy for noisy outdoor scenes, and an eighth audio configuration strategy for noisy indoor scenes.
[0073] The first audio configuration strategy is to set the media playback volume to a medium level and the timbre to be balanced.
[0074] The second audio configuration strategy is to set the media playback volume to medium and increase the low frequency of the sound effects.
[0075] The third audio configuration strategy is to increase the volume of media playback, reduce low frequencies in the sound effects, and increase mid- and high frequencies.
[0076] The fourth audio configuration strategy is to reduce the media playback volume.
[0077] The fifth audio configuration strategy is to enhance the low frequencies of Bluetooth playback.
[0078] The sixth audio configuration strategy is to reduce the media playback volume.
[0079] The seventh audio configuration strategy is to increase the volume of media playback and reduce the low frequency of sound effects.
[0080] The eighth audio configuration strategy is to increase the volume of media playback and reduce the low frequency of sound effects.
[0081] The environmental scene recognition and decision-making method based on smart devices in one embodiment of the present application can effectively improve the user experience by identifying different environmental scenes and adapting the audio configuration parameters of the corresponding smart devices in different environmental scenes.
[0082] The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An environmental scene recognition and decision-making system based on intelligent devices, characterized in that, Comprising: A collection module, disposed in the intelligent device, for collecting environmental data; A data processing module, connected to the collection module, for processing the environmental data to obtain environmental scene features; An environmental scene recognition module, connected to the data processing module, for recognizing the environmental scene features to obtain the corresponding environmental scene; A decision-making module, connected to the environmental scene recognition module, for calling a preset configuration strategy according to the environmental scene to automatically adjust the audio configuration parameters of the intelligent device.
2. The environmental scene recognition and decision-making system according to claim 1, wherein The environmental scene recognition module includes a finite state machine model or a classification network model.
3. The environmental scenario recognition and decision-making system according to claim 2, wherein When the environmental scene recognition module is a finite state machine model, the data processing module includes one or more of: a voice detection unit, a quiet / noisy detection unit, a long-term noise detection unit, an outdoor / indoor detection unit, a traffic mode detection unit, and a motion / still detection unit, and the environmental scene features are the environmental scene types corresponding to the outputs of the respective detection units.
4. The environmental scene recognition and decision-making system according to claim 2, wherein When the environmental scene recognition module is a classification network model, the data processing module includes a feature extraction unit or a feature extraction network, and the environmental scene features are features for identifying environmental scene types.
5. The environmental scene recognition and decision-making system according to claim 4, wherein The environmental scene features include one or more of: voice features, noise sound pressure level features, noise spectrum features, synthetic acceleration features, angular acceleration features, and carrier-to-noise ratio (CNR) of visible satellites.
6. The environmental scenario recognition and decision-making system according to claim 1, characterized in that The collection module includes one or more of: an accelerometer, a gyroscope, a magnetometer, a GPS / GNSS receiver, a light sensor, a proximity sensor, and a microphone.
7. The environmental scene recognition and decision-making system according to claim 1, characterized in that, The environmental scenes include a quiet indoor scene, a quiet outdoor scene, a long-term continuous noise scene, a quiet indoor scene with someone speaking, an outdoor sports scene, a public transportation scene, a noisy outdoor scene, and a noisy indoor scene.
8. The environmental scene recognition and decision-making system according to claim 7, wherein The configuration strategies include a first audio configuration strategy corresponding to the quiet indoor scene, a second audio configuration strategy corresponding to the quiet outdoor scene, a third audio configuration strategy corresponding to the long-term continuous noise scene, a fourth audio configuration strategy corresponding to the quiet indoor scene with someone speaking, a fifth audio configuration strategy corresponding to the outdoor sports scene, a sixth audio configuration strategy corresponding to the public transportation scene, a seventh audio configuration strategy corresponding to the noisy outdoor scene, and an eighth audio configuration strategy corresponding to the noisy indoor scene; The first audio configuration strategy is that the media playback volume is medium and the timbre is balanced; The second audio configuration strategy is that the media playback volume is medium and the sound effect increases the low frequency; The third audio configuration strategy is that the media playback volume is increased, and the sound effect reduces the low frequency and increases the medium and high frequencies; The fourth audio configuration strategy is that the media playback volume is reduced; The fifth audio configuration strategy is that the Bluetooth playback increases the low frequency; The sixth audio configuration strategy is that the media playback volume is reduced; The seventh audio configuration strategy is that the media playback volume is increased, and the sound effect reduces the low frequency; The eighth audio configuration strategy is that the media playback volume is increased, and the sound effect reduces the low frequency.
9. The environmental scene recognition and decision-making system according to claim 1, characterized in that The audio configuration parameters include volume and / or sound effect.
10. An environmental scene recognition and decision-making method based on intelligent devices, characterized in that, Applied to the environmental scenario recognition and decision-making system according to any one of claims 1-9, the environmental scenario recognition and decision-making method includes: Collect environmental data based on the acquisition module in the intelligent device; Process the environmental data to obtain environmental scenario features; Use a finite state machine model or a classification network model to identify the environmental scenario features to obtain the corresponding environmental scenario; Call a preset configuration policy according to the environmental scenario to automatically adjust the audio configuration parameters of the intelligent device.
Citation Information
Patent Citations
Method and system for automatically adjusting multimedia volume according to different scene modes
CN104135705A
Audio paying method and audio playing equipment
CN106648524A
Self-adaptive audio control device and method based on scene recognition
CN110049403A
Volume adjustment method and device of mobile terminal, mobile terminal and storage medium
CN110995933A
Method for adjusting volume of terminal and terminal
CN114845213A