Sound effect control method, system and device based on position sensing and medium
By obtaining user location data through location-aware devices and dynamically adjusting the playback parameters of audio devices, the problem of loss of stereoscopic perception caused by changes in the listener's position when playing on multiple devices is solved, thereby improving the listening experience.
Patent Information
- Application Number
- CN202510814607.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-19
AI Technical Summary
When multiple audio devices are playing simultaneously, the listener moves to different positions, causing the sound to lose its three-dimensional sense. Existing technologies fail to effectively dynamically adjust the playback parameters of the audio devices according to the user's position, affecting the listener's listening experience.
The current location data of the target user is obtained through the location-aware device, the target sound effect parameters are determined based on the data, and the playback of the audio device is controlled to dynamically adjust the sound effect.
It achieves real-time optimization of sound effects based on the user's position, improves the audience's listening experience, and ensures that stable and uniform sound or music can be felt at different positions.
Smart Images

Figure CN120676307A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of audio playback technology, and in particular to a method, system, device, and medium for sound effect control based on position perception. Background Art
[0002] Audio devices are used to play music and other sounds, and the quality of their playback directly impacts the user experience. When multiple audio devices are playing simultaneously, the sound waves emitted by these devices can interfere with each other, leading to a loss of three-dimensionality when the listener moves to a different location.
[0003] Therefore, it is necessary to provide a sound effect control method, system, device and medium based on position perception, which can dynamically adjust the playback parameters of the audio device according to the real-time location of the listener, thereby optimizing the sound effect of the audio device and improving the listener's listening experience. Summary of the Invention
[0004] One or more embodiments of this specification provide a location-aware sound effect control method. The method includes: obtaining current location data of a target user in an audio playback area through a location-aware device; determining target sound effect parameters based on the current location data; and controlling playback of an audio device based on the target sound effect parameters.
[0005] One or more embodiments of the present specification provide a sound effect control system based on position perception, the system including an acquisition module, a parameter determination module and a control module; the acquisition module is configured to obtain the current position data of the target user in the audio playback area through a position perception device; the parameter determination module is configured to determine the target sound effect parameters based on the current position data; and the control module is configured to control the playback of the audio device based on the target sound effect parameters.
[0006] One or more embodiments of the present specification provide a location-aware sound effect control device, which includes at least one processor and at least one memory; the at least one memory is used to store computer instructions; and the at least one processor is used to execute at least part of the computer instructions to implement a location-aware sound effect control method.
[0007] One or more embodiments of the present specification provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes a location-aware sound effect control method. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein: Figure 1 is a schematic diagram of an application scenario of a location-aware sound effect control method according to some embodiments of this specification; Figure 2 is an exemplary module diagram of a sound effect control system based on location awareness according to some embodiments of this specification; Figure 3 is an exemplary flow chart of a method for controlling sound effects based on location awareness according to some embodiments of this specification; Figure 4 is an exemplary schematic diagram of a relationship determination model according to some embodiments of this specification; Figure 5 This is an exemplary schematic diagram of controlling the playback of an audio device according to some embodiments of this specification. DETAILED DESCRIPTION
[0009] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0010] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0011] The playback quality of an audio device depends on a variety of factors, such as the device type and the user's location. Current methods for splitting audio playback into left and right channels fail to account for changes in the user's location, resulting in varying user experiences at different locations. Therefore, some embodiments of this specification dynamically adjust audio device playback parameters based on the user's real-time location, effectively enhancing the user's listening experience.
[0012] Figure 1 This is a schematic diagram of an application scenario of a location-aware sound control method according to some embodiments of this specification.
[0013] In some embodiments, as Figure 1As shown, an application scenario 100 of a location-aware sound effect control method (hereinafter referred to as application scenario 100 ) includes a target user 110 , an audio device 120 , a processor 130 , a network 140 , a storage device 150 and a location-aware device 160 .
[0014] In some embodiments, the application scenario 100 may include a scenario where audio needs to be played in a specific area. For example, the application scenario 100 may be a conference room where audio needs to be played. The processor 130 controls the audio device 120 based on the location of the target user 110 in the conference room to optimize the sound quality of the audio playback. For another example, the application scenario 100 may be a residential room where music needs to be played. The processor 130 controls the audio device 120 based on the location of the target user 110 in the room to optimize the sound quality of the music playback.
[0015] The target user 110 refers to the audience in the application scenario 100 .
[0016] The audio device 120 refers to an electronic device for playing sound. For example, the audio device 120 may include at least one of a speaker, a music player, etc. There may be multiple audio devices.
[0017] In some embodiments, a signal processing unit is provided on the audio device 120. The signal processing unit is used to receive data and / or signals.
[0018] The processor 130 may process data and / or information obtained from other devices or system components. The processor 130 may execute program instructions based on this data, information, and / or processing results to perform one or more functions described in this application. In some embodiments, the processor 130 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-core processing device). By way of example only, the processor may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a graphics processing unit (GPU), a physical processing unit (PPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a microcontroller unit, a reduced instruction set computer (RISC), a microprocessor, or any combination thereof.
[0019] Network 140 can connect the various components of application scenario 100 and / or external resources. Network 140 enables communication between the various components and with external resources, facilitating the exchange of data and / or information. In some embodiments, network 140 can include a fiber optic network, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), or any combination thereof.
[0020] The storage device 150 may be used to store data and / or instructions. In some embodiments, the storage device 150 may include a random access memory (RAM), a read-only memory (ROM), a mass storage device, or the like, or any combination thereof.
[0021] The location sensing device 160 refers to a hardware device used to sense the location of the target user 110 in real time, such as at least one of an ultrasonic device, a high-definition camera, and the like.
[0022] In some embodiments, the application process of the location-aware sound effect control method includes: the location-aware device 160 collects the current location data of the target user 110 and transmits the current location data to the processor 130 via the network 140. The processor 130 generates target sound effect parameters based on the current location data and transmits the target sound effect parameters to the audio device 120 via the network 140 to control the playback of the audio device 120. The storage device 150 can store the data and / or information generated in the above process.
[0023] For more information about the components of the above application scenario 100, please refer to Figures 2 to 5 The corresponding description.
[0024] Figure 2 It is an exemplary module diagram of a location-aware sound control system according to some embodiments of this specification.
[0025] In some embodiments, the location-aware sound effect control system 200 includes an acquisition module 210 , a parameter determination module 220 , and a control module 230 .
[0026] The acquisition module 210 is configured to acquire the current location data of the target user in the audio playback area through a location awareness device.
[0027] The parameter determination module 220 is configured to determine target sound effect parameters based on the current position data.
[0028] In some embodiments, the parameter determination module 220 is further configured to determine the target sound effect parameters based on the current position data and a plurality of parameter correspondences, where each parameter correspondence is a correspondence between a single listening position and a single preset sound effect parameter.
[0029] In some embodiments, the parameter determination module 220 is further configured to determine a target orientation parameter of the audio device based on the current position data.
[0030] In some embodiments, the parameter determination module is further configured to determine a target user among a plurality of listeners within the audio playback area.
[0031] The control module 230 is configured to control the playback of the audio device based on the target sound effect parameters.
[0032] In some embodiments, the control module 230 is further configured to control the playback of the audio device based on the target direction parameter and the target sound effect parameter.
[0033] In some embodiments, the acquisition module 210 may include the position sensing device 160, etc. The parameter determination module 220 may be integrated on the processor 130. The control module 230 may include a signal processing unit on the audio device 120. For a description of the position sensing device 160, the processor 130, and the audio device 120, see Figure 1 and its related descriptions.
[0034] For more information about the acquisition module 210, the parameter determination module 220, and the control module 230, please refer to Figures 3 to 5 The corresponding content.
[0035] In some embodiments of this specification, a sound effect control system based on position awareness can dynamically adjust the playback parameters of an audio device according to the real-time location of the listener, thereby optimizing the sound effect of the audio device and improving the listener's listening experience.
[0036] It should be noted that the above description of the location-aware sound control system 200 and its modules is for convenience only and does not limit this specification to the scope of the embodiments. It is understandable that those skilled in the art, after understanding the principles of the system, may arbitrarily combine the modules or form subsystems connected to other modules without departing from the principles. In some embodiments, Figure 2 The acquisition module 210, parameter determination module 220, and control module 230 disclosed herein may be different modules within a system, or a single module may implement the functions of two or more of the aforementioned modules. For example, the modules may share a storage module, or each module may have its own storage module. Such variations are within the scope of protection of this specification.
[0037] Figure 3 This is an exemplary flow chart of a method for controlling sound effects based on location perception according to some embodiments of this specification. Figure 3 As shown, the process 300 includes the following steps. In some embodiments, the process 300 can be executed by the processor 130. For an explanation of the processor 130, see Figure 1 and its related descriptions.
[0038] Step 310: Obtain the current location data of the target user in the audio playback area through a location sensing device.
[0039] For information about location-aware devices, see Figure 1 and its related descriptions.
[0040] Target users are specific listeners within the audio playback zone, such as conference speakers and singers. An audio playback zone refers to the area covered by the audio played by an audio device. In some embodiments, the audio playback zone can be pre-set and include at least one of a room, a conference room, and the like. Target users can also be pre-set.
[0041] For information about audio devices, see Figure 1 and its related descriptions.
[0042] Current location data refers to the current location data of the target user. The location data may be data representing the target user's location within the audio playback area. In some embodiments, the processor may establish a spatial coordinate system based on the audio playback area and represent the target user's location within the audio playback area using coordinates, i.e., the target user's current location data.
[0043] For example, the processor establishes a spatial coordinate system based on the size of the audio playback area, with the long side of the audio playback area parallel to the ground as the X-axis of the spatial coordinate system, the short side of the audio playback area parallel to the ground as the Y-axis of the spatial coordinate system, and the vertical axis of the audio playback area perpendicular to the ground as the Z-axis of the spatial coordinate system. Then the current location data of the target user can be ( , , ),in, and It can represent the corresponding values of the target user's position on the X-axis and Y-axis. It can represent the value corresponding to the target user's head height on the Z axis.
[0044] In some embodiments, the processor may receive location data of the target user uploaded by a location sensing device. A location sensing device (such as an ultrasonic device) may collect the location data of the target user in real time by transmitting ultrasonic signals and receiving reflected ultrasonic signals.
[0045] In some embodiments, there may be one or more listeners in the audio playback area. The processor may also determine the target user from the listeners in the audio playback area in various ways. For example, the processor may determine registered users from one or more listeners and determine the target user from the registered users.
[0046] In some embodiments, for each listener, the processor may query the user database to determine whether there is a reference feature vector that matches the listener's facial feature vector. If the reference feature vector exists, the processor determines that the listener is a registered user. If the reference feature vector does not exist, the processor determines that the listener is not a registered user.
[0047] In some embodiments, a user database may be pre-configured to include multiple reference feature vectors and registered users corresponding to a single reference feature vector. A reference feature vector refers to a facial feature vector of a registered user. A registered user refers to an audience member whose facial image has been collected in advance and whose facial feature vector has been constructed.
[0048] In some embodiments, the registered user corresponding to the reference feature vector may be represented by an identifier, etc. The identifier may include at least one of a name, an ID number, etc.
[0049] A facial feature vector is a feature vector constructed based on facial features. In some embodiments, the processor may capture facial images of the audience member using a location-aware device (such as a high-definition camera), extract the audience member's facial features from the facial images using a facial recognition algorithm, and construct a facial feature vector based on the facial features. Facial features may include at least one of face shape, eye distance, and skin color. The facial recognition algorithm may include at least one of the MTCNN algorithm and the RetinaFace algorithm.
[0050] In some embodiments, a location-aware device (such as a high-definition camera) can capture facial images of the audience from multiple angles and send the facial images to a processor.
[0051] In some embodiments, the matching condition may include the similarity between the vectors being greater than a similarity threshold. The similarity between the vectors may be negatively correlated with the distance between the vectors. The distance between the vectors may include Euclidean distance, etc. The similarity threshold may be pre-set based on historical experience.
[0052] In some embodiments, the processor may send a facial image capture request to a user terminal (e.g., a smartphone) of each listener when the listener enters the audio playback area. In response to the listener sending an instruction to agree to the request, the processor controls the location-aware device to capture the listener's facial image.
[0053] In some embodiments, if the determined registered user is one person, the processor may determine the registered user as the target user. If the determined registered user is multiple people, the processor may randomly determine one person from the multiple registered users as the target user.
[0054] In some embodiments, the processor may execute multiple user determination strategies in a preset order to determine the target user among the listeners in the audio playback area.
[0055] The user determination policy refers to a policy for determining a target user. In some embodiments, the user determination policy may include a first determination policy, a second determination policy, and a third determination policy. The first determination policy has a higher priority than the second determination policy, and the second determination policy has a higher priority than the third determination policy.
[0056] In some embodiments, the preset order can be pre-set. For example, the processor can continuously monitor the triggering conditions of multiple user-determined policies and preferentially execute the user-determined policies that meet the triggering conditions and have a higher priority.
[0057] In some embodiments, the triggering condition for the first determination strategy includes, for example, that there are multiple registered users among the audience members in the audio playback area. The triggering condition for the second determination strategy includes, for example, that there are no registered users among the audience members in the audio playback area, but there are audience members in the audio playback area who possess device controllers. The triggering condition for the third determination strategy includes, for example, that there are no registered users and no audience members in the audio playback area who possess device controllers.
[0058] In some embodiments, the first determination strategy includes the processor selecting the registered user with the longest residence time in the audio playback area as the target user based on the residence time of multiple registered users in the audio playback area at a preset period. The preset period can be pre-set based on historical experience, such as 30 seconds.
[0059] The dwell time refers to the length of time that the listener continues to reside in the audio playback area. In some embodiments, the processor can set a timer for each registered user when each registered user is confirmed, so as to record the time point when the registered user enters the audio playback area and the dwell time.
[0060] In some embodiments, the processor may capture a facial image of the registered user via the location sensing device at predetermined intervals. In response to failure to acquire the facial image of the registered user, the processor determines that the registered user has left the audio playback zone, and the registered user's timer is paused. If the registered user returns to the audio playback zone, the registered user's timer resumes.
[0061] In some embodiments, if the determined target user leaves the audio playback area, the processor may reselect a registered user with the longest stay time from among the registered users still in the audio playback area as a new target user.
[0062] In some embodiments, since the target user may temporarily leave the audio playback area, in order to avoid frequent switching of target users, the processor can determine whether a new target user needs to be determined after the original target user leaves the audio playback area based on the residence time of the registered users still in the audio playback area and the residence time of the original target user. For example, if the residence time of the registered users still in the audio playback area exceeds the residence time of the original target user, and the exceeded residence time is greater than the duration threshold, the processor can select the registered users who meet the aforementioned conditions with the longest residence time as the new target user. The duration threshold can be pre-set based on historical experience, such as 1 minute.
[0063] The paused dwell time of the original target user refers to the dwell time recorded after the timer of the original target user is paused.
[0064] In some embodiments, the second determination strategy includes the processor targeting listeners holding device controllers as users.
[0065] The device controller can be a controller for controlling the audio device. For example, at least one of a remote control, an application on a user terminal, or a control panel in the audio playback area can be used. The listener can use the device controller to control the audio device's power on and off, volume, and playback content.
[0066] In some embodiments, the device controller may be integrated with a positioning module (such as an ultrasonic transmitter light). When the listener operates the device controller, the processor may obtain the location data of the device controller through a position sensing device (such as a receiver array light in an ultrasonic device) and determine the listener holding the device controller in combination with a high-definition camera.
[0067] In some embodiments, the third determination strategy includes the processor determining listeners located in the preferred listening area as target users.
[0068] The preferred listening area refers to an area in the audio playback area where the listening effect is better. In some embodiments, one or more preferred listening areas may be pre-set in the audio playback area.
[0069] In some embodiments, the processor may obtain the location data of all listeners in the audio playback area through a location sensing device, determine one or more listeners located in one or more preferred listening areas, and select the listener with the longest stay time among the one or more listeners as the target user.
[0070] In some embodiments, if the original target user leaves the preferred listening area, the processor may reselect a listener with the longest stay time from the remaining listeners in the preferred listening area as the target user. The processor reselects the target user from the remaining listeners in the preferred listening area in a manner similar to reselecting the target user from the registered users in the audio playback area.
[0071] In some embodiments of this specification, by executing multiple user determination strategies, target users can be intelligently determined in different situations, thereby improving the level of intelligent sound control and user satisfaction in multi-person scenarios.
[0072] In some embodiments of this specification, compared to pre-setting target users, target users are determined among the audience in the audio playback area. This allows for intelligent determination of target users in different scenarios or when there are large changes in personnel within a scenario, thereby achieving personalized sound effect control.
[0073] Step 320: Determine target sound effect parameters based on the current position data.
[0074] Target sound effect parameters refer to the sound effect parameters to be used. Sound effect parameters refer to parameters related to controlling the playback of an audio device. In some embodiments, the sound effect parameters may include the volume distribution ratio between multiple channels of each audio device in multiple audio devices, the audio signal phase, the signal amplitude adjustment coefficient, and filter parameters. The signal amplitude adjustment coefficient may be an adjustment coefficient for amplifying or attenuating the amplitude of the audio signal.
[0075] The filter parameters may be parameters that control the filter's response to a specific frequency band, such as at least one of an amount of boost or attenuation of signal energy in the specific frequency band, a start point or end point of the specific frequency band, and a width of the specific frequency band. The specific frequency band may be a frequency band associated with the timbre of audio played by the audio device.
[0076] In some embodiments, the processor may determine the target sound effect parameters based on the current location data in a variety of ways. For example, the processor may determine the playback sub-area where the target user is located based on the current location data, and determine the sound effect parameters corresponding to the playback sub-area as the target sound effect parameters.
[0077] The playback sub-area refers to the area after the audio playback area is divided. In some embodiments, the playback sub-area corresponding to the audio playback area can be pre-set, such as different office areas in an office. The sound effect parameters corresponding to the playback sub-area can be pre-set based on historical experience.
[0078] In some embodiments, the processor may determine the target sound effect parameter based on the current position data and one or more parameter correspondences. A parameter correspondence is a correspondence between a listening position and a preset sound effect parameter.
[0079] The listening position refers to the location within the audio playback area used for listening. In some embodiments, the processor may grid the audio playback area to generate multiple grids, and use the coordinates of the geometric center point of each grid in the spatial coordinate system as the listening position. The grid size may be pre-set based on historical experience, such as 1 cubic meter.
[0080] For example, the size of the grid is 1 cubic meter, and the audio playback area can be divided into × × ( 、 、 are all integers not less than 1) grid areas, where In the X-axis direction of the spatial coordinate system, In the Y-axis direction of the spatial coordinate system, In the Z-axis direction of the spatial coordinate system. , , The coordinates of the geometric center points of the grids can be obtained by ( , , ) indicates that Refers to the grid number along the X-axis direction, The value range is 1 to ,in Refers to the grid number along the Y axis, The value range is 1 to ,in Refers to the grid number along the Z axis, The value range is 1 to . =( -0.5), =( -0.5), =( -0.5).
[0081] In some embodiments, if the processor performs grid processing on the audio playback area, the playback sub-area may be represented by one or more grids.
[0082] The preset sound effect parameters refer to the sound effect parameters corresponding to the listening position. In some embodiments, the preset sound effect parameters corresponding to each listening position can be the sound effect parameters pre-set and adjusted by a technician before the audio is played.
[0083] In some embodiments, the parameter mapping relationship can also be pre-set based on historical data. For example, for a listening position, the processor can retrieve multiple historical playback records of the target user at the listening position from a storage device, select the historical playback record with the longest stay time of the target user at the listening position, and use the sound effect parameters used in the historical playback record as the preset sound effect parameters corresponding to the listening position in the parameter mapping relationship to obtain the parameter mapping relationship corresponding to the listening position.
[0084] For example, using the three-channel L / C / R as an example, the parameter mapping relationship may include the coordinates of the listening position, the volume distribution ratio between L / C / R, the audio signal phase, the signal amplitude adjustment coefficient, and the filter parameters. The volume distribution ratio between L / C / R can be (+1dB / -1dB / -0.5dB), the audio signal phase between L / C / R can be (-0.2ms / +0.1ms / +0.3ms), and the signal amplitude adjustment coefficient between L / C / R can be (+1dB / -1dB / 0). The filter parameters can be (a=-2dB, f=6kHz, Q=1.5), where a represents the attenuation of signal energy in a specific frequency band, f represents the starting point of the specific frequency band, and Q represents the width of the specific frequency band.
[0085] In some embodiments, the processor can determine the listening position corresponding to the current position based on the current position data, and determine the parameter correspondence corresponding to the listening position based on the correspondence between the listening position and one and / or multiple parameters, and determine the preset sound effect parameters in the parameter correspondence as the target sound effect parameters.
[0086] In some embodiments, the processor may preprocess the current position data and calculate the spatial distances between multiple listening positions and the target user's current position using a spatial distance formula, and use the listening position with the closest spatial distance as the listening position corresponding to the current position data.
[0087] In some embodiments, the preprocessing may include normalizing the current position data and performing coordinate magnification on the normalized current position data.
[0088] For example, the length of the audio playback area is meters, width is Meters, height meters, the current location data is ( , , ), then the normalized current position data is ( ,y / , ), and then the normalized current position data is magnified. The current position data after coordinate magnification is ( , , ). The spatial distance formula can be shown as formula (1): (1) in, Represents the spatial distance between a single listening position and the current position of the target user. The coordinates of a single listening position are ( , , ).
[0089] In some embodiments, in response to the current location of the target user satisfying the parameter determination condition, the processor may determine the target sound effect parameters based on the current location data.
[0090] The parameter determination condition may be a condition for determining whether the target sound effect parameters need to be determined. In some embodiments, the parameter determination condition may include the spatial distance between the target user's current location and the historical location at the first historical moment being greater than a distance threshold. The spatial distance may be determined using a spatial distance formula, etc. The first historical moment may be pre-set, such as 10 seconds prior to the current moment.
[0091] It is understandable that if the target user moves a long distance, the originally determined sound effect parameters may not guarantee the user's listening experience, so the processor can re-determine the target sound effect parameters to adjust the playback of the audio device.
[0092] In some embodiments, the distance threshold may be pre-set based on historical experience, or the processor may adjust the distance threshold based on user preferences. User preferences refer to the target user's preferred music genre, such as at least one of rock music and symphony music. The processor may obtain the user preferences from a storage device.
[0093] For example, if the user prefers rock music, the target user is more likely to move, and the processor may lower the distance threshold. If the user prefers symphony music, the target user is less likely to move, and the processor may increase the distance threshold. The extent to which the processor lowers or increases the distance threshold may be pre-set based on historical experience.
[0094] In some embodiments, in response to the current location of the target user satisfying the parameter determination condition, the processor may determine the target sound effect parameters based on the current location data by the above-mentioned method of determining the target sound effect parameters.
[0095] In some embodiments of this specification, by judging whether the current location of the target user meets the parameter determination conditions, and then determining whether to determine the target sound effect parameters, unnecessary sound effect parameter updates are reduced and the smoothness of the user experience is improved.
[0096] In some embodiments of this specification, more specific target sound effect parameters can be determined based on the correspondence between the current position data and one or more parameters to ensure the listening experience of the target user in different positions.
[0097] Step 330: Control the playback of the audio device based on the target sound effect parameters.
[0098] In some embodiments, the processor may convert the target sound effect parameter into a control signal, and send the control signal to a signal processing unit of the audio device, thereby controlling the playback of the audio device.
[0099] In some embodiments, the processor can also control the playback of the audio device based on the target direction parameter and the target sound effect parameter. Figure 5 and its related descriptions.
[0100] In some embodiments of this specification, the playback parameters of the audio device are dynamically adjusted according to the real-time location of the target user, thereby optimizing the sound effects of the audio device so that the target user can still experience stable and uniform sound or music at different locations.
[0101] Figure 4 It is an exemplary schematic diagram of a relationship determination model according to some embodiments of this specification.
[0102] In some embodiments, the processor may determine one and / or multiple parameter correspondences (such as parameter correspondence 440 ) through a relationship determination model 430 based on multiple listening positions (such as listening position 410 ) and environmental data corresponding to the multiple listening positions (such as environmental data 420 ).
[0103] In some embodiments, in response to the environmental data satisfying the adjustment condition, the processor may further adjust one and / or multiple parameter correspondences through a relationship determination model.
[0104] For more information on the relationship between listening position and parameters, see Figure 3 and its related descriptions.
[0105] Environmental data refers to data related to the environment of the listening position. In some embodiments, the environmental data may include at least one of temperature, humidity, and air pressure.
[0106] In some embodiments, since listeners in the audio playback area may perform various actions, the environmental data at different listening positions in the audio playback area may be different. Therefore, each listening position may correspond to a set of environmental data. The various actions performed by listeners may include turning on a heater, humidifier, or air conditioner.
[0107] In some embodiments, if the environment in the audio playback area is relatively stable, multiple (eg, three) adjacent listening positions may correspond to a set of environmental data.
[0108] In some embodiments, the processor may obtain environmental data corresponding to one or more listening positions using a variety of sensors deployed at the one or more listening positions, wherein the sensors include at least one of a temperature sensor, a humidity sensor, and an air pressure sensor.
[0109] In some embodiments, the relationship determination model may be a machine learning model. For example, the relationship determination model may include any one or a combination of a neural network (NN) model or other custom model structures.
[0110] In some embodiments, the input of the relationship determination model may include a single listening position and environmental data corresponding to the listening position, and the output may include preset sound effect parameters corresponding to the listening position.
[0111] In some embodiments, the processor may determine the model based on a large number of first training samples with first labels using a training relationship such as a gradient descent method. The first training samples may include sample listening positions and sample environment data, and the first labels of the first training samples may be actual sound effect parameters used at the sample listening positions.
[0112] In some embodiments, the first training sample and the first label can be obtained based on historical playback records. For example, the processor can filter out a preferred playback record from multiple historical playback records, use the historical listening position and historical environment data corresponding to the preferred playback record as the first training sample, and use the sound effect parameters actually used in the preferred playback record as the first label.
[0113] In some embodiments, the processor can count the historical dwell time of historical users at historical listening positions in multiple historical playback records, and select historical playback records with a historical dwell time greater than a dwell time threshold as preferred playback records. The dwell time threshold can be pre-set based on historical experience. For an explanation of the dwell time, see Figure 3 and its related descriptions.
[0114] In some embodiments, the relationship determination model can be trained by inputting a plurality of first training samples with first labels into an initial relationship determination model, constructing a loss function based on the first labels and prediction results of the initial relationship determination model, iteratively updating the initial relationship determination model based on the loss function, and completing the relationship determination model training when the loss function of the initial relationship determination model satisfies a preset condition. The preset condition may be that the loss function converges, the number of iterations reaches a set value, etc.
[0115] The adjustment condition may be a condition for determining whether the parameter correspondence needs to be adjusted. In some embodiments, the adjustment condition may include a change in one or more environmental data values being greater than a corresponding environmental change threshold. The environmental change threshold may be pre-set based on historical experience, such as a temperature change threshold, a humidity change threshold, and an air pressure change threshold.
[0116] In some embodiments, the change in environmental data can be represented by the difference between the environmental data at the current moment and the environmental data at a second historical moment. The change in each type of environmental data is calculated independently. The second historical moment can be pre-set based on historical experience, such as 24 hours prior to the current moment.
[0117] In some embodiments, in response to the environmental data satisfying an adjustment condition, the processor may further adjust one or more parameter correspondences using a relationship determination model. For example, the processor inputs the environmental data satisfying the adjustment condition and the corresponding listening position into the relationship determination model, obtains preset sound effect parameters corresponding to the listening position output by the relationship determination model, and replaces the preset sound effect parameters in the parameter correspondence corresponding to the listening position with the preset sound effect parameters output by the relationship determination model.
[0118] In some embodiments, the processor may further determine the number of parameter correspondences based on the environmental data distribution.
[0119] The environmental data distribution may represent the distribution of environmental data at different listening positions. In some embodiments, the environmental data distribution may include at least one of temperature distribution, humidity distribution, and air pressure distribution.
[0120] In some embodiments, the processor may determine the environmental non-uniformity of the audio playback area based on the distribution of the environmental data, and determine the number of parameter correspondences based on the environmental non-uniformity.
[0121] The environmental non-uniformity can represent the stability of the environment in the audio playback area. The higher the environmental non-uniformity, the more unstable the environment in the audio playback area.
[0122] In some embodiments, the processor can calculate the standard deviation of each environmental data item, such as temperature, humidity, and air pressure, within the audio playback area based on the distribution of environmental data, and weightedly calculate the sum of the temperature standard deviation, humidity standard deviation, and air pressure standard deviation to obtain the environmental non-uniformity. The weights corresponding to environmental data items such as temperature, humidity, and air pressure can be pre-set based on historical experience.
[0123] In some embodiments, since one parameter correspondence corresponds to one listening position and one listening position corresponds to one grid, after determining the number of parameter correspondences, the processor can adjust the size of the grids to make the number of grids the same as the number of parameter correspondences.
[0124] In some embodiments of this specification, the number of one and / or multiple parameter correspondences is determined based on the distribution of environmental data, so that the fineness of the grid division of the audio playback area can be dynamically adjusted according to the stability of the environment, thereby improving the accuracy of determining the target sound effect parameters.
[0125] In some embodiments of this specification, by determining the relationship model, the sound effect parameters suitable for the listening position can be quickly determined according to the listening position and the environment around the listening position, thereby realizing intelligent generation of parameter correspondence and environmental adaptive update.
[0126] Figure 5 This is an exemplary schematic diagram of controlling the playback of an audio device according to some embodiments of this specification.
[0127] In some embodiments, the processor may determine a target orientation parameter 520 of the audio device 540 based on the current position data 510 ; and control the playback of the audio device 540 based on the target orientation parameter 520 and the target sound effect parameter 530 .
[0128] After receiving the target orientation parameter, the signal processing unit of the audio device may control a driving device of the audio device to drive the audio device to rotate, so as to change the orientation of the audio device.
[0129] For a description of the current position data and target sound effect parameters, see Figure 3 and its related descriptions.
[0130] The target orientation parameter refers to an orientation parameter used for confirmation. The orientation parameter refers to a parameter related to the angle of the audio device's orientation. In some embodiments, the orientation parameter may include at least one of a horizontal rotation angle and a vertical pitch angle of each of the multiple audio devices.
[0131] In some embodiments, the processor may determine the target orientation parameter of the audio device based on the current location data in various ways. For example, the processor may determine the playback sub-area where the target user is located based on the current location data, and determine the orientation parameter corresponding to the playback sub-area as the target orientation parameter. The orientation parameter corresponding to the playback sub-area may be pre-set.
[0132] For instructions on playing the sub-area, please refer to step 320 and its related description.
[0133] In some embodiments, the processor may determine a target orientation parameter based on the current position data and one or more orientation correspondences. Each orientation correspondence is a correspondence between a single listening position and a single preset orientation parameter. For an explanation of the listening position, see step 320 and its related description.
[0134] The preset orientation parameter refers to an orientation parameter corresponding to the listening position.
[0135] In some embodiments, the orientation correspondence can also be pre-set based on historical data. For example, for a listening position, the processor can retrieve multiple historical playback records of the target user at the listening position from a storage device, select the historical playback record with the longest stay time of the target user at the listening position, and use the orientation parameter used in the historical playback record as the preset orientation parameter corresponding to the listening position in the orientation correspondence, thereby obtaining the orientation correspondence corresponding to the listening position.
[0136] For example, using the three-channel L / C / R setup, the orientation correspondence can include the coordinates of the listening position and the orientation parameters of each audio device in the L / C / R setup. The orientation parameters for each audio device in the L / C / R setup can be (+10° / 0), (-25° / 0), or (-40° / 0), for example. (+10° / 0) indicates a +10° horizontal rotation angle and unchanged vertical pitch angle. (-25° / 0) indicates a -25° horizontal rotation angle and unchanged vertical pitch angle. The positive and negative directions of the horizontal and vertical pitch angles can be preset.
[0137] In some embodiments, the processor can determine the listening position corresponding to the current position based on the current position data, and determine the orientation correspondence corresponding to the listening position based on the listening position and one and / or multiple orientation correspondences, and determine the preset orientation parameter in the orientation correspondence as the target orientation parameter.
[0138] For an explanation of determining the listening position corresponding to the current position based on the current position data, refer to step 320 and its related description.
[0139] In some embodiments of the present specification, a target orientation parameter can be quickly determined through one and / or multiple orientation correspondences, ensuring that the audio device can be quickly and accurately aligned with the target user.
[0140] In some embodiments, the processor may determine one and / or multiple orientation correspondences through a relationship determination model based on multiple listening positions and environment data corresponding to the multiple listening positions.
[0141] For an explanation of the relationship determination model, see Figure 4 and its related descriptions.
[0142] In some embodiments, the output of the relationship determination model may further include a preset orientation parameter corresponding to the listening position. The processor may input each listening position and the environment data corresponding to each listening position into the relationship determination model to obtain the preset orientation parameter corresponding to the listening position output by the relationship determination model.
[0143] In some embodiments, if the output of the relationship determination model includes a preset orientation parameter corresponding to the listening position, the first label may also include the orientation parameter actually used by the sample listening position. For instructions on determining the first label, see Figure 4 and its related descriptions.
[0144] In some embodiments, the processor may determine preset orientation parameters output by the model based on the listening position and the relationship, and construct an orientation correspondence relationship corresponding to the listening position.
[0145] In some embodiments of this specification, a relationship determination model is used to determine preset orientation parameters corresponding to the listening position, and then a orientation correspondence is constructed, thereby realizing intelligent and adaptive determination of the orientation correspondence, which is conducive to determining target orientation parameters that can coordinate with target sound effect parameters to optimize the listening experience.
[0146] In some embodiments, the processor may convert the target direction parameter and the target sound effect parameter into a control signal, and send the control signal to the signal processing unit of the audio device, thereby controlling the playback of the audio device.
[0147] In some embodiments of this specification, controlling the playback of an audio device through target orientation parameters and target sound effect parameters can improve the accuracy of audio playback and the personalized listening experience of the target user.
[0148] Some embodiments of this specification also provide a location-aware sound effect control device, comprising at least one processor and at least one memory. The at least one memory is configured to store computer instructions. The at least one processor is configured to execute at least some of the computer instructions to implement any of the methods described in the above embodiments.
[0149] Some embodiments of this specification further provide a computer-readable storage medium, which stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes any one of the methods in the above embodiments.
[0150] Furthermore, certain features, structures, or characteristics in one or more embodiments of this specification may be appropriately combined.
[0151] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may vary according to the required features of the individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
[0152] If there is any inconsistency or conflict between the descriptions, definitions, and / or usage of terms in the materials cited in this specification and the contents described in this specification, the descriptions, definitions, and / or usage of terms in this specification shall prevail.
Claims
1. A sound effect control method based on position perception, characterized in that: The method comprises: Obtain the current location data of the target user in the audio playback area through the location-aware device; Determining target sound effect parameters based on the current position data; Based on the target sound effect parameters, the playback of the audio device is controlled.
2. The method according to claim 1, wherein The determining of target sound effect parameters based on the current position data includes: The target sound effect parameter is determined based on the current position data and one or more parameter correspondences, where a parameter correspondence is a correspondence between a listening position and a preset sound effect parameter.
3. The method according to claim 1, wherein The controlling the playing of the audio device based on the target sound effect parameter includes: determining a target orientation parameter of the audio device based on the current position data; Based on the target direction parameter and the target sound effect parameter, the playing of the audio device is controlled.
4. The method according to claim 1, wherein The method further comprises: The target user is determined among the listeners in the audio playback area.
5. A sound effect control system based on position perception, characterized in that: The system includes an acquisition module, a parameter determination module and a control module; The acquisition module is configured to acquire the current location data of the target user in the audio playback area through a location sensing device; The parameter determination module is configured to determine target sound effect parameters based on the current position data; The control module is configured to control the playing of the audio device based on the target sound effect parameter.
6. The system according to claim 1, wherein: The parameter determination module is further configured to: The target sound effect parameters are determined based on the current position data and a plurality of parameter correspondences, each parameter correspondence being a correspondence between a single listening position and a single preset sound effect parameter.
7. The system according to claim 1, wherein: The parameter determination module is further configured to determine a target orientation parameter of the audio device based on the current position data; the control module is further configured to control the playback of the audio device based on the target orientation parameter and the target sound effect parameter.
8. The system according to claim 5, wherein: The parameter determination module is further configured to: The target user is determined among a plurality of listeners in the audio playback area.
9. A sound effect control device based on position perception, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is for storing computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the method according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that The storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the method according to any one of claims 1 to 4.