Audio playing method, device, electronic device and storage medium

The audio playback method intelligently selects devices using multiple data dimensions, enhancing system intelligence and user experience by reducing manual device selection in multi-device systems.

CN120122912BActive Publication Date: 2025-07-15LINKPLAY TECHNOLOGY INC NANJING
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510609618.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-15
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

In the prior art, the selection of audio playback devices usually relies on default configurations or manual operation of users, which lacks intelligence, resulting in cumbersome and inflexible operations.

Method used

By obtaining multi-dimensional target data, such as location information, device configuration information, historical control information, audio attribute information and current time and space information, the algorithm calculates the matching degree, and automatically selects the most suitable audio playback device for playback.

Benefits of technology

It realizes intelligent selection of audio playback equipment, improves the intelligence of the system, reduces user manual operation steps, and improves user experience and system ease of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120122912B_ABST
    Figure CN120122912B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device, and storage medium for playing audio; among them, applied to an audio playback system, the audio playback system includes at least two audio playback devices, and the method includes: in response to a playback instruction for a target audio, obtaining at least one piece of target data; wherein, the at least one piece of target data includes any one of the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; determining at least one target audio playback device from the audio playback system according to the at least one piece of target data; and playing the target audio through the target audio playback device. The present disclosure can improve the intelligence level of determining the audio playback device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of audio playback, and in particular, to a method, apparatus, electronic device, and storage medium for playing audio. Background Art

[0002] In a playback system configured with multiple playback devices, when determining which one or more playback devices to use for each audio playback, usually one or more playback devices pre-specified by the user are quickly set as the default devices for each playback through default configuration, or the playback devices are specified through the user's operations. This method is not flexible. If the playback device that the user wants to use for the current playback is not the default device, complicated playback device selection operations are required, and the intelligence level is low. Summary of the Invention

[0003] In view of this, an object of the present disclosure is to provide a method, apparatus, electronic device, and storage medium for playing audio to improve the intelligence level of determining an audio playback device.

[0004] In a first aspect, an embodiment of the present disclosure provides a method for playing audio, which is applied to an audio playback system. The audio playback system includes at least two audio playback devices. The method includes: in response to a playback instruction for a target audio, obtaining at least one piece of target data; where the at least one piece of target data includes any one of the position information of a detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; determining at least one target audio playback device from the audio playback system according to the at least one piece of target data; and playing the target audio through the target audio playback device.

[0005] In a second aspect, an embodiment of the present disclosure provides an apparatus for playing audio, which is applied to an audio playback system. The audio playback system includes at least two audio playback devices. The apparatus includes: an obtaining module, configured to obtain at least one piece of target data in response to a playback instruction for a target audio; where the at least one piece of target data includes any one of the position information of a detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; a determining module, configured to determine at least one target audio playback device from the audio playback system according to the at least one piece of target data; and a playing module, configured to play the target audio through the target audio playback device.

[0006] In a third aspect, embodiments of the present disclosure provide an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned audio playing method.

[0007] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above-mentioned audio playing method.

[0008] Embodiments of the present disclosure bring the following beneficial effects:

[0009] The above-mentioned audio playing method, device, electronic device, and storage medium predict an audio playing device through target data in different dimensions, enabling the audio to intelligently select a playing device for playback, rather than simply presetting a playing device through default settings. As a result, the intelligence level of the audio playback system is improved.

[0010] Other features and advantages of the present disclosure will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0011] To make the above objectives, features, and advantages of the present disclosure more obvious and understandable, the following preferred embodiments are specifically given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0012] To more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0013] Figure 1 It is a flowchart of an embodiment of the audio playing method in the embodiments of the present disclosure;

[0014] Figure 2 It is a schematic diagram of an audio playing device provided by the embodiments of the present disclosure;

[0015] Figure 3 It is a schematic diagram of an electronic device provided by the embodiments of the present disclosure. Detailed Embodiments

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0017] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims, and above-mentioned accompanying drawings of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0018] For ease of understanding, the specific process of the embodiments of the present disclosure will be described below. The embodiments of the present disclosure are applied to an audio playback system, and the audio playback system includes at least two audio playback devices. Please refer to Figure 1 , an embodiment of the audio playback method in the embodiments of the present disclosure includes:

[0019] Step S10: In response to a playback instruction for a target audio, obtain at least one piece of target data; wherein, the at least one piece of target data includes any item of the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located.

[0020] The user can trigger a playback instruction for the target audio through a terminal device. For example, the user triggers a playback instruction for the target audio through a television, a control panel, a mobile device (such as a mobile phone), etc. in the audio playback system, so that the target audio is played through the audio playback devices in the audio playback system. The playback instruction for the target audio can also be triggered automatically, such as timed triggering, etc., and specific details are not limited herein.

[0021] Target data is data used to determine a target audio playback device. In this embodiment, one item of target data can be obtained, or more than one item of target data can be obtained. The target data can be any one or more of the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located. It can also be other data that can be used to determine the target audio playback device as listed, and specific details are not limited here.

[0022] The detection target can be any object whose position can be detected. By way of example and not limitation, the detection target can be a device, a living organism, or a designated marker. For example, the detection target can be a mobile phone, a bracelet, a watch, etc. that can be used to represent the position of the wearer, or any human body, a human body with a face feature that conforms to a predetermined standard, a human body with a voiceprint feature that conforms to a predetermined standard, a human body with a fingerprint feature that conforms to a predetermined standard, etc. that can be used to represent the position of the human body. It can also be a marker with specific specific markings such as a predetermined icon, a predetermined shape, or a predetermined weight. Specific details are not limited here.

[0023] For a detection target of a device type, the position information of the detection target can be obtained through methods such as satellite positioning, WiFi signal detection, Bluetooth Low Energy (BLE) signal strength detection, etc. Among them, WiFi signal detection and BLE signal strength detection have higher accuracy in indoor positioning and can improve the accuracy of determining the audio playback device.

[0024] The configuration information of the audio playback device can include the hardware configuration information, software configuration information, and related attribute information of the audio playback device. For example, the hardware configuration information can include the model of the digital-to-analog converter (DAC), the output power and signal-to-noise ratio of the amplifier, the parameters of the interface, the frequency response range, the impedance, the sensitivity, etc. The software configuration information can include the parameters of the operating system, the version of the driver, the sampling rate, the bit depth, the audio enhancement, the network streaming media protocol, the equalizer, etc. The related attribute information can include the identifier, the name, the spatial information where it is located, whether it is connected to an audio-visual device, etc. Specific details are not limited here.

[0025] The historical control information of the audio playback system includes the control information for controlling the audio playback system within a preset duration in the past. For example, the historical control information can include that at a certain moment in the past, a certain user triggered the playback of a certain audio on a certain audio playback device through a certain control device / method. Specific details are not limited here.

[0026] The attribute information of the target audio can include attributes related to sound quality, content type, style, or other attributes that can be used for selecting an audio playback device. The attribute information of the target audio can also be divided into technical attributes and perceptual attributes. For example, technical attributes can include sample rate (SampleRate), bit depth (Bit Depth), number of channels (Channels), bitrate (Bitrate), audio format (File Format), and frequency response (Frequency Response), etc. Perceptual attributes can include pitch, loudness, timbre, spatialization, clarity, etc. There is no specific limitation here.

[0027] The current spatio-temporal information refers to the spatio-temporal information of the environment where the audio playback system is located, which can include current time information, current weather information, current network environment information, current ambient light information, current ambient sound information, etc., which are information that can be used to represent the current environmental state. There is no specific limitation here.

[0028] Step S20: Determine at least one target audio playback device from the audio playback system according to at least one piece of target data;

[0029] In this embodiment, according to each piece of obtained target data, at least one audio playback device matching the target data can be determined from the audio playback system. In one embodiment, the matching degree between each piece of target data and each audio playback device can be calculated respectively to obtain the first matching degree of each audio playback device corresponding to each piece of target data, and then all the first matching degrees corresponding to each audio playback device are fused to obtain the target matching degree of each audio playback device. The audio playback device with a target matching degree higher than the preset matching degree threshold can be determined as the target audio playback device.

[0030] For example, if the audio playback system includes audio playback device A and audio playback device B, and the obtained target data includes target data A and target data B, then the first matching degrees of audio playback device A with target data A and target data B respectively, and the first matching degrees of audio playback device B with target data A and target data B respectively can be calculated first. Then, the first matching degrees of audio playback device A with target data A and target data B respectively are fused to obtain the target matching degree of audio playback device A, and the first matching degrees of audio playback device B with target data A and target data B respectively are fused to obtain the target matching degree of audio playback device B.

[0031] Next, determine whether the target matching degree of the audio playback device A and the target matching degree of the audio playback device B are higher than the preset matching degree threshold respectively. If the target matching degree of the audio playback device A is higher than the preset matching degree threshold, the audio playback device A is determined as the target audio playback device, and the same applies to the audio playback device B, and specific details are not limited here.

[0032] In one implementation, according to the above target data, through a preset device feature generation algorithm, audio playback device feature data matching the above target data can be generated, and then at least one target audio playback device matching the audio playback device feature data can be determined from the audio playback system according to the audio playback device feature data.

[0033] Step S30: Play the target audio through the target audio playback device.

[0034] After determining the target audio playback device, the target audio playback device can be controlled to play the target audio, so that the audio playback can intelligently select the playback device, rather than just intelligently presetting the playback device through the default settings, and the intelligence level of the audio playback system is improved.

[0035] In one implementation, the playback parameters of the audio can also be determined according to the target data, and when playing the target audio through the target audio playback device, play according to the determined playback parameters. Among them, the playback parameters can include parameters such as volume, equalizer, buffer size, mode, etc., and specific details are not limited here.

[0036] The audio playback method provided by the above implementation predicts the audio playback device through target data in different dimensions, so that the audio playback can intelligently select the playback device, rather than just intelligently presetting the playback device through the default settings, and the intelligence level of the audio playback system is improved.

[0037] Next, a specific description of the method for determining the audio playback device for each item of target data will be given.

[0038] In one implementation, at least one item of target data is: the position information of the detection target, and the detection target is the target device or the target organism; when determining at least one target audio playback device from the audio playback system according to at least one item of target data, it includes: calculating the proximity between the detection target and the preset audible range of each audio playback device in the audio playback system according to the position information; wherein, the preset audible range of each audio playback device is used to indicate the target connected space where the corresponding audio playback device is located; determine at least one target audio playback device whose proximity is greater than or equal to the preset proximity threshold from the audio playback system.

[0039] In the case where the target data is the position information of the detection target, when determining the target audio playback device, first, according to the position information of the detection target, calculate the proximity between the detection target and the audible ranges of each audio playback device, and then determine the audio playback devices with a proximity greater than or equal to the preset proximity threshold as the target audio playback devices.

[0040] The preset audible range of an audio playback device refers to the area where the loudness of the sound is greater than or equal to the preset loudness threshold at a specific volume (sound pressure level). That is to say, when playing audio at a certain volume, if the loudness of the sound measured at a certain position is less than the above-mentioned preset loudness threshold, then that position does not belong to the preset audible range. The preset audible range is related to the position where the audio playback device is placed, the connectivity of the space it is in, etc.

[0041] For example, assume an audio playback system in a home environment. Audio playback device A is placed in room A. Assume that when this audio playback device A plays audio at a certain volume in this room, the loudness heard at each position in the room is greater than or equal to the preset loudness threshold, while at any position outside this room, the detected loudness is less than the preset loudness threshold. Then, this room A is the preset audible range of this audio playback device A, and specific details are not limited here.

[0042] When calculating the proximity between the detection target and the preset audible ranges of each audio playback device in the audio playback system, the proximity corresponding to the detection target within the preset audible range of the audio playback device can be determined as 100%. For a detection target outside the preset audible range of the audio playback device, the proximity can be calculated according to the distance between the detection target and the nearest boundary position of the preset audible range. For example, for every 10 cm increase in distance, the proximity decreases by 1%, and specific details are not limited here.

[0043] The detection target can be a target device or a target organism. For example, it can be a target mobile phone, a target bracelet, a target watch, a human body, a target human body, etc. The target audio playback device determined by the position of the detection target can enable the playback of audio to sense the position of the target, achieving the effect of seamless control of playback and having a high degree of intelligence.

[0044] In one implementation, the proximity between the detection target and the preset audible range of the i-th audio playback device in the audio playback system can also be calculated through a preset proximity calculation formula :

[0045]

[0046] where d refers to the distance between the detection target and the preset audible range of the i-th audio playback device in the audio playback system. Refers to the maximum effective distance, when d exceeds the proximity is 0. When ≥ the preset proximity threshold, the i-th audio playback device can be determined as the target audio playback device.

[0047] For example, assume there are three smart speakers in a family, with speaker numbers being 1, 2, and 3 in sequence, located in the living room (i = 1), bedroom (i = 2), and kitchen (i = 3) respectively. When the user wears a smart watch and enters the living room, the system calculates the distances between the user and the preset audible ranges of the three speakers through WiFi signal strength detection as follows: = 2 meters, = 8 meters, = 5 meters. Set = 10 meters, and the preset proximity threshold is 60%.

[0048] Then, the proximity levels corresponding to the three smart speakers are:

[0049] = 100% × (1 - 2 / 10) = 80%

[0050] = 100% × (1 - 8 / 10) = 20%

[0051] = 100% × (1 - 5 / 10) = 50%

[0052] Since > 60%, it is determined that the speaker located in the living room (i.e., the 1st speaker) is the target audio playback device. When the user issues a voice command of "Play today's news", the system automatically plays through the 1st speaker without the user having to manually select the device.

[0053] This algorithm realizes intelligent device selection based on the user's location, reduces the operation steps of the user's manual device selection, and improves the usability and user experience of the system. At the same time, the algorithm uses relative distance percentage calculation, adapts to application scenarios of different space sizes, and has better generality.

[0054] In one embodiment, at least one piece of target data is: the configuration information of each audio playback device in the audio playback system; when determining at least one target audio playback device from the audio playback system according to at least one piece of target data, it includes: obtaining the audio-visual type of the target audio; wherein, the audio-visual type is used to indicate that the target audio is a movie and television audio or a non-movie and television audio; calculating the device adaptation degree of each audio playback device in the audio playback system according to the audio-visual type and the configuration information of each audio playback device; determining the audio playback devices with a device adaptation degree greater than a preset device adaptation degree threshold as the target audio playback devices, and obtaining at least one target audio playback device.

[0055] The target audio can be the audio of a movie and television or not. For movie and television audio and non-movie and television audio, the device adaptation degree of the target audio can be calculated through the configuration information of different audio playback devices to obtain a target audio playback device that better matches the audio-visual type of the target audio.

[0056] When calculating the device adaptation degree of each audio playback device, for the target audio with the audio-visual type of movie and television audio, the number of configuration parameters that meet the audio-visual playback in the configuration information of each audio playback device can be calculated, and the device adaptation degree of each audio playback device can be determined according to the number of configuration parameters; for the target audio with the audio-visual type of non-movie and television audio, the number of configuration parameters that meet the sound quality of the target audio in the configuration information of each audio playback device can be calculated, and then the device adaptation degree of each audio playback device can be determined according to the number of configuration parameters.

[0057] When determining the device adaptation degree of each audio playback device according to the number of configuration parameters, the ratio of the number of configuration parameters to a preset number threshold or the maximum number of parameters can be calculated, and then the ratio can be determined as the device adaptation degree of the corresponding audio playback device, where the maximum number of parameters refers to the maximum value among the number of configuration parameters, and specific details are not limited here.

[0058] In one embodiment, the configuration adaptation degree between different audio-visual types and different configuration parameters can be set through a preset corresponding relationship. When calculating the device adaptation degree of each audio playback device in the audio playback system according to the audio-visual type and the configuration information of each audio playback device, the configuration adaptation degree between each configuration parameter in the configuration information of each audio playback device and the audio-visual type of the target audio can be obtained from the above preset corresponding relationship first, and then the device adaptation degree of the mth audio playback device in the audio playback system can be calculated through a preset device adaptation degree calculation formula :

[0059]

[0060] Wherein, represents the weight of the ith configuration parameter of the mth audio playback device, Indicates the configuration adaptability between the i-th configuration parameter of the m-th audio playback device and the audio-visual type of the target audio.

[0061] For example, assume that the user requests to play the audio of a movie, and the audio-visual type of the target audio is movie audio. Assume that there are three smart speakers in the family, numbered 1, 2, and 3 in sequence, located in the living room (m = 1), bedroom (m = 2), and kitchen (m = 3) respectively. The configuration parameters of each smart speaker include the number of channels (i = 1), audio decoding ability (i = 2), power output (i = 3), and network connection performance (i = 4) in sequence. The weights of each configuration parameter are as follows: = 0.4, = 0.3, = 0.2, = 0.1.

[0062] Through the above preset correspondence, the configuration adaptabilities between the configuration parameters of each smart speaker and movie audio are obtained as follows:

[0063] Smart speaker No. 1 in the living room: 5.1 channels ( = 95%), Dolby decoding ( = 90%), 100W output ( = 85%), wired connection ( = 95%);

[0064] Smart speaker No. 2 in the bedroom: 2.0 channels ( = 60%), AAC decoding ( = 70%), 20W output ( = 60%), Bluetooth connection ( = 75%);

[0065] Smart speaker No. 3 in the kitchen: mono ( = 40%), basic decoding ( = 50%), 10W output ( = 40%), WiFi connection ( = 90%).

[0066] Furthermore, the device adaptabilities corresponding to the 3 smart speakers calculated are as follows:

[0067] =(0.4×95% + 0.3×90% + 0.2×85% + 0.1×95%) / 1 = 91.5%;

[0068] =(0.4×60% + 0.3×70% + 0.2×60% + 0.1×75%) / 1 = 64.5%;

[0069] =(0.4×40% + 0.3×50% + 0.2×40% + 0.1×90%) / 1 = 48%;

[0070] Assume the preset device adaptation threshold is 70%. Then, since =91.5%>70%, the smart speaker in the living room is determined as the target audio playback device. Through multi-parameter weighted calculation, this algorithm achieves precise matching of audio types and device configurations, ensuring the best sound quality experience for users. The parameter weights of the algorithm can be dynamically adjusted according to different audio types, realizing the self-adaptability of the system, while improving the professionalism of audio playback and user satisfaction.

[0071] In one implementation, at least one piece of target data is: historical control information of the audio playback system; when determining at least one target audio playback device from the audio playback system according to at least one piece of target data, it includes: obtaining the current time information and the style label information of the target audio; according to the current time information and the style label information, matching at least one target audio playback device with a historical preference degree higher than the preset preference degree threshold from the historical control information.

[0072] The style label information of the target audio is used to indicate the style of the target audio. For example, classical style, pop style, rock style, jazz style, movie and TV style, etc. According to the current time information and the style label information, a target audio playback device that conforms to the historical preference can be matched from the historical control information, so that the user's habit of playing audio in a specific style at a specific time can be replicated, which can enhance the user's experience in audio playback.

[0073] When matching at least one target audio playback device with a historical preference degree higher than the preset preference degree threshold from the historical control information according to the current time information and the style label information, it is possible to judge whether there is matching control information in the historical control information according to the current time information and the style label information. If so, the historical preference degrees of each audio playback device can be calculated according to the matching control information, and then the audio playback devices with historical preference degrees greater than the preset preference degree threshold are determined as the target audio playback devices.

[0074] When determining whether there is matching control information in the historical control information, it is possible to determine whether there is control information in the historical control information whose playback time is the same as the preset time period in which the current time is located and whose style label is the same. If there is, the audio playback device for playing corresponding to the control information is determined as the preferred playback device. When calculating the historical preference degree, the proportion of the number of each preferred playback device in the total number of all preferred playback devices is counted to obtain the historical preference degree corresponding to each preferred playback device. For non-preferred playback devices, their historical preference degree can be 0, and specific details are not limited here.

[0075] In one implementation manner, at least one piece of target data is: the attribute information of the target audio. When determining at least one target audio playback device from the audio playback system according to at least one piece of target data, it includes: determining the content type of the target audio according to the attribute information; and determining at least one target audio playback device that matches the preset adapted content type of the audio playback device from the audio playback system according to the content type.

[0076] The content type of the target audio refers to the type of the text content in the target audio. According to whether there is text in the target audio, it can be divided into a text type and a non-text type. The non-text type is usually pure music or white noise, etc. The text type can be further divided into a dialogue type or a non-dialogue type. The dialogue type is usually the audio of a movie or TV drama, and the non-dialogue type is usually a song. For the non-dialogue type, it can also be divided into content types according to the emotional tendency of the text content, such as sad type, cheerful type, angry type, plain type, etc., and specific details are not limited here.

[0077] After determining the content type of the target audio, the content type of the target audio can be matched with the preset adapted content types corresponding to each audio playback device, and the audio playback device including the same preset adapted content type is determined as the target audio playback device, so as to obtain at least one target audio playback device. Through content type matching, the playback device can have the type memory ability, avoid playing audio with inconsistent types on inappropriate playback devices, and improve the user experience.

[0078] In one implementation manner, at least one piece of target data is: the current spatio-temporal information where the audio playback system is located, and the current spatio-temporal information includes the current weather information and / or the current time information. When determining at least one target audio playback device from the audio playback system according to at least one piece of target data, it includes: predicting the behavior habit characteristics and emotional characteristics according to the current spatio-temporal information to obtain a characteristic prediction result; and matching at least one target audio playback device whose characteristic matching degree with the characteristic prediction result is greater than a preset characteristic matching degree threshold according to the preset characteristic information of each audio playback device in the audio playback system.

[0079] When determining at least one target audio playback device from an audio playback system according to the current spatio-temporal information where the audio playback system is located, the user's behavior habits and emotions can be predicted based on the current weather information and / or current time information in the current spatio-temporal information, and the user's current behavior habit characteristics and emotion characteristics, that is, the feature prediction result, can be obtained.

[0080] Next, the feature prediction result is matched with the preset feature information corresponding to each audio playback device, the feature matching degree is calculated, the target feature matching degree corresponding to each audio playback device is obtained, and then the audio playback device with the target feature matching degree greater than the preset feature matching degree threshold is determined as the target audio playback device.

[0081] In one implementation, when predicting the behavior habit characteristics and emotion characteristics, the prediction can be performed through a pre-trained artificial intelligence model to obtain the feature prediction result. When performing the matching, the feature matching degree can also be calculated through a preset feature difference algorithm, and the feature difference algorithm can be the Pearson Correlation algorithm, the Spearman's Rank Correlation algorithm, the Mutual Information algorithm, etc., and specific details are not limited here.

[0082] The above describes the specific method for determining the target audio playback device for each item of target data. For the case where the number of items of target data is greater than one, the target audio playback devices determined by each item of target data can be fused and calculated to determine the final target audio playback device for playing the target audio.

[0083] When performing the fusion calculation, the number of target audio playback devices determined by each item of target data can be counted, and the audio playback device with the total number greater than the preset total threshold or the largest total number is determined as the target audio playback device for playing the target audio.

[0084] In one implementation, when determining at least one target audio playback device from an audio playback system according to at least one item of target data, it includes: calculating the matching degree between each item of target data and each audio playback device to obtain the matching degree calculation result corresponding to each item of target data; weighting and superimposing the matching degree calculation result corresponding to each item of target data according to the preset weight information corresponding to each item of target data to obtain the target matching degree of each audio playback device; and determining the audio playback device with the target matching degree higher than the preset device matching degree threshold as the target audio playback device.

[0085] When calculating the matching degree between each piece of target data and each audio playback device, the specific method of determining the target audio playback device using each piece of the above target data can be adopted for calculation to obtain the matching degree calculation result corresponding to each piece of target data.

[0086] For example, if the target data is the position information of the detection target, the corresponding matching degree calculation result can be the above proximity degree; if the target data is the configuration information of each audio playback device in the audio playback system, the corresponding matching degree calculation result can be the above device adaptation degree; if the target data is the historical control information of the audio playback system, the corresponding matching degree calculation result can be the above historical preference degree; if the target data is the attribute information of the target audio, the corresponding matching degree calculation result can be the matching degree of the content type; if the target data is the current spatio-temporal information where the audio playback system is located, the corresponding matching degree calculation result can be the above feature matching degree.

[0087] After obtaining the matching degree calculation results, the preset weight information corresponding to each piece of target data can be weighted and superimposed with the corresponding matching degree calculation results to obtain the target matching degree corresponding to each audio playback device. Then, the audio playback devices with a target matching degree higher than the preset device matching degree threshold can be determined as the target audio playback devices, so that the determination of the target audio playback devices can fuse multi-dimensional data for calculation, improving the accuracy of the determination of the playback devices and enhancing the intelligence level of the system.

[0088] In one implementation, the preset weight information corresponding to the j-th piece of target data can be determined by the following formula:

[0089]

[0090] where, is the preset basic weight of the j-th piece of target data, is the learning rate coefficient of the j-th piece of target data (belonging to [0, 1]), which is a hyperparameter used to control the degree of adjusting parameters in each step of the optimization algorithm of the machine learning model, is the historical success rate index (belonging to [0, 1]), which is related to whether the matching degree calculation result of the corresponding target data successfully predicts the device actually used by the user for audio playback.

[0091] For example, assume that the 1st (j = 1) piece of target data is the position information of the detection target, the 2nd (j = 2) piece of target data is the configuration information of each audio playback device in the audio playback system, the 3rd (j = 3) piece of target data is the historical control information of the audio playback system, and the 4th (j = 4) piece of target data is the current spatio-temporal information where the audio playback system is located.

[0092] Let the preset basic weights be: = 0.3, = 0.3, = 0.2, = 0.2; The historical success rate index is: = 0.6, = 0.8, = 0.9, = 0.7; The learning rate coefficient is: = 0.5, = 0.5, = 0.7, = 0.6.

[0093] Furthermore, calculate the preset weight information corresponding to the jth target data as:

[0094] = 0.3 × (1 + 0.5 × 0.6) = 0.39;

[0095] = 0.3 × (1 + 0.5 × 0.8) = 0.42;

[0096] = 0.2 × (1 + 0.7 × 0.9) = 0.326;

[0097] = 0.2 × (1 + 0.6 × 0.7) = 0.284;

[0098] It is also possible to normalize it to obtain:

[0099] = 0.27, = 0.29, = 0.23, = 0.21;

[0100] The above algorithm for preset weight information realizes the intelligent fusion of multi-dimensional data through dynamic weight coefficients, can adaptively adjust the influence weights of each dimension according to historical usage conditions, and forms an intelligent system that continuously learns and optimizes.

[0101] After obtaining the dynamic preset weight information, the weighted superposition formula can be used to perform weighted superposition on the calculation results of the matching degree corresponding to each target data, so as to obtain the target matching degree of the ith audio playback device :

[0102]

[0103] Among them, represents the matching degree calculation result corresponding to the j-th target data of the i-th audio playback device.

[0104] Suppose there are three smart speakers in the family, numbered 1, 2, and 3 in sequence, located in the living room (i = 1), bedroom (i = 2), and kitchen (i = 3) respectively. For the smart speaker in the living room (i.e., the first smart speaker), are respectively: = 70%, = 95%, = 60%, = 80%. For the smart speaker in the bedroom (i.e., the second smart speaker), are respectively: = 85%, = 75%, = 90%, = 85%. For the smart speaker in the kitchen (i.e., the third smart speaker), are respectively: = 50%, = 70%, = 40%, = 60%.

[0105] Then, the target matching degrees corresponding to these 3 smart speakers in sequence are:

[0106] = 0.27×70% + 0.29×95% + 0.23×60% + 0.21×80% = 77.15%;

[0107] = 0.27×85% + 0.29×75% + 0.23×90% + 0.21×85% = 83.2%;

[0108] = 0.27×50% + 0.29×70% + 0.23×40% + 0.21×60% = 55.9%;

[0109] If the preset device matching degree threshold is 70%, then, since =77.15% > 70%, so the smart speaker in the living room is determined as the target audio playback device. After that, after the actual audio playback device used for playback is selected on the terminal, it can be determined whether the target audio playback device is successfully predicted. If the target audio playback device is the same as the actual audio playback device used for playback selected by the terminal, it indicates that the prediction is successful this time; otherwise, it indicates that the prediction fails. The data of this prediction result (success or failure) can be fed back to the learning module to update the historical success rate index and learning rate coefficient of each dimension.

[0110] In one implementation, when playing the target audio through the target audio playback device, it includes: sending the identification information of the target audio playback device to the target terminal; in response to the audio playback device selection made by the terminal based on the target audio playback device, determining the first audio playback device selected by the target terminal for playing the target audio; playing the target audio through the first audio playback device; and adjusting the preset weight information corresponding to each piece of target data according to the first audio playback device.

[0111] Before playing the target audio, the identification information of the determined target audio playback device can also be sent to the target terminal, so that the target terminal can quickly confirm to play the target audio using the target audio playback device or switch to other audio playback devices to play the target audio. The audio playback device selected by the user through the target terminal for playing the target audio is determined as the first audio playback device.

[0112] The target terminal can be the terminal that triggers the audio playback instruction or any terminal in the audio playback system, and no specific limitation is made here. After the user selects the first audio playback device, the preset weight information corresponding to each piece of target data can be adjusted according to the selected first audio playback device. When the calculation results of the matching degrees corresponding to each piece of target data are weighted and superimposed according to the adjusted preset weight information, among the target matching degrees of each audio playback device obtained, only the target matching degree of the first audio playback device is higher than the preset device matching degree threshold, so that the determination of the playback device can adjust the parameters of the algorithm through the user's selection, and the user participates in the training and optimization of the algorithm, making the algorithm closer and closer to the user's expectations.

[0113] Corresponding to the above method embodiments, refer to Figure 2Schematic diagram of a playback device for audio, which is applied to an audio playback system. The audio playback system includes at least two audio playback devices. The device includes: an acquisition module 20, configured to acquire at least one piece of target data in response to a playback instruction for a target audio; wherein, the at least one piece of target data includes any item of location information of a detection target, configuration information of each audio playback device in the audio playback system, historical control information of the audio playback system, attribute information of the target audio, and current spatio-temporal information in which the audio playback system is located; a determination module 22, configured to determine at least one target audio playback device from the audio playback system according to the at least one piece of target data; a playback module 24, configured to play the target audio through the target audio playback device.

[0114] For the above-mentioned audio playback device, by predicting the audio playback device through target data in different dimensions, the audio playback can intelligently select the playback device, rather than simply presetting the playback device through the default settings, thereby improving the intelligence level of the audio playback system.

[0115] Optionally, the at least one piece of target data is: location information of a detection target, and the detection target is a target device or a target organism; the above-mentioned determination module 22 is configured to: calculate the proximity between the detection target and the preset audible ranges of each audio playback device in the audio playback system according to the location information; determine at least one target audio playback device from the audio playback system whose proximity is greater than or equal to a preset proximity threshold.

[0116] Optionally, the at least one piece of target data is: configuration information of each audio playback device in the audio playback system; the above-mentioned determination module 22 is configured to: obtain the video and audio type of the target audio; wherein, the video and audio type is used to indicate that the target audio is a video and audio for a movie or not; calculate the device adaptation degree of each audio playback device in the audio playback system according to the video and audio type and the configuration information of each audio playback device; determine the audio playback device with a device adaptation degree greater than a preset device adaptation degree threshold as the target audio playback device, and obtain at least one target audio playback device.

[0117] Optionally, the at least one piece of target data is: historical control information of the audio playback system; the above-mentioned determination module 22 is configured to: obtain the current time information and the style label information of the target audio; match at least one target audio playback device with a historical preference degree higher than a preset preference degree threshold from the historical control information according to the current time information and the style label information.

[0118] Optionally, the at least one piece of target data is: attribute information of the target audio; the determining module 22 is configured to: determine the content type of the target audio according to the attribute information; and determine at least one target audio playback device that matches the preset adaptation content type of the audio playback device from the audio playback system according to the content type.

[0119] Optionally, the at least one piece of target data is: the current spatio-temporal information where the audio playback system is located, and the current spatio-temporal information includes current weather information and / or current time information; the determining module 22 is configured to: predict behavior habit characteristics and emotional characteristics according to the current spatio-temporal information to obtain a feature prediction result; and match at least one target audio playback device whose feature matching degree with the feature prediction result is greater than a preset feature matching degree threshold according to the preset feature information of each audio playback device in the audio playback system.

[0120] Optionally, the determining module 22 is configured to: calculate the matching degree between each piece of target data and each audio playback device to obtain a matching degree calculation result corresponding to each piece of target data; perform weighted superposition on the matching degree calculation result corresponding to each piece of target data according to the preset weight information corresponding to each piece of target data to obtain the target matching degree of each audio playback device; and determine the audio playback device with a target matching degree higher than the preset device matching degree threshold as the target audio playback device.

[0121] Optionally, the playback module 24 is specifically configured to: send the identification information of the target audio playback device to the target terminal; in response to the audio playback device selection made by the terminal based on the target audio playback device, determine the first audio playback device selected by the target terminal to play the target audio; play the target audio through the first audio playback device; and adjust the preset weight information corresponding to each piece of target data according to the first audio playback device.

[0122] This embodiment further provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above audio playback method. The electronic device can be a server or a terminal device.

[0123] See Figure 3 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above audio playback method.

[0124] Further, Figure 3The electronic device shown also includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0125] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is realized between this system network element and at least one other network element. The Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0126] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 100 or by instructions in software form. The above-mentioned processor 100 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments, for example:

[0127] In response to a play instruction for a target audio, obtain at least one piece of target data; wherein, the at least one piece of target data includes any one of the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; determine at least one target audio playback device from the audio playback system according to the at least one piece of target data; and play the target audio through the target audio playback device.

[0128] In this method, the audio playback device is predicted through target data in different dimensions, so that the audio playback can intelligently select the playback device, rather than just presetting the playback device through the default settings. The intelligence level of the audio playback system is improved.

[0129] Optionally, the at least one piece of target data is the position information of the detection target, and the detection target is a target device or a target organism; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: calculating the proximity between the detection target and the preset audible range of each audio playback device in the audio playback system according to the position information; and determining at least one target audio playback device from the audio playback system whose proximity is greater than or equal to a preset proximity threshold.

[0130] Optionally, the at least one piece of target data is the configuration information of each audio playback device in the audio playback system; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: obtaining the audio-visual type of the target audio; wherein, the audio-visual type is used to indicate that the target audio is a film and television audio or a non-film and television audio; calculating the device adaptation degree of each audio playback device in the audio playback system according to the audio-visual type and the configuration information of each audio playback device; and determining the audio playback device with a device adaptation degree greater than a preset device adaptation degree threshold as the target audio playback device to obtain at least one target audio playback device.

[0131] Optionally, the at least one piece of target data is the historical control information of the audio playback system; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: obtaining the current time information and the style label information of the target audio; and matching at least one target audio playback device with a historical preference degree higher than a preset preference degree threshold from the historical control information according to the current time information and the style label information.

[0132] Optionally, the at least one item of target data is: attribute information of the target audio; the step of determining at least one target audio playback device from the audio playback system according to the at least one item of target data includes: determining the content type of the target audio according to the attribute information; according to the content type, determining at least one target audio playback device in the audio playback system that matches the preset adapted content type of the audio playback device.

[0133] Optionally, the at least one item of target data is: the current spatio-temporal information where the audio playback system is located, and the current spatio-temporal information includes current weather information and / or current time information; the step of determining at least one target audio playback device from the audio playback system according to the at least one item of target data includes: predicting behavior habit characteristics and emotional characteristics according to the current spatio-temporal information to obtain a characteristic prediction result; according to the preset characteristic information of each audio playback device in the audio playback system, matching at least one target audio playback device whose characteristic matching degree with the characteristic prediction result is greater than a preset characteristic matching degree threshold.

[0134] Optionally, the step of determining at least one target audio playback device from the audio playback system according to the at least one item of target data includes: calculating the matching degree between each item of target data and each audio playback device to obtain a matching degree calculation result corresponding to each item of target data; according to the preset weight information corresponding to each item of target data, performing weighted superposition on the matching degree calculation result corresponding to each item of target data to obtain the target matching degree of each audio playback device; determining the audio playback device with a target matching degree higher than a preset device matching degree threshold as the target audio playback device.

[0135] Optionally, the step of playing the target audio through the target audio playback device includes: sending the identification information of the target audio playback device to the target terminal; in response to the audio playback device selection made by the terminal based on the target audio playback device, determining a first audio playback device selected by the target terminal to play the target audio; playing the target audio through the first audio playback device; adjusting the preset weight information corresponding to each item of target data according to the first audio playback device.

[0136] This embodiment also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above audio playback method, for example:

[0137] In response to a play instruction for a target audio, obtain at least one piece of target data; wherein, the at least one piece of target data includes any one of the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; determine at least one target audio playback device from the audio playback system according to the at least one piece of target data; and play the target audio through the target audio playback device.

[0138] In this method, the audio playback device is predicted through target data in different dimensions, so that the audio playback can intelligently select the playback device, rather than just presetting the playback device through the default settings. The intelligence level of the audio playback system is improved.

[0139] Optionally, the at least one piece of target data is the position information of the detection target, and the detection target is a target device or a target organism; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: calculating the proximity between the detection target and the preset audible range of each audio playback device in the audio playback system according to the position information; and determining at least one target audio playback device in the audio playback system whose proximity is greater than or equal to a preset proximity threshold.

[0140] Optionally, the at least one piece of target data is the configuration information of each audio playback device in the audio playback system; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: obtaining the audio-visual type of the target audio; wherein, the audio-visual type is used to indicate that the target audio is a film and television audio or a non-film and television audio; calculating the device adaptation degree of each audio playback device in the audio playback system according to the audio-visual type and the configuration information of each audio playback device; and determining the audio playback device with a device adaptation degree greater than a preset device adaptation degree threshold as the target audio playback device to obtain at least one target audio playback device.

[0141] Optionally, the at least one piece of target data is the historical control information of the audio playback system; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: obtaining the current time information and the style label information of the target audio; and matching at least one target audio playback device with a historical preference degree higher than a preset preference degree threshold from the historical control information according to the current time information and the style label information.

[0142] Optionally, the at least one piece of target data is: attribute information of the target audio; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: determining the content type of the target audio according to the attribute information; determining at least one target audio playback device that matches the preset adapted content type of the audio playback device from the audio playback system according to the content type.

[0143] Optionally, the at least one piece of target data is: the current spatio-temporal information where the audio playback system is located, and the current spatio-temporal information includes current weather information and / or current time information; the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: predicting behavior habit characteristics and emotional characteristics according to the current spatio-temporal information to obtain a characteristic prediction result; matching at least one target audio playback device whose characteristic matching degree with the characteristic prediction result is greater than a preset characteristic matching degree threshold according to the preset characteristic information of each audio playback device in the audio playback system.

[0144] Optionally, the step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: calculating the matching degree between each piece of target data and each audio playback device to obtain a matching degree calculation result corresponding to each piece of target data; performing weighted superposition on the matching degree calculation results corresponding to each piece of target data according to the preset weight information corresponding to each piece of target data to obtain the target matching degree of each audio playback device; determining the audio playback devices with the target matching degree higher than the preset device matching degree threshold as the target audio playback devices.

[0145] Optionally, the step of playing the target audio through the target audio playback device includes: sending the identification information of the target audio playback device to the target terminal; in response to the audio playback device selection made by the terminal based on the target audio playback device, determining the first audio playback device selected by the target terminal to play the target audio; playing the target audio through the first audio playback device; adjusting the preset weight information corresponding to each piece of target data according to the first audio playback device.

[0146] The computer program product of the audio playback method, device, electronic device and storage medium provided by the embodiments of the present disclosure includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated here.

[0147] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0148] In addition, in the description of the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms "installed", "connected", and "coupled" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific situations.

[0149] If the above function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0150] In the description of the present disclosure, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present disclosure. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0151] Finally, it should be noted that the above embodiments are only specific implementation manners of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A method for playing audio, characterized in that, Applied to an audio playback system, the audio playback system includes at least two audio playback devices, and the method includes: In response to a playback instruction for a target audio, obtain at least one piece of target data; wherein, the at least one piece of target data includes: the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; Determine at least one target audio playback device from the audio playback system according to the at least one piece of target data; Play the target audio through the target audio playback device; The step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: Calculate the matching degree between each piece of target data and each audio playback device to obtain a matching degree calculation result corresponding to each piece of target data; According to the preset weight information corresponding to each piece of target data, perform weighted superposition on the matching degree calculation results corresponding to each piece of target data to obtain the target matching degree of each audio playback device; Determine the audio playback devices with a target matching degree higher than the preset device matching degree threshold as target audio playback devices; Among them, the preset weight information corresponding to the j-th target data is determined by the following formula: Among them, is the preset basic weight of the j-th target data, is the learning rate coefficient of the j-th target data, belongs to [0, 1], and is a hyperparameter used to control the degree to which the machine learning model adjusts parameters in each step of the optimization algorithm, is the historical success rate index, belongs to [0, 1], and is related to whether the calculation result of the matching degree with the corresponding target data successfully predicts the device actually used by the user for audio playback.

2. The method according to claim 1, characterized in that, The at least one piece of target data is: the position information of the detection target, and the detection target is a target device or a target organism; The step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: Calculate the proximity between the detection target and the preset audible range of each audio playback device in the audio playback system according to the position information; Determine at least one target audio playback device from the audio playback system whose proximity is greater than or equal to the preset proximity threshold.

3. The method according to claim 1, wherein The at least one piece of target data is: the configuration information of each audio playback device in the audio playback system; The step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: Obtain the audio-visual type of the target audio; wherein, the audio-visual type is used to indicate that the target audio is a movie and television audio or a non-movie and television audio; Calculate the device adaptability of each audio playback device in the audio playback system according to the audio-visual type and the configuration information of each audio playback device; Determine the audio playback devices with a device adaptability greater than the preset device adaptability threshold as target audio playback devices to obtain at least one target audio playback device.

4. The method according to claim 1, wherein The at least one piece of target data is: the historical control information of the audio playback system; The step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: Obtain the current time information and the style label information of the target audio; Match at least one target audio playback device with a historical preference degree higher than the preset preference degree threshold from the historical control information according to the current time information and the style label information.

5. The method according to claim 1, wherein The at least one piece of target data is: the attribute information of the target audio; The step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: Determine the content type of the target audio according to the attribute information; According to the content type, determine at least one target audio playback device from the audio playback system that matches the preset adapted content type of the audio playback device.

6. The method according to claim 1, wherein The at least one piece of target data is: the current spatio-temporal information where the audio playback system is located, and the current spatio-temporal information includes the current weather information and / or the current time information; The step of determining at least one target audio playback device from the audio playback system according to the at least one piece of target data includes: According to the current spatio-temporal information, predict the behavior habit characteristics and emotional characteristics to obtain a characteristic prediction result; According to the preset characteristic information of each audio playback device in the audio playback system, match at least one target audio playback device whose characteristic matching degree with the characteristic prediction result is greater than the preset characteristic matching degree threshold.

7. The method according to claim 1, characterized in that, The step of playing the target audio through the target audio playback device includes: Send the identification information of the target audio playback device to the target terminal; In response to the audio playback device selection made by the terminal based on the target audio playback device, determine the first audio playback device selected by the target terminal to play the target audio; Play the target audio through the first audio playback device; Adjust the preset weight information corresponding to each piece of target data according to the first audio playback device.

8. An audio playback device, characterized in that, Applied to an audio playback system, the audio playback system includes at least two audio playback devices, and the device includes: An acquisition module, configured to acquire at least one piece of target data in response to a playback instruction for a target audio; wherein, the at least one piece of target data includes any one of the position information of the detection target, the configuration information of each audio playback device in the audio playback system, the historical control information of the audio playback system, the attribute information of the target audio, and the current spatio-temporal information where the audio playback system is located; A determination module, configured to determine at least one target audio playback device from the audio playback system according to the at least one piece of target data; A playback module, configured to play the target audio through the target audio playback device; The above-mentioned determination module is further configured to: calculate the matching degree between each piece of target data and each audio playback device to obtain a matching degree calculation result corresponding to each piece of target data; perform weighted superposition on the matching degree calculation result corresponding to each piece of target data according to the preset weight information corresponding to each piece of target data to obtain the target matching degree of each audio playback device; determine the audio playback device with a target matching degree higher than the preset device matching degree threshold as the target audio playback device; wherein, the preset weight information corresponding to the j-th piece of target data is determined by the following formula: Among them, is the preset basic weight of the j-th target data, is the learning rate coefficient of the j-th target data, belongs to [0, 1], and is a hyperparameter used to control the degree to which the machine learning model adjusts parameters in each step of the optimization algorithm, is the historical success rate index, belongs to [0, 1], and is related to whether the calculation result of the matching degree with the corresponding target data successfully predicts the device actually used by the user for audio playback.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the audio playback method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by the processor, the computer-executable instructions cause the processor to implement the audio playback method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Playing equipment selection method, device and equipment and computer readable storage medium

    CN112333533A

  • Audio playing method and device, electronic equipment and storage medium

    CN113296728A

  • Playing equipment control method and device, equipment and medium

    CN115550755A

  • Audio playing method and electronic equipment

    CN117992008A