Dynamic multi-stream implementation method and system, and medium
Generate grouping and playback instructions through user operations, and automatically group and control multi-audio devices, solving the problem that users find it difficult to quickly realize multi-device grouping playback and improving user experience.
Patent Information
- Application Number
- PCT/CN2024/071825
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-01-11
- Publication Date
- 2025-05-22
AI Technical Summary
In a multi-audio device environment, it is difficult for users to quickly realize group playback of multiple audio devices in a specific area. Traditional methods are cumbersome and naming is inconvenient for quick selection.
The user generates a packet instruction through the first operation, and the grouping module groups the audio devices in the target area, generates device grouping information, and then generates playback instructions based on the user's second operation, and controls the target device to play the target audio source.
It realizes the automatic and rapid implementation of user playback needs, improves user experience, and simplifies the grouping and playback operations of multi-audio devices.
Smart Images

Figure CN2024071825_22052025_PF_FP_ABST
Abstract
Description
A method, system and medium for implementing dynamic multi-stream
[0001] Cross-references
[0002] This application claims priority to Chinese application No. 202311544276.1, filed on November 17, 2023, and the entire contents of the above application are incorporated herein by reference. Technical Field
[0003] This specification relates to the field of audio playback technology, and in particular to a method, system, and medium for implementing dynamic multi-streaming. Background Art
[0004] When a user wants to play one or more specific audio or video sources on audio devices in a designated area of a space with multiple audio devices (such as a home, shopping mall, exhibition center, or conference hall), traditional methods often do not support grouping multiple audio devices. Playing audio or video sources by audio device group requires users to operate a single audio device multiple times to achieve the playback requirement, which is relatively cumbersome. Furthermore, audio devices are often named using addresses or random symbols, which makes it difficult for users to quickly select the appropriate device to play the audio or video source.
[0005] Therefore, it is necessary to provide a method for implementing dynamic multi-streaming to automatically and quickly meet the user's playback needs and improve the user's experience.
[0006] Summary of the Invention
[0007] One or more embodiments of this specification provide a method for implementing dynamic multi-streaming. The method includes: generating a grouping instruction based on a first user operation, wherein the first user operation is related to the user's grouping requirement; grouping audio devices in a target area based on the grouping instruction to obtain device grouping information; generating a playback instruction based on the device grouping information and a second user operation, wherein the second user operation is related to the user's playback requirement, wherein the playback requirement includes a target audio source; and controlling the target device to play the target audio source based on the playback instruction.
[0008] One or more embodiments of the present specification provide a dynamic multi-stream implementation system, including: a grouping instruction generation module, configured to generate a grouping instruction based on a first user operation, wherein the first user operation is related to the user's grouping requirement; a grouping module, configured to group audio devices in a target area based on the grouping instruction to obtain device grouping information; a playback instruction generation module, configured to generate a playback instruction based on the device grouping information and a second user operation, wherein the second user operation is related to the user's playback requirement, wherein the playback requirement includes a target sound source; and a playback module, configured to control the target device to play the target sound source based on the playback instruction.
[0009] One or more embodiments of this specification provide a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes a method for implementing dynamic multi-streaming. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, like numbers represent like structures, wherein:
[0011] FIG1 is a schematic diagram of an application scenario of a system for implementing dynamic multi-flow according to some embodiments of this specification;
[0012] FIG2 is an exemplary module diagram of a system for implementing dynamic multi-streaming according to some embodiments of this specification;
[0013] FIG3 is an exemplary flow chart of a method for implementing dynamic multi-streaming according to some embodiments of this specification;
[0014] FIG4 is an exemplary flowchart of determining device grouping information according to some embodiments of this specification;
[0015] FIG5 is an exemplary flowchart of determining a candidate grouping scheme according to some embodiments of this specification;
[0016] FIG6 is an exemplary flow chart of determining the estimated playback quality of candidate grouping solutions according to some embodiments of this specification;
[0017] FIG. 7 is a schematic diagram of another method for determining estimated playback quality of a candidate grouping solution according to some embodiments of this specification. DETAILED DESCRIPTION
[0018] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.
[0019] It should be understood that the terms "system," "device," "unit," and / or "module" used herein are a method for distinguishing different components, elements, parts, portions, or assemblies at different levels. However, if other terms can achieve the same purpose, the terms may be replaced by other expressions.
[0020] As used in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not refer to the singular but also include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0021] Flowcharts are used throughout this specification to illustrate the operations performed by systems according to embodiments of this specification. It should be understood that preceding or following operations do not necessarily need to be performed in exact order. Instead, the steps may be processed in reverse order or simultaneously. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0022] FIG1 is a schematic diagram of an application scenario of a system for implementing dynamic multi-streaming according to some embodiments of this specification.
[0023] As shown in FIG1 , an application scenario 100 of a dynamic multi-streaming implementation system may include audio devices 110 in various areas, users 120 , user terminals 130 , processors 140 , storage devices 150 , networks 160 , and target audio sources 170 .
[0024] The audio devices 110 in each zone refer to devices installed in each zone (e.g., Zone 1, Zone 2, ..., Zone N) that have the function of playing audio sources. For example, the audio devices may be speakers, music players (iPads, MP3 players, etc.). In some embodiments, the audio devices 110 may also have video playback capabilities, for example, the audio devices may be televisions.
[0025] In some embodiments, the audio device 110 may have one or more speakers. This application does not limit the type of audio device. The audio device 110 may send basic information of the audio device 110 to the processor 140 via the network 160 for subsequent processing. The basic information of the audio device 110 may include a MAC address, account number, device type, device name, etc. In some embodiments, the audio device 110 may also receive and execute operating instructions issued by the processor 140. For example, the audio device 110 may receive a play instruction issued by the processor 140 and play based on the play instruction.
[0026] User 120 refers to a user who uses an audio device. For example, in a home scenario, user 120 may be a family member, etc.; for another example, in a business scenario, user 120 may be a property management personnel, etc. User 120 may issue user instructions through the user terminal 130. For example, user 120 may perform operations such as grouping audio devices, selecting at least one audio device to play music, etc. through the user terminal 130. Among them, user terminal 130 refers to one or more terminal devices or software used by user 120. For example, user terminal 130 may be a device with input and / or output functions, such as a mobile phone and a computer. In some embodiments, user terminal 130 may obtain user input in a variety of ways (such as voice or text, etc.).
[0027] User terminal 130 refers to a terminal device that provides operational and display functions for user interaction. In some embodiments, user terminal 130 may obtain user instructions based on user input or other operations and send the user instructions to storage device 150 and / or processor 140 for storage and / or subsequent processing. User instructions may include grouping instructions, playback instructions, etc. In some embodiments, user instructions may also include user operation instructions. For more information about user operation instructions, please refer to the relevant description of step 340 in Figure 3.
[0028] In some embodiments, the user terminal 130 may also obtain the streaming media group information sent by the processor 140, process the streaming media group information to obtain a full list of audio devices, and then display the full list of audio devices to the user 120 on the interactive interface of the user terminal 130. In some embodiments, the user terminal 130 may also directly obtain the full list of audio devices sent by the processor 140 and display it to the user on the interactive interface.
[0029] The processor 140 may be used to manage data resources and process data and / or information from at least one component involved in the application scenario 100 or an external data source (e.g., a cloud data center). The processor 140 may execute program instructions based on the data, information, and / or processing results, thereby performing one or more functions described in this specification.
[0030] In some embodiments, the processor 140 may receive a grouping instruction sent by the user terminal 130 and generate device grouping information based on the grouping instruction, and then send the device grouping information to the user terminal 130. For more information, please see Figure 3 and its related description. In some embodiments, the processor 140 may receive a play instruction sent by the user terminal 130 and then send the play instruction to the corresponding audio device 110. In some embodiments, the processor 140 may receive a user operation instruction sent by the user terminal 130 and automatically perform audio source pairing and playback based on the user operation instruction.
[0031] In some embodiments, processor 140 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-core processing device). As an example only, processor 140 may include a central processing unit (CPU), a graphics processing unit (GPU), or any combination thereof.
[0032] The storage device 150 can be used to store data and / or instructions. For example, the storage device 150 can be used to store user instructions transmitted by the user terminal 130 via the network 160. For another example, the storage device 150 can also be used to store one or more instruction data issued by the processor 140 to the user terminal 130 or the audio device 110. In some embodiments, the storage device 150 can also store data reported by the audio device 110, such as basic information reported by the audio device 110. In some embodiments, data communication can be performed between the storage device 150 and the processor 140 via the network 160, and the storage device 150 can also be part of the processor 140.
[0033] In some embodiments, the storage device 150 may include random access memory (RAM), read only memory (ROM), mass storage, the like, or any combination thereof.
[0034] In some embodiments, one or more components of the application scenario 100 may transmit data to other components of the application scenario 100 via the network 160. For example, the processor 140 may obtain information and / or data from the user terminal 130, the audio device 110, or the storage device 150 via the network 160, or may send information and / or data to the user terminal 130 or the storage device 150 via the network 160. In some embodiments, one or more components of the application scenario 100 may also directly communicate data with each other.
[0035] In some embodiments, target audio source 170 refers to an audio source that the user desires to play, such as music, recordings, etc. In some embodiments, one or more audio playback devices in each zone can play the target audio source. For more information about the target audio source, please refer to the relevant description in Figure 3.
[0036] FIG2 is a module diagram of a system for implementing dynamic multi-streaming according to some embodiments of this specification.
[0037] In some embodiments, the dynamic multi-streaming implementation system 200 may include a grouping instruction generation module 210 , a grouping module 220 , a play instruction generation module 230 , and a play module 240 . In some embodiments, the dynamic multi-streaming implementation system 200 may be integrated into the processor 140 .
[0038] In some embodiments, the grouping instruction generation module 210 is configured to generate a grouping instruction based on a first user operation, where the first user operation is related to a grouping requirement of the user.
[0039] In some embodiments, the grouping module 220 is configured to group the audio devices in the target area based on the grouping instruction to obtain device grouping information.
[0040] In some embodiments, the grouping module 220 is further configured to group the audio devices in the target area based on the grouping instruction to obtain first grouping information; and to determine first streaming media group information based on the first grouping information. The grouping module 220 is further configured to obtain a first control instruction from the user, the first control instruction including update information for the first streaming media group information. The grouping module 220 is further configured to update the first streaming media group information based on the first control instruction to obtain second streaming media group information. The grouping module 220 is further configured to determine device group information based on the second streaming media group information.
[0041] In some embodiments, the grouping module 220 is further configured to generate a full list of audio devices based on the first streaming media group information; display the full list information to the user; and determine a first control instruction based on the user's adjustment information on the full list information.
[0042] In some embodiments, the user's grouping requirements include a content feature set and matching information corresponding to the content feature set, where the content feature set includes one or more content features. The grouping module 220 is further configured to generate one or more candidate grouping schemes based on the user's grouping requirements; the candidate grouping schemes include content features and audio devices that play the content features.
[0043] In some embodiments, the grouping module 220 is further configured to determine the estimated playback quality of the candidate grouping scheme based on historical playback information of the audio devices in the candidate grouping scheme.
[0044] In some embodiments, the grouping module 220 is further configured to determine a target grouping scheme based on the estimated playback quality; and determine device grouping information based on the target grouping scheme.
[0045] In some embodiments, the matching information corresponding to the content feature set includes the target playback area of the content feature and the device requirement information corresponding to the content feature. In some embodiments, the grouping module 220 is further configured to determine the device quantity range of the target devices corresponding to the content feature set based on the number of content features in the content feature set and the device requirement information corresponding to the content features. In some embodiments, the grouping module 220 is further configured to determine the number of target devices to be enabled based on the device quantity range. In some embodiments, the grouping module 220 is further configured to determine the audio devices corresponding to the number to be enabled as the audio devices to be enabled. In some embodiments, the grouping module 220 is further configured to determine a set of first matching relationships that meet preset matching conditions based on the target playback area of the content feature; the first matching relationship includes each content feature corresponding to one or more audio devices to be enabled.
[0046] In some embodiments, the grouping module 220 is further configured to generate streaming media to be matched that meets a preset number requirement; the preset number requirement includes that the number of streaming media to be matched is not less than the number of content features in the content feature set. In some embodiments, the grouping module 220 is further configured to determine a correspondence between the streaming media to be matched and the content features to obtain a second matching relationship; and to generate one or more streaming media groups. In some embodiments, the grouping module 220 is further configured to determine a correspondence between the streaming media to be matched and the streaming media groups to obtain a third matching relationship. In some embodiments, the grouping module 220 is further configured to obtain a candidate grouping scheme based on the first matching relationship, the second matching relationship, and the third matching relationship.
[0047] In some embodiments, the historical playback information of the audio device includes the signal transmission delay and packet loss rate during the historical playback. In some embodiments, the grouping module 220 is further configured to determine the signal transmission quality of the audio device based on the signal transmission delay and packet loss rate during the historical playback of the audio device; and to determine the estimated playback quality of the candidate grouping scheme based on the signal transmission quality of each audio device in the candidate grouping scheme.
[0048] In some embodiments, the grouping module 220 is further configured to determine multiple sub-evaluation values of the audio device based on the signal transmission delay and packet loss rate during multiple historical playbacks of the audio device; and determine the signal transmission quality of the audio device based on the weighted sum of the multiple sub-evaluation values.
[0049] In some embodiments, when multiple sub-evaluation values are weighted and summed, the weight of the sub-evaluation value is negatively correlated with the time interval between the occurrence time of the historical playback corresponding to the sub-evaluation value and the current time.
[0050] In some embodiments, the grouping module 220 is further configured to determine the estimated playback quality of the candidate grouping scheme based on the candidate grouping scheme, the signal transmission quality of each audio device in the candidate grouping scheme, and the first evaluation information of each audio device through a quality prediction model; the quality prediction model is a machine learning model; the first evaluation information includes the signal transmission delay and packet loss rate of the audio device at multiple time points within a preset time period.
[0051] In some embodiments, the play instruction generation module 230 is configured to generate a play instruction based on the device grouping information and the user's second operation, where the user's second operation is related to the user's play requirement, which includes a target sound source.
[0052] In some embodiments, the playing module 240 is configured to control the target device to play the target audio source based on the playing instruction.
[0053] In some embodiments, each target device plays a corresponding channel audio signal.
[0054] In some embodiments, the target audio source includes a mixed audio source obtained by mixing at least two streaming media types of audio sources.
[0055] In some embodiments, the dynamic multi-stream implementation system 200 also includes an instruction update module, which is configured to determine an audio source update instruction based on a third operation of the user; the third operation of the user is related to the user's playback update requirement, and the playback update requirement includes a target device group and an updated audio source of the target device group; and the playback module 240 is also configured to control the target device group to play the updated audio source based on the audio source update instruction.
[0056] In some embodiments, the playback module 240 is further configured to determine a user operation instruction based on the user's first operation and / or the user's second operation; the user operation instruction includes a target sound source and a target playback area; and based on the user operation instruction, pair the sound source and determine the target device.
[0057] In some embodiments, the playback module 240 is further configured to pair the target audio source with the target transmitting channel; pair the target transmitting channel with the target receiving channel of the target playback area; and determine the target device based on the target receiving channel.
[0058] It should be noted that the above description of the dynamic multi-stream implementation system and its modules is for convenience of description only and does not limit this specification to the scope of the embodiments cited. It is understandable that for those skilled in the art, after understanding the principle of the system, it is possible to arbitrarily combine the various modules, or form a subsystem to connect with other modules without deviating from this principle. In some embodiments, the grouping instruction generation module 210, grouping module 220, play instruction generation module 230 and play module 240 disclosed in Figure 1 can be different modules in a system, or a module can implement the functions of two or more of the above modules. For example, each module can share a storage module, or each module can have its own storage module. Such variations are all within the scope of protection of this specification.
[0059] Figure 3 is an exemplary flow chart of a method for implementing dynamic multi-streaming according to some embodiments of this specification. In some embodiments, process 300 may be executed by processor 140 or dynamic multi-streaming implementation system 200. As shown in Figure 3, process 300 includes the following steps 310 to 360.
[0060] Step 310: Generate a grouping instruction based on the user's first operation.
[0061] The first user operation refers to a physical operation performed by the user on the user terminal related to audio device grouping. For example, the first user operation may be an operation in which the user inputs a grouping request on the user terminal through various input windows of an application, mini-program, webpage, etc. on the terminal device in various ways (including but not limited to text, voice, touch screen, etc.).
[0062] In some embodiments, the processor may obtain the user's first operation through the user's terminal device. In some embodiments, the user's first operation is related to the user's grouping requirement.
[0063] Grouping requirements refer to the user's expected grouping of audio devices in the target area.
[0064] In some embodiments, the grouping requirements may include the number of groups in the target area and the audio devices included in each group.
[0065] A target area refers to the entire or partial spatial area where the audio device is deployed. For example, if the audio device is deployed in a residential location, the target area may be the entire spatial area where the user lives, or a portion of a room within it. For another example, if the audio device is deployed in a commercial location, the target area may be the entire shopping mall or a portion of a floor within it. In some embodiments, the location of the audio device within the target area may be fixed or movable.
[0066] In some embodiments, the processor may divide the target area into one or more sub-areas, so that the user can select an audio device in a corresponding sub-area for playback. For example, when the target area is the entire space where the user lives, the sub-areas may include the bedroom, living room, bathroom, balcony, kitchen, etc. For another example, when the target area is the entire shopping mall where the user lives, the sub-areas may include each store, bathroom, rest area, etc. in the mall.
[0067] In some embodiments, a sub-zone may include one or more audio devices.
[0068] For example, assuming that there are 10 audio devices in the target area (recorded as audio devices 1 to 10), the entire target area (taking the residential scene as an example) is divided into 3 sub-areas (bedroom area, living room area, and bathroom area), among which audio devices 1 to 2 are deployed in the bedroom area, audio devices 3 to 8 are deployed in the living room area, and audio devices 9 to 10 are deployed in the bathroom area.
[0069] In some embodiments, the sub-areas of the target area and the groups of audio devices can correspond one to one. For example, a user can group audio devices 1-2 deployed in the bedroom area into group A, audio devices 3-8 deployed in the living room area into group B, and audio devices 9-10 deployed in the bathroom area into group C.
[0070] In some embodiments, the sub-areas of the target area and the groups of audio devices can have a one-to-many relationship. For example, if audio devices 3 to 8 are deployed in the living room, the user can divide audio devices 3 to 8 into multiple groups, such as audio devices 3 to 4 into group B1 and audio devices 5 to 8 into group B2.
[0071] In some embodiments, the sub-areas of the target area and the groups of audio devices can have a many-to-one relationship. For example, a user can select at least one audio device from each of at least two sub-areas to form at least one group. Continuing with the previous example, the user can group audio device 1 in the bedroom area and audio device 3 in the living room area into a group D.
[0072] In some embodiments, the grouping requirement may further include other information related to the grouping of audio devices. For more details, refer to the corresponding content of FIG. 4 .
[0073] In some embodiments, the grouping requirement can be determined based on the user's specific audio playback requirements and / or the spatial layout of the target area, and the user's specific audio playback requirements can be determined by the user's first operation. For example, assuming that there are 10 audio devices in the target area (recorded as audio devices 1 to 10), the target area includes the bedroom, living room and kitchen, and the user's audio playback requirement is to play "Rice Fragrance" on at least 3 audio devices in the living room and "Nocturne" on at least 2 audio devices in the bedroom, then the grouping requirement can be based on the playback requirements and / or the spatial layout of the target area. The audio devices are divided into three groups, the first group including audio device 1 and audio device 2 in the bedroom, the second group including audio devices 3 to 8 in the living room, and the third group including audio devices 9 to 10 in the kitchen.
[0074] The grouping command is a command for grouping all audio devices in a target area.
[0075] In some embodiments, the processor may generate a grouping instruction based on the grouping requirement determined by the user's first operation. For example, the processor may generate a corresponding grouping instruction based on the number of groups required in the grouping requirement and the audio devices included in each group.
[0076] Step 320: Group the audio devices in the target area based on the grouping instruction to obtain device grouping information.
[0077] For the description of audio devices, please refer to the relevant description in Figure 1.
[0078] Device grouping information refers to information about the audio device groups that are ultimately displayed to the user. For example, the device grouping information may be the three groups of audio devices mentioned above: the first group includes audio device 1 and audio device 2 in the bedroom, the second group includes audio devices 3-8 in the living room, and the third group includes audio devices 9-10 in the kitchen.
[0079] In some embodiments, the process of the processor grouping audio devices in the target area based on the grouping instruction to obtain device grouping information includes:
[0080] Step S10: grouping the audio devices in the target area based on the grouping instruction to obtain first grouping information.
[0081] The first grouping information refers to information related to the grouping of audio devices in the target area. For example, the first grouping information may include the number of audio devices in the target area, the group names of the audio devices, and the streaming media types supported by each audio device group. The streaming media types may include AirPlay 2, Spotify, Roon, DLNA, Airable, etc.
[0082] In some embodiments, the processor may group the audio devices in the target area based on the grouping instruction, and obtain first grouping information corresponding to one or more groups of audio devices after grouping.
[0083] Step S20: Determine first streaming media group information based on the first grouping information.
[0084] The first streaming media group information refers to the streaming media group information corresponding to the first grouping information.
[0085] Streaming media group information refers to information related to the grouping of multiple audio devices to play audio sources. For example, the streaming media group information may include basic information about the audio devices included in each audio device group.
[0086] In some embodiments, the streaming media group information includes one or more groups of streaming media information. A group of streaming media information refers to information related to the audio source played by a single audio device group. For example, a group of streaming media information may include an audio device group name and basic information about one or more audio devices in the corresponding group.
[0087] One set of streaming media information corresponds to one audio device group. For example, all audio devices in the target area are divided into four groups (recorded as audio device groups 1 to 4). Each audio device group generates a corresponding set of streaming media information, resulting in a total of four sets of streaming media information (recorded as streaming media information groups 1 to 4).
[0088] In some embodiments, each set of streaming media information includes one or more sub-information items, each of which corresponds to a target streaming media type.
[0089] Streaming media types may include airplay2, spotify, roon, DLNA, airable, etc. The target streaming media type refers to the union of the streaming media types supported by all audio devices in the audio device group corresponding to the set of streaming media information.
[0090] In some embodiments, the streaming media types supported by the same group of audio devices can be the same. For example, if the first group of audio devices includes audio devices 1-3, and audio devices 1-3 all support five streaming media types (AirPlay2, Spotify, Roon, DLNA, and Airable), then the target streaming media types include the five streaming media types (AirPlay2, Spotify, Roon, DLNA, and Airable). This group of streaming media information (i.e., the first group of streaming media information) contains five sub-information items, each corresponding to the five streaming media types in the target streaming media types.
[0091] In some embodiments, the streaming media types supported by the same group of audio devices can be different. In this case, the target streaming media type is the union of the streaming media types supported by each audio device in the group. For example, if group 1 includes audio devices 1 through 3, audio device 1 supports AirPlay 2 and Spotify, audio device 2 supports Spotify and Roon, and audio device 3 supports Airable, then the target streaming media types are AirPlay 2, Spotify, Roon, and Airable. This group of streaming media information then contains four sub-information items, one corresponding to each of the four target streaming media types.
[0092] In some embodiments, the content contained in each sub-information may include the supported target streaming media type, basic information of one or more audio devices in the corresponding group, etc. In some embodiments, the basic information of the audio device may include the MAC address, account number, device type, device name, etc. Among them, the audio device name may include the name of the sub-area where the audio device is located. For example, the audio device name may be "Audio Device 1-Group A-Bedroom". In some embodiments, the content contained in each sub-information may also include programs corresponding to one or more services (for example, device discovery, encoding, decoding, etc.) required to implement the audio device to play the audio source.
[0093] In some embodiments, each sub-information in the streaming media group information corresponds to a complete function, which may include multiple sub-functions. In some embodiments, the complete function may be divided into a device function of the audio device and an audio playback function. Based on the playback status of the target audio source, the processor may call and execute the program corresponding to the sub-function required to play the target audio source from the storage device 150.
[0094] For example, assuming that the streaming media group information includes 4 groups of streaming media information (corresponding to 4 audio device groups), each group of streaming media group information includes 5 sub-information (corresponding to 5 streaming media types), and the device functions of the audio device corresponding to each sub-information (such as discovery function and connection function, etc.) and the audio playback function require 100 megabytes of running space. Then, directly running the device functions and audio playback functions of all audio devices corresponding to 20 sub-information requires a total running space of 2000 megabytes.
[0095] In some embodiments of this specification, the execution of the device function and audio playback function of the audio device corresponding to each sub-information is divided into two stages. For example, the first stage is the execution of the audio device's device functions (such as the discovery function and the connection function), and the second stage is the execution of the audio playback function after the audio device is connected. If the running space of the first stage function of an audio device is 10 megabytes and the running space of the second stage function is 90 megabytes, when all audio devices are running, each audio device will first run the first stage function (a total running space of 20*10=200 megabytes). Then, after determining which audio devices need to be connected and the connection is established, only the second stage function of the audio devices that need to be connected (e.g., 5 audio devices) will be executed (a total running space of 5*90=450 megabytes). This saves space (i.e., the second stage function will not be executed for unused audio devices).
[0096] In some embodiments of this specification, when playing a target audio source, the operation of the device function and the audio playback function of the audio device are divided into two sections, and the device functions of all audio devices are fully operated. After the corresponding audio device is connected, the audio playback function will wait until the actual demand arrives before starting, thereby realizing multi-stream resource optimization and saving computing resources.
[0097] In some embodiments, the processor may determine, based on the first grouping information, the streaming media group information corresponding to each group of audio devices in the first grouping information, thereby obtaining the first streaming media group information.
[0098] Step S30: Obtain the user's first control instruction.
[0099] The first control instruction refers to an instruction for adjusting the first streaming media group information.
[0100] In some embodiments, the first control instruction includes update information of the first streaming media group information.
[0101] The update information of the first streaming media group information refers to information for modifying, supplementing or replacing the first streaming media group information.
[0102] In some embodiments, the user can perform operations such as dragging or sliding on the interactive interface of the terminal device to issue a first control instruction, and the processor obtains the user's first control instruction through the terminal device.
[0103] In some embodiments, the processor may also generate full list information of audio devices based on the first streaming media group information; display the full list information to the user; and determine the first control instruction based on the user's adjustment information of the full list information.
[0104] The full list of audio devices refers to the relevant information of the audio devices that can be adjusted and displayed to the user.
[0105] In some embodiments, the full list information includes one or more sub-data, each sub-data corresponds to a sub-information in a group of streaming media information, and different sub-data in the full list information corresponds to different sub-information in the streaming media group information.
[0106] In some embodiments, each sub-data in the full list information may include the name of the corresponding audio device group and a supported target streaming media type. The name of the corresponding audio device group includes the name of the sub-area in which the audio device in the group is located. In some embodiments, if multiple audio devices in the audio device group are in the same sub-area (such as the living room area), the name of the audio device group may be "Group A-Living Room"; if multiple audio devices are in different sub-areas (for example, the living room area and the bedroom area), the name of the audio device group may be "Group A-Living Room and Bedroom".
[0107] In some embodiments, the processor may generate full list information of audio devices based on the first streaming media group information.
[0108] In some embodiments, the processor may send the full list information of the determined audio devices to the user terminal, and then display the full list information to the user on the interactive interface of the user terminal.
[0109] The adjustment information of the full list information refers to the adjustment information used to adjust the first streaming media group information.
[0110] In some embodiments, the user can send adjustment information of the entire list information by dragging or sliding any set of first streaming media group information on the interactive interface of the terminal device, and the processor obtains the adjustment information of the entire list information through the terminal device.
[0111] In some embodiments, the processor may generate a first control instruction based on the user's adjustment information on the full list information.
[0112] In some embodiments of this specification, a full list of information is displayed to the user for easy understanding. The user can then adjust the grouping by directly dragging or other actions on the interactive interface of the user terminal, thereby reducing the difficulty of operation and improving the user experience.
[0113] Step S40: Based on the first control instruction, update the first streaming media group information to obtain the second streaming media group information.
[0114] The second streaming media group information refers to the updated first streaming media group information.
[0115] In some embodiments, the processor may adjust the first streaming media group information based on the first control instruction to obtain updated first streaming media group information as the second streaming media group information.
[0116] Step S50: Determine device grouping information based on the second streaming media group information.
[0117] In some embodiments, the processor may determine one or more corresponding device groups based on the second streaming media group information, and further determine one or more device group information.
[0118] In some embodiments, the processor may also determine the device grouping information based on other methods, as shown in FIG. 4 for details.
[0119] Step 330: Generate a play instruction based on the device group information and the second user operation.
[0120] The second user operation refers to a physical operation performed by the user on the user terminal related to the audio device playback. For example, the second user operation can be an operation in which the user inputs a playback request on the user terminal through various means (including but not limited to text, voice, touch screen, etc.).
[0121] In some embodiments, the processor may obtain the second user operation through the user's terminal device. In some embodiments, the second user operation is related to the user's playback demand.
[0122] Playback requirements refer to user requirements for audio source playback. For example, playback requirements include the playback time and duration.
[0123] In some embodiments, the playback request may further include a target audio source.
[0124] The target audio source refers to the audio source that the user expects to play, such as music, recordings, etc.
[0125] In some embodiments, the target audio source may also be a video source.
[0126] In some embodiments, the target audio source may be a single audio source of a streaming type, for example, a piece of music played on airplay2.
[0127] In some embodiments, the target audio source may also be a mixed audio source comprising at least two streaming media types, for example, a song from AirPlay 2 and a song from Spotify mixed together.
[0128] In some embodiments, the processor may directly mix at least two streaming media types of audio sources to obtain a mixed audio source, and send the mixed audio source as a target audio source to a target device for playback.
[0129] The target device may be one or more audio devices that will play the audio source. For more information about the target device, see step 340.
[0130] If the target device includes multiple audio devices, the multiple audio devices start playing the target sound source at the same time point, so that the multiple audio devices play the target sound source simultaneously.
[0131] In some embodiments of the present specification, a processor mixes at least two types of streaming media audio sources to obtain a mixed audio source, and then each audio device in the target device receives and plays the mixed audio source, thereby ensuring that the mixed audio source reaches each audio device in the target device at the same time, thereby ensuring consistency in playback across multiple audio devices in the target device.
[0132] In some embodiments, each audio device in the target device may mix the received audio sources of at least two streaming media types to obtain a mixed audio source and play the mixed audio source.
[0133] In some embodiments of the present specification, since the computing power of each audio device in the target device is consistent, the consistency of playback of multiple audio devices in the target device is ensured.
[0134] In some embodiments, when the target audio source includes multi-channel audio signals, the playback request may further include the audio signals of the corresponding playback channels of each target device. For example, when the target audio source is a stereo source, which includes left and right channel audio signals, the playback request may be to have some of the audio devices in the target devices play the left channel audio signal, and have the remaining audio devices in the target devices play the right channel audio signal.
[0135] For another example, when the target sound source is a surround sound source, such as a 5.1 surround sound source including six channels of audio signals, the playback requirement may be to have a specified audio device in the target device play the audio signal of a specified channel.
[0136] In some embodiments of the present specification, by assigning each target device to a corresponding audio signal of a playback channel, different needs of different users can be better met, thereby improving user experience.
[0137] Playback commands are commands used to control a target device to play a target audio source.
[0138] In some embodiments, the processor can determine the target audio source for playback based on the user's second operation; determine the streaming media type corresponding to the target audio source for playback based on the target audio source; determine the target device for playback based on the streaming media type corresponding to the target audio source and the device grouping information; and determine the playback instruction based on the target audio source for playback and the target device for playback.
[0139] In some embodiments, the method by which the processor determines the type of streaming media corresponding to the target audio source is similar to the method for determining the media type of the streaming media to be matched in the target grouping scheme, see the description of step 590 in FIG. 5 for details.
[0140] In some embodiments, the processor may select a matching audio device from the device grouping information based on the streaming media type corresponding to the target audio source, and determine the device as the target device for playback.
[0141] Step 340: Based on the play instruction, control the target device to play the target audio source.
[0142] In some embodiments, the target device may be one or more preset audio devices that can support playback of the target audio source.
[0143] In some embodiments, the processor can determine the target device based on a variety of ways. In some embodiments, the target device can be a device specified by the user. For example, the target device can be all or part of the audio equipment in the living room specified by the user.
[0144] In some embodiments, the target device can be determined based on the target grouping scheme and playback requirements. For example, a corresponding number of audio devices (the audio devices need to be able to support the playback of the target sound source) are randomly determined within the corresponding group as target devices for playback. For more information about the target grouping scheme, please refer to the corresponding content of Figure 4.
[0145] In some embodiments, the processor may determine a user operation instruction based on the user's first operation and / or the user's second operation; and perform audio source pairing and determine a target device based on the user operation instruction.
[0146] User operation instructions refer to operation instructions related to playing audio sources issued by users on user terminals.
[0147] In some embodiments, the user operation instruction may include a target sound source and a target playback area. For an explanation of the target playback area, please refer to the description of step 410 in FIG. 4 .
[0148] In some embodiments, the processor may directly determine the user operation instruction based on the first user operation and / or the second user operation.
[0149] In some embodiments, the processor performs audio source pairing and determines a target device based on a user operation instruction. The specific implementation method can be implemented using the following steps h10 to h30:
[0150] Step h10: Pair the target audio source with the target sending channel.
[0151] The target sending channel refers to the sending channel that is ultimately paired with the target audio source.
[0152] In some embodiments, the processor may obtain pairing information of all channels of the current sending channel and the number of channels required by the target sound source, and then pair the target sound source with the target sending channel based on the aforementioned information.
[0153] In some embodiments, the pairing information of all channels of the sending channel may include a paired audio source for each sending channel. A paired audio source refers to an audio source bound to a sending channel.
[0154] In some embodiments, when there is an idle audio source among the paired audio sources, the processor can unbind the idle audio source and release the corresponding transmission channel. An idle audio source refers to an audio source that is not in use or is not being played.
[0155] In some embodiments, when the paired audio sources do not include the target audio source and the number of available transmission channels is not less than the number of channels required by the target audio source, the processor may determine the target transmission channel from the available transmission channels and pair the target audio source with the target transmission channel. The available transmission channels refer to the remaining unpaired transmission channels among the current transmission channels.
[0156] In some embodiments, when the number of available transmission channels is less than the number of channels required by the target audio source, the processor may issue a prompt to the user, such as "the current playback channels are insufficient".
[0157] Step h20: Pair the target sending channel with the target receiving channel of the target playback area.
[0158] The target receiving channel refers to the receiving channel that is ultimately paired with the target sending channel.
[0159] In some embodiments, the processor may pair the target transmitting channel with the target receiving channel of the target playback area according to a preset channel pairing rule. The preset channel pairing rule refers to a pairing rule for the target transmitting channel and the target receiving channel of the target playback area, and may be preset by a person skilled in the art based on experience.
[0160] In some embodiments, after the target audio source is paired with the target sending channel, the processor may obtain pairing information of all channels in the current receiving channel.
[0161] The pairing information for all channels of the receiving channel may include the transmitting channel paired with each receiving channel. Based on user operations, the partition information of the target area (each sub-area and the audio devices in the sub-area), and the pairing information for all channels of the transmitting channel, the processor can obtain available receiving channels from the receiving channel, determine the target receiving channel from the available receiving channels, and then pair the target receiving channel with the target transmitting channel. The available receiving channels refer to the remaining unpaired receiving channels in the current receiving channel.
[0162] Step h30: Determine the target device based on the target receiving channel.
[0163] In some embodiments, the processor may select one or more audio devices corresponding to the target receiving channel as target devices.
[0164] In some embodiments, the user pre-play processor determines the adjustments required for the current play operation and completes the audio source pairing operation through the operations of steps h10 to h30, thereby avoiding confusion in audio source pairing.
[0165] In addition, users only need to input the target audio source and the area or device where the target audio source is played. The processor can automatically pair the target audio source with the sending channel, the sending channel with the receiving channel, and the receiving channel with the audio device, meeting user needs while improving user experience.
[0166] In some embodiments, the processor may control the target device to play the target audio source based on the target device and the target audio source displayed by the play instruction.
[0167] In some embodiments, the method for implementing dynamic multi-streaming may further include the following steps 350 to 360 .
[0168] Step 350: Determine the audio source update instruction according to the third user operation.
[0169] The third user operation refers to a physical operation performed by the user on the user terminal related to updating the audio source played by the audio device. For example, the third user operation can be an operation in which the user inputs a playback update request on the user terminal through various means (including but not limited to text, voice, touch screen, etc.).
[0170] In some embodiments, the processor may obtain the user's third operation through the user's terminal device.
[0171] In some embodiments, the third user operation may be related to the user's playback update requirement.
[0172] In some embodiments, the playback update request includes a target device group and an updated audio source of the target device group.
[0173] The updated sound source refers to another sound source different from the target sound source.
[0174] In some embodiments, the target device group may include multiple target devices. After determining to update the audio source and the target device group based on the user's third operation, if an audio source is playing in the target device group, the updated audio source may replace the originally playing audio source, causing the target device group to play the updated audio source.
[0175] In some embodiments, multiple target devices in a target device group can be user-specified audio devices. For example, multiple audio devices in the same target device group can play various types of audio sources, including any combination of these types. For another example, each time a sound is played, at least one audio device in the specified target device group participates in the playback. For example, all audio devices can play the sound source, or only one audio device can play the sound source.
[0176] The audio source update command can be used to adjust the audio source played by the current target device group.
[0177] In some embodiments, the processor may directly determine the audio source update instruction according to the user's third operation.
[0178] Step 360: Based on the audio source update instruction, control the target device group to play the updated audio source.
[0179] In some embodiments, the processor may send an audio source update instruction to the target device to control the target device group to play the updated audio source.
[0180] In some embodiments of the present application, users can dynamically group all audio devices in a target area based on grouping requirements and dynamically create device grouping information that includes streaming media group information, allowing users to freely select and play various streaming media types of audio / video sources. The device grouping information ultimately displayed to the user includes the name of the audio device in the partition name, meeting user needs while improving the user experience. Furthermore, users can update the audio source played by the target device at any time to meet their changing playback needs.
[0181] FIG4 is an exemplary flow chart of determining device grouping information according to some embodiments of this specification. In some embodiments, process 400 may be executed by processor 140 or dynamic multi-flow implementation system 200. As shown in FIG4, process 400 includes the following steps 410 to 440.
[0182] Step 410: Generate one or more candidate grouping solutions based on the user's grouping requirements.
[0183] In some embodiments, the user's grouping requirements may further include a content feature set and matching information corresponding to the content feature set.
[0184] In some embodiments, a content feature set includes one or more content features.
[0185] Content features refer to the content that the user wants to listen to. For example, if the user wants to listen to two songs, "Nocturne" and "Rice Fragrance", then "Nocturne" is one content feature and "Rice Fragrance" is another content feature.
[0186] In some embodiments, each content feature corresponds to at least one selectable streaming media type.
[0187] Optional streaming media types refer to the streaming media types that can be selected according to the content characteristics. For example, if the song "Nocturne" mentioned above includes two streaming media types of audio sources, Spotify and Roon, then the optional streaming media types for "Nocturne" are Spotify and Roon.
[0188] In some embodiments, the content feature set also includes the number of channels corresponding to each content feature under the corresponding optional streaming media type, etc. The number of channels corresponding to the streaming media type can be a default value or can be preset by those skilled in the art based on experience.
[0189] The matching information corresponding to the content feature set refers to information that matches the playback of each content feature in the content feature set, such as an optional streaming media type corresponding to each content feature in the content feature set.
[0190] In some embodiments, the matching information corresponding to the content feature set may further include a target playback area of the content feature and device requirement information corresponding to the content feature.
[0191] The target playback area refers to the area where the user needs to play the audio. For example, if the user needs to play audio 1 in the living room and audio 2 in the living room and bedroom, the target playback area is the living room and bedroom.
[0192] Device requirement information refers to the number of devices required for different content features in different planned playback areas. For example, if a user requires two audio devices to play audio 1 in the living room and one audio device to play audio 2 in the living room, the device requirement information includes that audio 1 requires two audio devices in the living room and audio 2 requires one audio device in the living room.
[0193] For instructions on how to obtain the user's grouping requirements, please refer to the instructions in step 310 of FIG. 3 .
[0194] The candidate grouping scheme refers to a grouping scheme of audio devices in a candidate target area.
[0195] In some embodiments, the total number of streaming media groups in different candidate grouping schemes may be different, so the consumption of system performance for implementing dynamic multi-streaming may be different, resulting in differences in playback quality.
[0196] In some embodiments, the total number of streaming media may be different in different candidate grouping schemes, that is, the total number of virtual channels divided out may be different. Therefore, the interference between multiple content features and the consumption of system performance for implementing dynamic multi-streaming may be different, resulting in differences in playback quality.
[0197] In some embodiments, the number of streaming media included in each streaming media group may be different in different candidate grouping schemes. Therefore, different audio devices may be in different groups, which may also lead to differences in playback quality.
[0198] In addition, since the more consistent (similar) the playback quality of multiple audio devices in the same group is, the better the playback quality is, if the playback quality of two audio devices in the same sub-group is too different, audio mixing problems may occur, which may also lead to differences in playback quality of different candidate grouping schemes.
[0199] In some embodiments, the candidate grouping schemes include content characteristics and audio devices that play the content characteristics determined based on the grouping requirements.
[0200] In some embodiments, the processor may randomly generate one or more candidate grouping solutions based on the user's grouping requirements.
[0201] In some embodiments, the processor may also determine candidate grouping schemes based on other methods, as shown in FIG. 5 for details.
[0202] Step 420 : Determine the estimated playback quality of the candidate grouping scheme based on the historical playback information of the audio devices in the candidate grouping scheme.
[0203] Historical playback information refers to information related to the historically played audio source, such as the number of times the audio source has been played.
[0204] In some embodiments, the historical playback information of the audio device includes the signal transmission delay and packet loss rate during the historical playback.
[0205] Signal transmission latency refers to the time it takes from the processor issuing a play command to the audio device playing the audio source. In some embodiments, the processor can obtain signal transmission latency through various methods (e.g., the Ping command, network performance testing tools, etc.). The Ping command is a network diagnostic tool that can send ICMP echo request packets to the target audio device and measure their return time to estimate signal transmission latency. Network performance testing tools may include iperf, SPeedtest, etc.
[0206] The packet loss rate refers to the ratio of the number of data packets lost in signal transmission to the number of data packets sent. In some embodiments, the processor can use multiple methods (such as Ping command, tracert command, etc.) to determine the packet loss rate.
[0207] The estimated playback quality refers to the estimated quality of the audio device playing the audio in the future. In some embodiments, the estimated playback quality can be expressed as a percentage.
[0208] In some embodiments, the processor may determine the estimated playback quality of the candidate grouping scheme based on historical playback information of the audio devices in the candidate grouping scheme using a first preset comparison table. The first preset comparison table includes a correspondence between the historical playback information of the audio devices in the reference grouping scheme and the estimated playback quality of the reference grouping scheme. The first preset comparison table may be constructed based on prior knowledge or historical data.
[0209] In some embodiments, the processor may also determine the estimated playback quality of the candidate grouping schemes based on other methods, as shown in FIG. 6 for details.
[0210] Step 430: Determine a target grouping scheme based on the estimated playback quality.
[0211] The target grouping scheme refers to the ultimately determined grouping scheme for audio devices in the target area.
[0212] In some embodiments, the processor may directly use the candidate grouping scheme with the highest estimated playback quality as the target grouping scheme.
[0213] In some embodiments, the processor can also sort the estimated playback quality of all candidate grouping schemes from high to low; then send multiple (e.g., 3) candidate grouping schemes with the highest ranking to the user's terminal device for the user to select; and finally determine the candidate grouping scheme finally confirmed by the user as the target grouping scheme.
[0214] Step 440: Determine device grouping information based on the target grouping scheme.
[0215] For description of device grouping information, please refer to the relevant description in step 320 of FIG. 3 .
[0216] In some embodiments, the processor may determine the device grouping information corresponding to the target grouping scheme as the device grouping information.
[0217] In some embodiments, the processor may further update the candidate grouping scheme in response to a target device failure, and re-determine the target grouping scheme according to the aforementioned steps 410 to 440 .
[0218] In some embodiments, the processor may eliminate faulty audio devices in the target area, and then regenerate one or more candidate grouping schemes for the remaining audio devices in the target area based on the user's grouping requirements.
[0219] In some embodiments of this specification, the audio devices in the target area are automatically grouped based on the user's needs, and the device grouping information is automatically determined. Then, based on the user's playback needs, the target device is automatically controlled to play the target audio source, thereby automatically and quickly meeting the user's playback needs and improving the user experience.
[0220] In addition, when the playback instructions are generated based on the automatically determined device grouping information and user playback needs to control the target device to play the target audio source, if the target device suddenly fails, the faulty target device can be eliminated, the device grouping information can be automatically re-determined, and then based on the user's playback needs, the target device can be automatically controlled to play the target audio source again, thereby automatically and quickly realizing the user's playback needs and further improving the user experience.
[0221] It should be noted that the above description of process 400 is for illustration and purpose only and does not limit the scope of application of this specification. Those skilled in the art may make various modifications and variations to process 400 under the guidance of this specification. However, such modifications and variations are still within the scope of this specification.
[0222] Figure 5 is an exemplary flow chart of a method for determining a candidate grouping scheme according to some embodiments of this specification. In some embodiments, process 500 may be executed by processor 140 or dynamic multi-flow implementation system 200. As shown in Figure 5, process 500 includes the following steps 510-590.
[0223] Step 510: Determine the device quantity range of the target devices corresponding to the content feature set based on the quantity of content features in the content feature set and the device requirement information corresponding to the content features, wherein the device quantity range includes the minimum activation value and the maximum activation value of the target devices.
[0224] For descriptions of the content feature set, content features, and device requirement information, please refer to the description of step 410 in FIG4 . For descriptions of the target device, please refer to the description of step 330 in FIG3 .
[0225] In some embodiments, the processor may directly determine the device quantity range of the target devices corresponding to the content feature set based on the number of content features in the content feature set and the device requirement information corresponding to the content features. For example, if the number of content features in the content feature set is 3 and the device requirement information corresponding to the content features is no requirement, then the minimum enabled value of the target devices corresponding to the content feature set is determined to be 3, and the maximum enabled value is determined to be the total number of audio devices in the target area.
[0226] For another example, if the number of content features in the content feature set is 3, and the device requirement information corresponding to the content features is that each content feature requires at least 2 audio devices to play together, then the minimum activation value of the target device corresponding to the content feature set is determined to be 6, and the maximum activation value is the total number of audio devices in the target area.
[0227] Step 520: Determine the number of target devices to be enabled based on the device quantity range.
[0228] The number of target devices to be enabled refers to the actual number of enabled target devices.
[0229] In some embodiments, the processor may randomly select a value within a range of device quantities and determine the value as the number of target devices to be enabled.
[0230] Step 530: Determine the audio devices corresponding to the number to be enabled as the audio devices to be enabled.
[0231] In some embodiments, the processor may randomly determine the audio devices corresponding to the number to be enabled as the audio devices to be enabled.
[0232] The audio device to be enabled is the audio device that is ready to be enabled.
[0233] In some embodiments, the processor may directly use the audio devices corresponding to the determined number to be enabled as the audio devices to be enabled.
[0234] Step 540 : Determine a set of first matching relationships that meet a preset matching condition based on the target playback area of the content feature.
[0235] For a description of the target playback area, please refer to the description in step 410 of FIG. 4 .
[0236] The first matching relationship refers to the matching relationship between the content feature and the audio device to be enabled.
[0237] In some embodiments, the first matching relationship includes one or more audio devices to be enabled corresponding to each content feature.
[0238] In some embodiments, the processor may randomly determine a set of first matching relationships that meet preset matching conditions based on the target playback area of the content feature.
[0239] Preset matching conditions refer to the preset matching conditions between the audio device to be enabled and the content feature. In some embodiments, the preset matching conditions may include: the audio device to be enabled corresponding to the content feature supports all streaming media types corresponding to the content feature (to avoid playback issues caused by the device not supporting the streaming media type of a content feature); and the number of audio devices to be enabled corresponding to the content feature is not less than the minimum number of enabled audio devices corresponding to each content feature determined based on the user's grouping requirements.
[0240] Step 550: Generate streaming media to be matched that meets the preset quantity requirement.
[0241] The "matching stream" refers to the stream to be determined that corresponds to the content characteristics. Each matching stream corresponds to a streaming type. The specific streaming type corresponding to the matching stream can be determined based on the target audio source. See below for details.
[0242] In some embodiments, the preset quantity requirement includes that the number of streaming media to be matched is not less than the number of content features in the content feature set, that is, each content feature must correspond to at least one streaming media to be matched.
[0243] In some embodiments, the processor may randomly generate streaming media to be matched, the number of which is greater than or equal to the number of content features in the content feature set.
[0244] Step 560: Determine the correspondence between the streaming media to be matched and the content features to obtain a second matching relationship.
[0245] The second matching relationship refers to the corresponding relationship between the streaming media to be matched and the content features.
[0246] In some embodiments, the processor may randomly correspond each to-be-matched streaming media to a content feature in a content feature set to obtain a second matching relationship.
[0247] Step 570: Generate one or more streaming media groups.
[0248] A streaming media group is a collection of one or more streaming media to be matched.
[0249] In some embodiments, the processor may randomly generate one or more streaming media groups. The number of generated streaming media groups may be less than or equal to the number of streaming media to be matched.
[0250] Step 580: Determine the correspondence between the streaming media to be matched and the streaming media group to obtain a third matching relationship.
[0251] The third matching relationship refers to the corresponding relationship between the streaming media to be matched and the streaming media group.
[0252] In some embodiments, the processor may randomly generate a certain number (e.g., less than or equal to the number of streaming media to be matched) of streaming media groups, and randomly assign each streaming media to be matched to one of the streaming media groups, so that each streaming media group includes at least one streaming media to be matched, thereby obtaining a third matching relationship.
[0253] Step 590: Obtain a candidate grouping solution based on the first matching relationship, the second matching relationship, and the third matching relationship.
[0254] For the definition of candidate grouping schemes, please refer to the description in step 410 of FIG. 4 .
[0255] In some embodiments, since a streaming media group corresponds to a candidate grouping scheme, the processor can determine all the streaming media to be matched contained in a streaming media group based on the third matching relationship, and then the processor can determine the matching relationship between each streaming media to be matched and the content feature based on the second matching relationship; then the processor determines the matching relationship between the content feature and the audio device to be enabled based on the first matching relationship, so as to obtain all the audio devices to be enabled contained in each streaming media group, and then obtain a corresponding candidate grouping scheme; for example, all the audio devices to be enabled contained in a streaming media group correspond to all the audio devices to be enabled contained in a device group in a candidate grouping scheme.
[0256] In some embodiments, the processor may repeatedly execute steps 510 to 590 to generate multiple candidate grouping solutions.
[0257] In some embodiments, the processor may further determine the target grouping scheme and device grouping information based on the generated multiple candidate grouping schemes using the method of steps 420 to 440 in FIG. 4 .
[0258] In some embodiments, the processor may further determine, based on the target audio source, a media type of the streaming media to be matched in the target grouping scheme; and generate a play instruction based on the target audio source and the target grouping scheme.
[0259] In some embodiments, the processor can determine, based on the target audio source, a reference streaming media type corresponding to the target audio source using a second preset comparison table; and determine the reference streaming media type corresponding to the target audio source as the media type of the streaming media to be matched corresponding to the corresponding audio device to be enabled in the target grouping scheme. The second preset comparison table includes a correspondence between the reference target audio source and the reference streaming media type corresponding to the reference target audio source. The second preset comparison table can be constructed based on prior knowledge or historical data.
[0260] In some embodiments of the present specification, one or more candidate grouping schemes are automatically generated based on the user's grouping requirements, and a target grouping scheme is automatically determined among the candidate schemes to obtain device grouping information. Then, based on the user's playback requirements, the target device is automatically controlled to play the target audio source, thereby automatically and quickly meeting the user's playback requirements and improving the user's usage experience.
[0261] FIG6 is an exemplary flow chart of determining the estimated playback quality of candidate grouping schemes according to some embodiments of this specification. In some embodiments, process 600 may be executed by processor 140 or dynamic multi-streaming implementation system 200. As shown in FIG6, process 600 includes the following steps 610-620.
[0262] Step 610 : Determine the signal transmission quality 613 of the audio device based on the signal transmission delay 611 and packet loss rate 612 of the audio device during historical playback.
[0263] For a description of signal transmission delay, packet loss rate, and signal transmission quality, please refer to the description of step 420 in FIG. 4 .
[0264] In some embodiments, the processor may determine the signal transmission quality of the audio device based on the signal transmission delay and packet loss rate of the audio device during historical playback using a third preset comparison table. The third preset comparison table includes a correspondence between the signal transmission delay and packet loss rate of a reference audio device during historical playback and the signal transmission quality of the reference audio device. The third preset comparison table may be constructed based on prior knowledge or historical data.
[0265] In some embodiments, the processor may also determine multiple sub-evaluation values of the audio device based on the signal transmission delay and packet loss rate during multiple historical playbacks of the audio device; and determine the signal transmission quality of the audio device based on the weighted sum of the multiple sub-evaluation values.
[0266] The sub-evaluation value refers to the evaluation value of the signal transmission delay and packet loss rate during a historical playback of the audio device.
[0267] In some embodiments, the processor may determine a sub-evaluation value corresponding to a historical playback based on the signal transmission delay and packet loss rate during a historical playback, in accordance with the relationship that the sub-evaluation value is negatively correlated with the signal transmission delay and the packet loss rate.
[0268] As an example only, the processor may use the following first calculation formula to determine the Nth sub-evaluation value p: p=exp -(a+b) a is the average signal transmission delay during the Nth historical playback of the audio device, and b is the average packet loss rate during the Nth historical playback of the audio device.
[0269] The average signal transmission delay and average packet loss rate refer to the average transmission delay and packet loss rate at multiple preset time points within a preset collection period. The preset collection period and multiple preset time points are historical time. The preset collection period can be the entire or partial duration of the corresponding historical playback and can be preset by those skilled in the art based on experience.
[0270] In some embodiments, the processor may determine the signal transmission quality of the audio device based on a weighted sum of the multiple sub-evaluation values.
[0271] In some embodiments, when multiple sub-evaluation values are weighted and summed, the weight of the sub-evaluation value is negatively correlated to the time interval between the occurrence time of the historical playback corresponding to the sub-evaluation value and the current time. For example, the weight N corresponding to the Nth sub-evaluation value corresponding to the Nth historical playback is exp (-时间间隔n) The time interval n is the interval between the historical playback time (eg, the end time of the audio device playback during the Nth historical playback) n and the current time point, and the unit can be hours h.
[0272] In some embodiments, the weight N may be normalized, for example, the normalized weight N1 = weight N / (weight 1 + weight 2 + ... + weight N).
[0273] As an example only, the processor may use the following second calculation formula to determine the signal transmission quality:
[0274] Signal transmission quality=1st sub-evaluation value*weight 1+2nd sub-evaluation value*weight 2+...+Nth sub-evaluation value*weight N.
[0275] In some embodiments of the present specification, the signal transmission quality of the audio device is determined by weighted summing of multiple sub-evaluation values, thereby avoiding the influence of accidental factors and improving the accuracy of the ultimately determined signal transmission quality of the audio device.
[0276] Step 620 : Based on the signal transmission quality 613 of each audio device in the candidate grouping scheme, an estimated playback quality 621 of the candidate grouping scheme is determined.
[0277] For the description of estimating the playback quality, please refer to the description in step 420 of FIG. 4 .
[0278] In some embodiments, the processor may determine the estimated playback quality of the candidate grouping scheme based on various methods. For example, the processor may determine the sum of the signal transmission qualities of the audio devices in the candidate grouping scheme as the estimated playback quality of the candidate grouping scheme.
[0279] In some embodiments, the processor may also use other methods to determine the estimated playback quality of the candidate grouping schemes. For details, see the description in FIG. 7 .
[0280] In some embodiments, in response to a target device failure, the processor may determine a new playback solution based on the signal transmission quality of the remaining audio devices and the number of channels of the target audio source, wherein the new playback solution includes a new target device and playback mode.
[0281] The number of channels of the target audio source refers to the number of channels played by the target audio source. For example, if the target audio source is a surround sound source, the number of channels of the target audio source is 6. For another example, if the target audio source is a stereo source, the number of channels of the target audio source is 2.
[0282] The new playback plan refers to another playback plan different from the current playback plan.
[0283] In some embodiments, the new playback scheme includes a new target device and a new playback mode. The new playback mode refers to a playback mode that is different from the current playback mode. For example, the target audio source can be switched from playing a certain number of channels (e.g., 6 channels) to playing a different number of channels (e.g., 2 channels).
[0284] In some embodiments, in response to a failure of the target device, the processor may determine a new playback solution based on the signal transmission quality of the remaining audio devices and the number of channels of the target audio source.
[0285] For example, when there are 6 audio devices (audio device 1, audio device 2...audio device 6) in the bedroom, the target devices are audio device 1 and audio device 2, and the target sound source is a 6-channel surround sound source, the processor can respond to a failure of audio device 1 and / or audio device 2. The processor can then select 2 audio devices whose preset transmission quality meets the preset transmission quality requirements from the remaining non-faulty audio devices 3, audio device 4, audio device 5 and audio device 6 in the bedroom, and replay the target sound source in 2-channel mode.
[0286] The preset quality requirement may be that the difference in signal transmission quality between any two new audio devices is no greater than a signal transmission quality threshold. The signal transmission quality threshold may be preset by those skilled in the art based on experience.
[0287] In some embodiments of this specification, in response to a target device failure, a new playback scheme is determined based on the signal transmission quality of the remaining audio devices and the number of channels of the target audio source. This can avoid the asymmetric sound playback caused by the audio device failure, resulting in mixed and unharmonious sound, and ensure audio playback quality. Furthermore, the signal transmission quality differences between the newly determined audio devices are minimal, ensuring consistent playback quality.
[0288] FIG. 7 is a schematic diagram of another method for determining estimated playback quality of a candidate grouping solution according to some embodiments of this specification.
[0289] In some embodiments, the processor may determine an estimated playback quality of the candidate grouping scheme based on a quality prediction model 720. The quality prediction model 720 may be used to process a candidate grouping scheme 710-1, signal transmission quality 710-2 of each audio device in the candidate grouping scheme, and first evaluation information 710-3 of each audio device to determine an estimated playback quality 730 of the candidate grouping scheme.
[0290] In some embodiments, the first evaluation information 710 - 3 of each audio device in the candidate grouping scheme includes the signal transmission delay and packet loss rate of the audio device at multiple time points within a preset time period.
[0291] The preset time period can be the period between the current time point and a preset historical time point. For example, if the current time point is 12:00 PM and the preset historical time point is 1:00 AM on the same day, the preset time period is the period between 1:00 AM and 12:00 PM. Multiple time points can be multiple time points within the preset time period; if the time period is between 1:00 AM and 12:00 PM, multiple time points are selected every half hour, starting at 1:00 AM.
[0292] For an explanation of candidate grouping schemes, estimated playback quality, signal transmission delay, and packet loss rate, please refer to the explanation in FIG4 . For an explanation of signal transmission quality, please refer to the explanation in step 610 in FIG6 .
[0293] In some embodiments, the quality prediction model may be a machine learning model. In some embodiments, the quality prediction model may include a neural network model (NN), a deep neural network (DNN), etc.
[0294] In some embodiments, the quality prediction model 720 may be obtained through training based on a plurality of labeled training samples.
[0295] In some embodiments, each group of training samples may include a historical sample grouping scheme, signal transmission quality of each audio device in the historical sample grouping scheme, and first evaluation information. The label may be an estimated playback quality of the candidate historical sample grouping scheme corresponding to the training sample.
[0296] In some embodiments, the historical sample grouping scheme and the first evaluation information of each audio device in the historical sample grouping scheme can be obtained through historical data or simulation. In some embodiments, the signal transmission quality of each audio device in the historical sample grouping scheme can be calculated based on the first evaluation information of each audio device in the historical sample grouping scheme, using the signal transmission quality calculation method described in step 610 of Figure 6.
[0297] In some embodiments, the label is the actual playback quality of the historical sample grouping scheme. The processor can determine the actual playback quality of the historical sample grouping scheme based on the freeze frequency, freeze duration, distortion frequency, and frequency of hum / noise of each audio device in the historical playback situation. For example, when the frequency of freezes, hum / noise, etc. occurring during the historical playback of an audio device exceeds a preset frequency threshold, the label is 0; when the frequency of freezes, hum / noise, etc. occurring during the historical playback of an audio device is lower than a preset frequency threshold, the label is 1. The preset frequency threshold can be preset by a person skilled in the art based on experience.
[0298] In some embodiments of this specification, the quality prediction model can be used to quickly and accurately predict the estimated playback quality of candidate grouping schemes, obtain the accurate target device to play the target sound source, and automatically and quickly meet the user's playback needs, further improving the user's experience.
[0299] While the basic concepts have been described above, it will be apparent to those skilled in the art that the detailed disclosure is merely illustrative and does not limit this specification. Although not explicitly stated herein, various modifications, improvements, and revisions to this specification may be made by those skilled in the art. Such modifications, improvements, and revisions are suggested in this specification and remain within the spirit and scope of the exemplary embodiments of this specification.
[0300] This specification also uses specific terms to describe the embodiments of this specification. For example, "one embodiment," "an embodiment," and / or "some embodiments" refer to a feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "one embodiment," "an embodiment," or "an alternative embodiment" two or more times in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, certain features, structures, or characteristics of one or more embodiments of this specification may be appropriately combined.
[0301] In addition, unless expressly stated in the claims, the order of the processing elements and sequences, the use of alphanumeric characters, or the use of other names described in this specification are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some of the invention embodiments currently considered useful through various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the spirit and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0302] Similarly, it should be noted that, in order to simplify the presentation of this specification and thus facilitate understanding of one or more embodiments of the invention, the foregoing descriptions of the embodiments of this specification sometimes combine multiple features into a single embodiment, figure, or description thereof. However, this disclosure method does not imply that the subject matter of this specification requires more features than those recited in the claims. In fact, an embodiment may have fewer features than all of the features of a single disclosed embodiment.
[0303] In some embodiments, numbers are used to describe the quantity of components and attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise stated, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the description and claims are approximate values, which may change according to the required characteristics of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the settings of such numerical values are as accurate as possible within the feasible range.
[0304] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, and documents, cited in this specification is hereby incorporated by reference in its entirety. This includes application history documents that are inconsistent with or conflict with the content of this specification, as well as documents (currently or subsequently attached to this specification) that limit the broadest scope of the claims of this specification. It should be noted that if the descriptions, definitions, and / or terminology used in the accompanying materials are inconsistent or conflicting with the content of this specification, the descriptions, definitions, and / or terminology used in this specification will control.
[0305] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.
Claims
1. A method for implementing dynamic multi-streaming, characterized in that: include: generating a grouping instruction based on a first operation of the user, wherein the first operation of the user is related to a grouping requirement of the user; Grouping the audio devices in the target area based on the grouping instruction to obtain device grouping information; Generate a play instruction based on the device grouping information and a second user operation, wherein the second user operation is related to a user's play requirement, and the play requirement includes a target sound source; as well as Based on the play instruction, the target device is controlled to play the target sound source.
2. The method according to claim 1, characterized in that The grouping of the audio devices in the target area based on the grouping instruction to obtain device grouping information includes: Grouping the audio devices in the target area based on the grouping instruction to obtain first grouping information; Determine first streaming media group information based on the first grouping information; Acquire a first control instruction of a user, where the first control instruction includes update information of the first streaming media group information; Based on the first control instruction, update the first streaming media group information to obtain second streaming media group information; and The device grouping information is determined based on the second streaming media group information.
3. The method according to claim 2, characterized in that The obtaining of the first control instruction of the user comprises: Generate full list information of the audio devices based on the first streaming media group information; Displaying the full list of information to the user; and The first control instruction is determined based on the user's adjustment information on the full list information.
4. The method according to claim 1, characterized in that: The user's grouping requirement includes a content feature set and matching information corresponding to the content feature set, wherein the content feature set includes one or more content features; The grouping of the audio devices in the target area based on the grouping instruction to obtain device grouping information includes: Based on the grouping requirements of the user, one or more candidate grouping schemes are generated; the candidate grouping schemes include the content features and the audio devices that play the content features; Determining an estimated playback quality of the candidate grouping scheme based on historical playback information of the audio devices in the candidate grouping scheme; Determining a target grouping scheme based on the estimated playback quality; and Based on the target grouping scheme, the device grouping information is determined.
5. The method according to claim 4, characterized in that The matching information corresponding to the content feature set includes a target playback area of the content feature and device requirement information corresponding to the content feature.
6. The method according to claim 5, characterized in that The generating one or more candidate grouping solutions based on the grouping requirement of the user comprises: Determining a device quantity range of the target devices corresponding to the content feature set based on the quantity of the content features in the content feature set and the device requirement information corresponding to the content features; Based on the device quantity range, determining the number of the target devices to be enabled; Determine the audio devices corresponding to the number to be enabled as the audio devices to be enabled; Based on the target playback area of the content feature, determining a set of first matching relationships that meet a preset matching condition; the first matching relationship includes each of the content features corresponding to one or more of the audio devices to be enabled; Generate streaming media to be matched that meets a preset quantity requirement; the preset quantity requirement includes that the number of streaming media to be matched is not less than the number of content features in the content feature set; Determine the corresponding relationship between the streaming media to be matched and the content feature to obtain a second matching relationship; Generate one or more streaming media groups; Determine the correspondence between the to-be-matched streaming media and the streaming media group to obtain a third matching relationship; and Based on the first matching relationship, the second matching relationship, and the third matching relationship, a candidate grouping scheme is obtained.
7. The method according to claim 4, characterized in that The historical playback information of the audio device includes signal transmission delay and packet loss rate during historical playback.
8. The method according to claim 7, characterized in that The determining the estimated playback quality of the candidate grouping scheme based on the historical playback information of the audio device in the candidate grouping scheme includes: The signal transmission delay and the packet loss rate of the audio device during the historical playback are used to determine the signal transmission delay of the audio device. Output quality; and Based on the signal transmission quality of each of the audio devices in the candidate grouping scheme, an estimated playback quality of the candidate grouping scheme is determined.
9. The method according to claim 8, characterized in that The determining the signal transmission quality of the audio device based on the signal transmission delay and the packet loss rate of the audio device during the historical playback includes: Determining a plurality of sub-evaluation values of the audio device based on the signal transmission delay and the packet loss rate during the multiple historical playbacks of the audio device; and The signal transmission quality of the audio device is determined based on a weighted sum of the plurality of sub-evaluation values.
10. The method according to claim 9, characterized in that When the multiple sub-evaluation values are weighted and summed, the weight of the sub-evaluation value is negatively correlated to the time interval between the occurrence time of the historical playback corresponding to the sub-evaluation value and the current time.
11. The method according to claim 8, characterized in that The determining the estimated playback quality of the candidate grouping scheme based on the signal transmission quality of each of the audio devices in the candidate grouping scheme comprises: Based on the candidate grouping scheme, the signal transmission quality of each audio device in the candidate grouping scheme and the first evaluation information of each audio device, the estimated playback quality of the candidate grouping scheme is determined through a quality prediction model; the quality prediction model is a machine learning model; the first evaluation information includes the signal transmission delay and the packet loss rate of the audio device at multiple time points within a preset time period.
12. The method according to claim 1, wherein the playback requirement further comprises: The channel audio signals played by each of the target devices correspond to each other.
13. The method according to claim 1, wherein the target sound source comprises a mixed sound source obtained by mixing at least two sound sources of streaming media types.
14. The method according to claim 1, characterized in that The method further comprises: Determine a sound source update instruction according to a third operation of the user; the third operation of the user is related to the user's playback update requirement, and the playback update requirement includes a target device group and an updated sound source of the target device group; and Based on the sound source update instruction, the target device group is controlled to play the updated sound source.
15. The method according to claim 1, characterized in that The determination of the target device includes: Determining a user operation instruction based on the first user operation and / or the second user operation; the user operation instruction includes a target sound source and a target playback area; and Based on the user operation instruction, audio source pairing is performed and the target device is determined.
16. The method according to claim 15, characterized in that The performing audio source pairing and determining the target device based on the user operation instruction includes: Pairing the target audio source with the target sending channel; Pairing the target transmission channel with the target receiving channel of the target playback area; and Based on the target receiving channel, the target device is determined.
17. A dynamic multi-stream implementation system, characterized in that: include: A grouping instruction generating module, configured to generate a grouping instruction based on a first operation of a user, wherein the first operation of the user is related to a grouping requirement of the user; a grouping module, configured to group the audio devices in the target area based on the grouping instruction to obtain device grouping information; A play instruction generating module is configured to generate a play instruction based on the device grouping information and a second user operation, wherein the second user operation is related to a user's play requirement, and the play requirement includes a target sound source; The playing module is configured to control the target device to play the target sound source based on the playing instruction.
18. The system according to claim 17, characterized in that The grouping module is further configured to: Grouping the audio devices in the target area based on the grouping instruction to obtain first grouping information; Determine first streaming media group information based on the first grouping information; Acquire a first control instruction of a user, where the first control instruction includes update information of the first streaming media group information; Based on the first control instruction, update the first streaming media group information to obtain the second streaming media group information; as well as The device grouping information is determined based on the second streaming media group information.
19. The system according to claim 17, characterized in that The system also includes an instruction update module; The instruction update module is configured to determine a sound source update instruction according to a third operation of the user; the third operation of the user is related to the user's playback update requirement, and the playback update requirement includes a target device group and an updated sound source of the target device group; The playing module is further configured to control the target device group to play the updated sound source based on the sound source update instruction.
20. A computer-readable storage medium storing computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the method for implementing dynamic multi-streaming as claimed in claim 1.
Citation Information
Patent Citations
Broadcast control method and terminal
CN104810032A
Playing control method and device, and terminal
CN106452644A
Multicast sending method and device of streaming media, multicast server and medium
CN113099259A
Media playing scheme determination method and system
CN116828235A
Playing equipment control method and device, equipment and medium
CN117014674A