Audio track switching method and electronic equipment

By collecting multimodal data and utilizing convolutional neural networks and context prediction models, the target audio track is automatically matched, solving the problems of low audio track switching efficiency and auditory disconnection in existing technologies, and achieving seamless and lossless audio track switching.

CN120932669APending Publication Date: 2025-11-11VIDAA (NETHERLANDS) INT HLDG LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511071564.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing audio track switching methods require manual user intervention or cannot accurately match user preferences, resulting in low switching efficiency and auditory gaps.

Method used

By collecting multimodal data and utilizing convolutional neural networks and context prediction models, the system automatically matches the target audio track and selects the switching timing based on audio feature analysis, achieving seamless and lossless audio track switching.

Benefits of technology

It improves the efficiency and accuracy of audio track switching, eliminates auditory gaps, reduces latency, and meets users' actual preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932669A_ABST
    Figure CN120932669A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio track switching method and electronic equipment, and the method comprises the steps: collecting multi-modal data through a detection device when a first audio stream corresponding to a first audio track is played; performing normalization processing on the multi-modal data by using a convolutional neural network model to obtain a multi-modal feature vector, inputting the multi-modal feature vector into a context prediction model, and predicting a target audio track matched with the multi-modal data by the context prediction model; detecting a mute section, a zero-crossing rate point, an energy stable section and a harmonic stable section of the first audio stream to obtain a candidate point set; selecting a target candidate point from the candidate point set; and performing phase alignment on the first audio track and the target audio track, and switching to the target audio track at the target candidate point to play a second audio stream corresponding to the target audio track. Therefore, the target audio track conforming to the actual preference of the user is predicted, the switching opportunity of the target audio track is automatically matched, auditory faults are eliminated, and the audio track switching efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of display device technology, and in particular to an audio track switching method and electronic device. Background Technology

[0002] An audio track is used to store audio signals. When creating an audio track, you can define the attributes of the audio signal, such as timbre, volume, and number of channels. Electronic devices can switch the audio signal being played by switching audio tracks.

[0003] In some implementations, users can manually switch audio tracks to listen to other audio. For example, when an electronic device is playing rock music, if the user wants to rest, they can manually switch to a target audio track, causing the electronic device to play the corresponding sleep mode audio, such as white noise or soothing light music. This method requires manual intervention from the user when switching tracks, resulting in low efficiency.

[0004] In some implementations, electronic devices can execute audio track switching schemes based on audio feature matching. By extracting audio features and matching them with features contained in a preset track feature library, the device can locate and switch to the target track. Electronic devices can also execute audio track switching schemes based on audio timing. This involves dynamically capturing the timing feature vector of the audio signal through a sliding window, comparing this timing feature vector with a preset track template, and then matching the target track with the highest similarity. While these two audio track switching schemes do not require manual track switching by the user, they have drawbacks: the matched target track may not accurately match the user's actual preferences, they cannot adapt to complex application environments, and there may be perceptible auditory gaps when switching audio across tracks, resulting in significant delays in track switching. Summary of the Invention

[0005] Some embodiments of this application provide an audio track switching method and electronic device. By collecting multimodal data including factors such as user behavior and environmental status, and performing context prediction on the multimodal data, the method predicts a target audio track that matches the user's actual preferences, automatically matches the switching timing of the target audio track, and automatically switches to the target audio track based on the switching timing. This eliminates the user's auditory blockage, reduces audio track switching delay, and eliminates the need for manual intervention in audio track switching. Instead, it automatically matches and switches to the target audio track based on multimodal data, thereby improving audio track switching efficiency.

[0006] Firstly, some embodiments of this application provide an audio track switching method, including:

[0007] When playing the first audio stream corresponding to the first audio track, multimodal data is collected by a detection device. The multimodal data includes visual modal data, auditory modal data and motion modal data.

[0008] The multimodal data is normalized using a convolutional neural network model to obtain multimodal feature vectors;

[0009] The multimodal feature vector is input into the context prediction model, which then predicts the target audio track that matches the multimodal data.

[0010] Detect the silence segment, zero-crossing point, energy stability segment, and harmonic stability segment of the first audio stream;

[0011] Obtain a candidate point set, the candidate point set including at least one of a first candidate point, a second candidate point, a third candidate point, and a fourth candidate point; wherein, the first candidate point is selected from the silent segment, the second candidate point is selected from the zero-crossing rate point, the third candidate point is selected from the energy stable segment, and the fourth candidate point is selected from the harmonic stable segment;

[0012] Select a target candidate point from the candidate point set, wherein the target candidate point is the switching point from the first audio track to the target audio track;

[0013] The first audio track and the target audio track are phase aligned, and the target audio track is switched at the target candidate point to play the second audio stream corresponding to the target audio track.

[0014] The beneficial effects of the embodiments in the first aspect above are as follows: By capturing user behavior, environmental conditions, and user voice through a detection device, multimodal data is divided into visual modal data (e.g., user body movements and visual direction captured by visual sensors / cameras), auditory modal data (e.g., rhythm, emotion, and semantic coherence involved in user dialogue), and motion modal data (e.g., changes in environmental scenes (including changes in environmental conditions such as noise, temperature, humidity, and lighting), changes in the number of people in the space, changes in user movement trajectory, and changes in user heart rate during exercise). This multimodal data collection provides multiple dimensions and factors for subsequent prediction of the target audio track, making the automatically matched target audio track more consistent with the user's actual habits and preferences. Furthermore, the context prediction model can perform multidimensional analysis of the multimodal feature vectors, thereby automatically predicting and matching the target audio track without requiring manual user intervention in track switching, which can improve... Audio track switching efficiency: Under the current first audio track, audio attributes such as silent segments, zero-crossing points, energy-stable segments, and harmonic-stable segments of the first video stream are detected. This provides a basis for obtaining a candidate point set, enabling audio track switching at silent frames, continuous consonant frames, energy-stable frames, or harmonic-stable frames. This avoids switching audio tracks in vocal frequency bands, strong rhythmic intervals, vowel formant intervals, and emotionally intense areas. A target candidate point is selected from the candidate point set, and the timing of switching to the target audio track is determined based on this target candidate point. This improves the audio quality of audio transitions and connections during track switching. Phase alignment is performed between the first and target audio tracks to ensure better audio transitions during track switching, avoiding audio attenuation or distortion caused by phase cancellation. This achieves seamless, lossless, and smooth switching between audio tracks, eliminating auditory gaps during audio switching, reducing track switching latency, and improving the speed and accuracy of track switching.

[0015] In some embodiments of the first aspect, the step of obtaining the candidate point set includes: calculating the zero-crossing rate, energy fluctuation range, and fundamental frequency change rate of audio frames in the first audio stream; when the length of the silent segment is greater than a first threshold, taking the midpoint of the silent segment as the first candidate point; and / or, when the zero-crossing rate of a consecutive preset number of audio frames is less than a second threshold, taking the valley point among the zero-crossing rate points as the second candidate point; and / or, when the energy fluctuation range is within the threshold range, taking the starting point of the energy stable segment as the third candidate point; and / or, when the fundamental frequency change rate is less than a third threshold, taking the midpoint of the harmonic stable segment as the fourth candidate point. The beneficial effects of this embodiment are as follows: When the length of the silent segment is greater than the first threshold, it indicates that the first audio stream has a relatively long continuous silent area. The midpoint of the silent segment is a relatively stable silent frame, which can be selected as the first candidate point. This allows the silent frame to be used as the switching point, avoiding switching audio tracks in non-continuous silent areas such as the vocal frequency band. When the zero-crossing rate of a preset number of consecutive audio frames is less than the second threshold, it indicates that the preset number of audio frames are consonant frames, and the first audio stream exhibits a discontinuous and unclear playback state. Among the zero-crossing rate points of the first audio stream, the valley point belongs to... For natural pauses, the zero-crossing rate trough can be selected as the second candidate point, thus avoiding track switching in the vowel formant range. If the energy fluctuation range is within the threshold range, it indicates that the energy fluctuation of the first audio stream is small and relatively stable; therefore, the starting point of the stable energy segment can be selected as the third candidate point, thus avoiding track switching in the strong rhythm range. If the fundamental frequency change rate is less than the third threshold, it indicates that the harmonics of the first audio stream are relatively stable, and the first audio stream presents a relatively soothing or calm emotional state rather than a strong emotional state; therefore, the midpoint of the harmonic stability segment can be selected as the fourth candidate point. In this way, track switching can be achieved at audio silence frames, continuous consonant frames, energy stability frames, or harmonic stability frames, avoiding track switching in the vocal frequency range, strong rhythm range, vowel formant range, and emotionally intense areas. By selecting a target candidate point from the candidate point set and determining the timing of switching to the target track based on this target candidate point, the audio quality of audio transitions and connections during track switching is improved.

[0016] In some embodiments of the first aspect, the step of selecting a target candidate point from the candidate point set includes: sequentially extracting audio features of a preset length before and after a preliminary candidate point; wherein the preliminary candidate point is a candidate point included in the candidate point set; processing the audio features of the preset length before and after the preliminary candidate point using an encoder of a converter model to calculate the quality score of the candidate point; filtering out unqualified candidate points from the candidate point set, wherein the quality score of the unqualified candidate point is less than a fourth threshold; merging adjacent candidate points in the filtered candidate point set whose spacing is less than a fifth threshold to obtain the target candidate point. The beneficial effects of this embodiment are as follows: For each candidate point in the candidate point set, audio features of a preset length before and after the candidate point are extracted. The audio features of the preset length before and after the candidate point are processed using a Transformer encoder to predict the quality score of the candidate point. The quality score of the candidate point is compared with a fourth threshold. If the quality score of the candidate point is less than the fourth threshold, it indicates that the candidate point is a candidate point with unqualified quality. The unqualified candidate point is then filtered out from the candidate point set, so that the candidate point set only contains candidate points with qualified quality. If there are multiple candidate points remaining in the candidate point set, adjacent candidate points with a spacing less than a fifth threshold can be merged. In this way, through the audio attributes of the first audio stream (such as silence, filter rate valley, harmonic stability, energy stability), candidate point quality screening, and merging of adjacent candidate points, a target candidate point is finally obtained. This target candidate point is the optimal audio track switching point, thereby improving the audio track switching effect.

[0017] In some embodiments of the first aspect, the step of inputting the multimodal feature vector into a context prediction model and having the context prediction model predict a target audio track matching the multimodal data includes: the context prediction model performing joint modeling from spatial and temporal dimensions based on the multimodal feature vector and a context memory pool to obtain spatiotemporal joint features; wherein, the context memory pool is used to store playback history information, the playback history information including audio track information related to different user habits and environmental states; the context prediction model using a fully connected layer and a normalized exponential function to predict user behavior on the spatiotemporal joint features to obtain a probability distribution of user behavior categories; the context prediction model using a long short-term memory network decoder to predict a user trajectory sequence within a future preset time period; the context prediction model determining a first candidate audio track based on the probability distribution of the user behavior categories and the user trajectory sequence; the context prediction model outputting the confidence score of the first candidate audio track based on a regression network; obtaining a second candidate audio track, the confidence score of the second candidate audio track being greater than a confidence score threshold; and determining the second candidate audio track with the highest confidence score as the target audio track. The beneficial effects of this embodiment are as follows: The context prediction model can call the context memory pool, which is prior knowledge used to store audio track information related to factors such as user habits and environmental conditions during historical usage, thereby providing knowledge reserves for subsequent prediction of target audio tracks. In this way, by performing spatiotemporal fusion based on multimodal feature vectors and context memory pool, spatiotemporal joint features are obtained. Then, user behavior prediction can be performed, outputting the probability distribution of user behavior categories and predicting the user trajectory sequence within a preset time period in the future. In this way, based on user behavior and action trajectory, it is possible to predict which audio playback trends may exist based on the user behavior and action trajectory, and thus provide at least one first candidate audio track. The confidence of the first candidate audio track is calculated. If there is a second candidate audio track with a confidence value greater than the confidence threshold, the second candidate audio track with the highest confidence value is selected as the target audio track. This can filter out the audio track category with the highest probability / trend, improve the accuracy of target audio track prediction, and thus improve the accuracy of automatic audio track matching and switching.

[0018] In some embodiments of the first aspect, the step of obtaining spatiotemporal joint features by the context prediction model based on the multimodal feature vector and the context memory pool from the spatial and temporal dimensions includes: the context prediction model retrieving target playback history information associated with the multimodal data from the context memory pool, and generating a context feature vector based on the target playback history information; the context prediction model concatenating the multimodal feature vector and the context feature vector to obtain a concatenated vector; the context prediction model acquiring spatial and temporal features from the concatenated vector; the context prediction model performing spatial graph convolution operation on the spatial features and a graph convolutional neural network to obtain spatial modal features; the context prediction model performing temporal transformation operation on the temporal features and a recurrent neural network to obtain temporal modal features; and the context prediction model fusing the spatial modal features and the temporal modal features to obtain the spatiotemporal joint features. The beneficial effects of this embodiment are as follows: when generating spatiotemporal joint features, correlation retrieval is performed based on the context memory pool, and after concatenating the multimodal feature vector and the context feature vector, it is mapped to the spatial and temporal dimensions, and finally spatiotemporal fusion is performed to obtain spatiotemporal joint features. In this way, the target playback history information associated in the context memory pool is used as prior knowledge for subsequent prediction of target audio tracks, thereby improving the accuracy of target audio track prediction and thus improving the accuracy of automatic audio track matching and switching.

[0019] In some embodiments of the first aspect, after inputting the multimodal feature vector into a context prediction model and having the context prediction model predict a target audio track matching the multimodal data, the method further includes: preloading the second audio stream corresponding to the target audio track from an audio buffer pool, wherein the audio buffer pool is used to cache audio streams corresponding to different audio tracks. The beneficial effect of this embodiment is that after determining the target audio track, the second audio stream corresponding to the target audio track can be preloaded, improving the start-up speed of the second video stream corresponding to the target audio track, thereby improving audio track switching efficiency.

[0020] In some embodiments of the first aspect, the step of switching to the target audio track at the target candidate point and playing the second audio stream corresponding to the target audio track includes: generating a first audio track switching instruction to be executed based on the target audio track and the target candidate point; sending the first audio track switching instruction to a switching executor; and having the switching executor respond to the first audio track switching instruction by transitioning from the first audio track to the target audio track at the target candidate point using a preset transition method to play the second audio stream corresponding to the target audio track. The beneficial effects of this embodiment are: to achieve precise control of audio track switching, a first audio track switching instruction to be executed can be generated based on the category information of the target audio track and the target candidate point, and the instruction can be sent to the underlying switching executor; the switching executor determines the timing of audio track switching corresponding to the target candidate point and can use a preset transition method to transition from the first audio track to the target audio track at the target candidate point, wherein the preset transition method is, for example, fade-in / fade-out, crossfade, etc., to achieve seamless and smooth audio track switching and improve the transition effect during audio track switching.

[0021] In some embodiments of the first aspect, before sending the first audio track switching instruction to the switching executor, the method further includes: if a second audio track switching instruction currently being executed by the switching executor is detected, obtaining a first instruction priority corresponding to the first audio track switching instruction and a second instruction priority corresponding to the second audio track switching instruction set by the user; writing the first instruction priority into the first audio track switching instruction; and marking the second audio track switching instruction as corresponding to the second instruction priority. The beneficial effects of this embodiment are: since user behavior and environmental state are variables, multimodal data may be updated following changes in user behavior / trajectory and environmental state. When multimodal data is updated, a new audio track switching instruction will be sent to the switching executor. Thus, the switching executor may receive multiple audio track switching instructions in sequence. If the switching executor is currently executing audio track switching instruction A and then receives a new audio track switching instruction B, the user can set the priorities of audio track switching instructions A and B, thereby providing a reference for the switching executor to execute the audio track switching instructions, improving the accuracy of multi-instruction execution in dynamic multimodal scenarios, and making audio track switching more in line with user habits and preferences.

[0022] In some embodiments of the first aspect, the step of the switching executor responding to the first audio track switching instruction, transitioning from the first audio track to the target audio track at the target candidate point through a preset transition method, and playing the second audio stream corresponding to the target audio track, includes: the switching executor obtaining the first instruction priority from the first audio track switching instructions to be executed; the switching executor obtaining the second instruction priority corresponding to the currently executing second audio track switching instruction; if the first instruction priority is higher than the second instruction priority, the switching executor stops executing the second audio track switching instruction and executes the first audio track switching instruction, writing first playback history information to the context memory pool; wherein, the first playback history information is used to indicate that the user tends to switch to the target audio track under the user behavior and environmental state corresponding to the multimodal data; if the first instruction priority is lower than the second instruction priority, the switching executor continues to execute the second audio track switching instruction, does not execute the first audio track switching instruction, and writes second playback history information to the context memory pool; wherein, the second playback history information is used to indicate that the user tends to switch to the audio track corresponding to the second audio track switching instruction under the user behavior and environmental state corresponding to the multimodal data. The beneficial effects of this embodiment are as follows: If the switching actuator is currently executing audio track switching instruction A, and a new audio track switching instruction B is received, the priorities of audio track switching instruction A and audio track switching instruction B are compared. If the priority of audio track switching instruction A is higher than the priority of audio track switching instruction B, that is, the currently executing instruction takes precedence over the instruction to be executed, the switching actuator continues to execute audio track switching instruction A and does not execute audio track switching instruction B, thereby prioritizing the playback of audio stream A corresponding to the audio track indicated by audio track switching instruction A, that is, without interrupting the playback of audio stream A; if the priority of audio track switching instruction A is lower than the priority of audio track switching instruction B, that is, the instruction to be executed takes precedence over the currently executing instruction, then the execution of audio track switching instruction A is stopped, and audio track switching instruction B is executed, that is, the playback of audio stream A is interrupted, and the audio stream B corresponding to the audio track indicated by audio track switching instruction B is switched. When the switching actuator controls the response and execution of instructions based on instruction priority, it synchronously writes the execution status into the context memory pool in the form of playback history information, thereby updating the context memory pool. This allows the context memory pool to continuously accumulate and correct prior knowledge, thereby improving the accuracy of target audio track matching and making the target audio track more in line with the user's habits and preferences.

[0023] Secondly, some embodiments of this application also provide an electronic device, including:

[0024] The detection device is configured to acquire multimodal data, including visual modal data, auditory modal data, and motion modal data;

[0025] The audio output device is configured to output the audio stream corresponding to the audio track.

[0026] The controller, coupled to the detection device and the audio output device, is configured as follows:

[0027] When the audio output device outputs the first audio stream corresponding to the first audio track, the multimodal data collected by the detection device is acquired;

[0028] The multimodal data is normalized using a convolutional neural network model to obtain multimodal feature vectors;

[0029] The multimodal feature vector is input into the context prediction model, which then predicts the target audio track that matches the multimodal data.

[0030] Detect the silence segment, zero-crossing point, energy stability segment, and harmonic stability segment of the first audio stream;

[0031] Obtain a candidate point set, the candidate point set including at least one of a first candidate point, a second candidate point, a third candidate point, and a fourth candidate point; wherein, the first candidate point is selected from the silent segment, the second candidate point is selected from the zero-crossing rate point, the third candidate point is selected from the energy stable segment, and the fourth candidate point is selected from the harmonic stable segment;

[0032] Select a target candidate point from the candidate point set, wherein the target candidate point is the switching point from the first audio track to the target audio track;

[0033] The first audio track and the target audio track are phase aligned, and the target audio track is switched at the target candidate point so that the audio output device outputs the second audio stream corresponding to the target audio track.

[0034] The beneficial effects of the embodiments in the second aspect above are as follows: By capturing user behavior, environmental conditions, and user voice through a detection device, multimodal data is divided into visual modal data (e.g., user body movements and visual direction collected by visual sensors / cameras), auditory modal data (e.g., rhythm, emotion, and semantic coherence involved in user dialogue), and motion modal data (e.g., changes in environmental scenes (including changes in environmental conditions such as noise, temperature, humidity, and lighting), changes in the number of people in the space, changes in user movement trajectory, and changes in user heart rate during exercise). This multimodal data collection provides multiple dimensions and factors for subsequent prediction of the target audio track, making the automatically matched target audio track more consistent with the user's actual habits and preferences. Furthermore, the context prediction model can perform multidimensional analysis of the multimodal feature vectors, thereby automatically predicting and matching the target audio track without requiring manual user intervention in track switching, which can improve audio quality. Track switching efficiency; Under the current first audio track, audio attributes such as silent segments, zero-crossing points, energy-stable segments, and harmonic-stable segments of the first video stream are detected, providing a basis for obtaining a candidate point set. This enables audio track switching at audio silence frames, continuous consonant frames, energy-stable frames, or harmonic-stable frames, avoiding switching in vocal frequency bands, strong rhythm ranges, vowel formant ranges, and emotionally intense areas. Target candidate points are selected from the candidate point set, and the timing of switching to the target audio track is determined based on these target candidate points, thereby improving the audio quality of audio transitions and connections during track switching. Phase alignment is performed between the first and target audio tracks to improve audio continuity during track switching, avoiding audio attenuation or distortion caused by phase cancellation. This achieves seamless, lossless, and smooth switching between audio tracks, eliminating auditory gaps during audio switching, reducing track switching latency, and improving the speed and accuracy of track switching. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in some embodiments of this application or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This application provides operational scenario diagrams for audio service processing in some embodiments.

[0037] Figure 2 A schematic diagram of the hardware configuration of an electronic device provided in some embodiments of this application;

[0038] Figure 3 A schematic diagram illustrating the software configuration of an electronic device provided in some embodiments of this application;

[0039] Figure 4 A schematic diagram of a media asset playback interface provided in some embodiments of this application;

[0040] Figure 5 A flowchart of an audio track switching method provided in some embodiments of this application;

[0041] Figure 6 A schematic diagram illustrating the principle of feature fusion for multimodal data provided in some embodiments of this application;

[0042] Figure 7 Prediction process of the context prediction model provided for some embodiments of this application Figure 1 ;

[0043] Figure 8 Prediction process of the context prediction model provided for some embodiments of this application Figure 2 ;

[0044] Figure 9 A flowchart for obtaining a set of candidate points provided in some embodiments of this application;

[0045] Figure 10 A flowchart for selecting audio track switching points based on a set of candidate points is provided for some embodiments of this application;

[0046] Figure 11 UI diagrams illustrating the acquisition instruction priority provided in some embodiments of this application;

[0047] Figure 12 This is a timing interaction diagram of the audio track switching method provided in some embodiments of this application. Detailed Implementation

[0048] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0049] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0050] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0051] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0052] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0053] Figure 1 This is an operational scenario diagram for audio service processing provided in some embodiments of this application.

[0054] like Figure 1 As shown, the operation scenario may include a server 100 and an electronic device 200. Examples of electronic devices 200 include smart TVs 200a, mobile terminals 200b, smart speakers 200c, etc.

[0055] In this application, server 100 and electronic device 200 can interact with each other through various communication methods. Electronic device 200 can be allowed to communicate via a local area network (LAN), wireless local area network (WLAN), or other networks. Server 100 can provide data related to audio services to electronic device 200; for example, electronic device 200 can preload or download audio data corresponding to an audio track from the server.

[0056] Server 100 can be a server that provides services such as audio playback and audio track switching. Server 100 can receive service requests sent by electronic device 200, such as audio loading requests. In response to the service requests sent by electronic device 200, server 100 sends the corresponding service data (e.g., audio track configuration information, audio data, etc.) to electronic device 200. Server 100 can be a server cluster or multiple server clusters, and can include one or more types of servers.

[0057] Electronic device 200 can be a hardware device or a software device. When electronic device 200 is a hardware device, electronic device generally refers to a device that has at least computing and processing capabilities, audio track switching, and audio output / playback capabilities. Electronic devices may also be configured with other capabilities, such as UI display capabilities. Electronic devices include, but are not limited to: smart TVs, smart speakers, other smart home devices (such as smart refrigerators, smart range hoods, etc.), smartphones, tablets, computers, smart wearable devices, virtual reality devices, augmented reality devices, smart in-vehicle terminals, etc.

[0058] When the electronic device 200 is a software device, it may include at least one software function module / service / model (e.g., multimodal data acquisition module, audio track prediction module, audio track switching module, audio playback module, etc.), and the software device may be applied to the hardware electronic device 200 listed above.

[0059] Figure 2 This is a schematic diagram of the hardware configuration of an electronic device provided in some embodiments of this application.

[0060] like Figure 2 As shown, the electronic device 200 may include at least one of the following: a communication device 210, a detection device 220, a device interface 230, a controller 240, a display 250, an audio output device 260, a memory, a power supply, and a user input interface.

[0061] In some embodiments, the communication device 210 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The electronic device 200 may have multiple communication devices 210 depending on the supported communication methods. For example, when the electronic device 200 supports wireless network communication, it may have a communication device 210 with WiFi functionality. When the electronic device 200 supports Bluetooth connectivity, it needs to have a communication device 210 with Bluetooth functionality.

[0062] In some embodiments, the detection device 220 is used to collect multimodal data. The detection device 220 may include an image acquisition device, a sound acquisition device, a physiological sensor, and an environmental sensor. The image acquisition device, such as a camera, is used to acquire images of the external environment and user-related images (e.g., images of user body movements, user gestures, etc.). The sound acquisition device, such as a microphone, is used to acquire sound data from the external environment and user dialogue data. The physiological sensor is used to detect the user's physiological state (e.g., heart rate, pulse, etc.). The environmental sensor is used to collect environmental data; the environmental sensor may include, but is not limited to, a light receiver, a temperature sensor, and a humidity sensor, to detect light, temperature, humidity, etc. in the environment.

[0063] In some embodiments, if the electronic device has a UI display capability, it further includes a display 250. The display 250 includes display function components for presenting an image and driving components for driving the image display. The display 250 is used to receive and display image signals output from the controller 240. For example, the display 250 can be used to display a media playback interface, an audio playback interface, an audio track list page, etc.

[0064] In some embodiments, the controller 240 may include at least one of a central processing unit, an audio processor, and a power processor, and a first to an nth interface for input / output. If the electronic device has a UI display capability, the controller 240 may also include a video processor and a graphics processor. The controller 240 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 240 is also used to control the overall operation of the electronic device 200, enabling audio track switching and audio playback.

[0065] In some embodiments, the audio output device 260 can be a built-in speaker of the electronic device 200 or an external audio output device connected to the electronic device 200, used to output and play audio. For the external audio output device connected to the electronic device 200, the electronic device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the electronic device 200 to output sound from the electronic device 200.

[0066] Figure 3 This is a schematic diagram illustrating the software configuration of an electronic device provided in some embodiments of this application.

[0067] In some embodiments, such as Figure 3 As shown, the electronic device 200 system can be configured as a three-layer system, from top to bottom as the application layer, middleware layer, and hardware layer.

[0068] In some embodiments, the application layer mainly includes commonly used applications of the electronic device 200, as well as an application framework. The commonly used applications are primarily browser-based applications, such as HTML5 apps, native apps, media playback applications, and audio playback applications.

[0069] In some embodiments, an application framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interface for these functions (toolbar, status bar, menu, dialog box).

[0070] In some embodiments, native apps can support online or offline access, push notifications, or access to local resources.

[0071] In some embodiments, the middleware layer includes various device protocols, multimedia protocols, and system components. Middleware can use the basic services (functions) provided by system software to connect different parts of an application system or different applications on the network, achieving resource sharing and function sharing.

[0072] In some embodiments, the hardware layer mainly includes a HAL interface, hardware, and drivers. The HAL interface is a unified interface for all device chips, with the specific logic implemented by each chip. The drivers mainly include: audio drivers, display drivers, Bluetooth drivers, camera drivers, Wi-Fi drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, pressure sensors, etc.), and power drivers.

[0073] In this embodiment, to improve audio playback quality, operators can configure multiple parallel audio tracks (AudioTracks, or simply tracks) for the audio content of media assets based on acoustic characteristics. These tracks store pre-recorded audio signals, such as vocals, instruments, and different musical styles (e.g., rock, jazz), carrying audio waveforms and MIDI (Musical Instrument Digital Interface). Track attributes can be defined during track creation, such as timbre, volume, input / output ports, and the number of channels. Electronic devices can switch between tracks to change the played audio signal.

[0074] Taking the electronic device 200, which also has display capabilities, as an example, when a user requests the first media asset, the electronic device 200 can display the media asset playback interface corresponding to the first media asset. The first media asset can be an audio / video type or a purely audio type.

[0075] Figure 4 This is a schematic diagram of a media asset playback interface provided in some embodiments of this application.

[0076] See Figure 4 Figure (a) shows an audio settings control 41 set in the media playback interface. The electronic device responds to the instruction to trigger the audio settings control 41, see [reference needed]. Figure 4 Figure (b) shows that the electronic device controls the display 250 to display an audio track list page 41a on the media playback interface. The audio track list page 41a includes at least one audio track category, such as audio track 1, audio track 2, and audio track 3. Each audio track category has a corresponding selection control, allowing the user to select the track category of interest according to their listening preferences.

[0077] See Figure 4 In Figure (b), assuming the current audio track category is track 1, if the user wants to switch to track 3, they click the selection control corresponding to track 3. See Figure (b). Figure 4 The diagram (c) shows how the current audio track category can be switched to track 3. For example, if track 1 is rock background music, and the user wants to relax, they can switch to track 3, which is white noise or soothing light music. This method requires manual intervention from the user when switching tracks, resulting in low efficiency.

[0078] In some embodiments, electronic devices can be configured with audio feature matching-based track switching technology. By extracting audio features and matching them with features contained in a preset track feature library, the target track can be located and switched to. The drawbacks of this track switching method include: single-modal dependence, relying solely on volume or simple timing for audio switching, failing to perceive changes in user behavior and environmental states, resulting in the matched target track not accurately matching the user's actual preferences and being unsuitable for complex application environments; lack of scenario prediction, unable to predict the target track based on user behavior trends, leading to low track switching efficiency; and abrupt switching, with perceptible auditory gaps and significant track switching delays during cross-track switching. In some embodiments, electronic devices can be configured with an audio timing-based track switching scheme. This scheme dynamically captures the timing feature vector of the audio signal through a sliding window, compares the similarity of this timing feature vector with a preset track template, and then matches the target track with the highest similarity. The drawbacks of this audio track switching method include: single-modal dependence, which only performs audio switching based on volume or simple timing, and cannot perceive changes in user behavior and environmental state, resulting in the target audio track not accurately matching the user's actual preferences and being unable to adapt to complex application environments; the use of a fixed-time sliding window for audio track switching ignores the semantic coherence of audio content, and the error switching rate is high under environmental noise interference, resulting in low accuracy of audio track switching.

[0079] Figure 5 A flowchart illustrating an audio track switching method provided in some embodiments of this application.

[0080] The purpose of this audio track switching method is to collect multimodal data including user behavior and environmental conditions, perform contextual prediction on the multimodal data to predict the target audio track that matches the user's actual preferences, automatically match the switching timing of the target audio track, and automatically switch to the target audio track based on the switching timing. This eliminates the user's auditory blockage, reduces audio track switching latency, and eliminates the need for manual user intervention in audio track switching. Instead, it automatically matches and switches to the target audio track based on multimodal data, improving audio track switching efficiency. This method can be executed by the controller 240 by running relevant applications (such as media asset playback applications, audio playback applications, etc.). The method includes:

[0081] Step S51: When playing the first audio stream corresponding to the first audio track, multimodal data is collected by the detection device.

[0082] Step S52: Normalize the multimodal data using a convolutional neural network model to obtain multimodal feature vectors. In some embodiments, see... Figure 2 The detection device 220 may include an image acquisition unit, a sound acquisition unit, a physiological sensor, and an environmental sensor. The image acquisition unit, such as a camera, is used to acquire images of the external environment and user-related images (e.g., images of user body movements and gestures). The sound acquisition unit, such as a microphone, is used to acquire sound data from the external environment and user conversation data. The physiological sensor is used to detect the user's physiological state (e.g., heart rate, pulse). The environmental sensor is used to acquire environmental data; the environmental sensor may include, but is not limited to, a light receiver, a temperature sensor, and a humidity sensor to detect light, temperature, and humidity in the environment.

[0083] By capturing user behavior, environmental conditions, and user voice through detection devices, multimodal data is divided into visual modal data (such as user body movements and visual direction captured by visual sensors / cameras), auditory modal data (such as rhythm, emotion, and semantic coherence involved in user dialogue), and motion modal data (such as changes in environmental scenes (including changes in environmental conditions such as noise, temperature, humidity, and lighting), changes in the number of people in the space, changes in user movement trajectory, and changes in user heart rate during exercise). This multimodal data can provide multiple dimensions and factors for subsequent prediction of target audio tracks, making the automatically matched target audio tracks more in line with the user's actual habits and preferences.

[0084] Figure 6 This is a schematic diagram illustrating the principle of feature fusion for multimodal data provided in some embodiments of this application.

[0085] See Figure 6 The data collected by the image acquisition device is referred to as visual modal data, the data detected by the physiological and environmental sensors are collectively referred to as motion modal data, and the data collected by the sound acquisition device is referred to as auditory modal data. Thus, the multimodal data collected by the detection device 220 includes visual modal data, auditory modal data, and motion modal data. This multimodal data contains multidimensional information such as user behavior, user movement trajectory, user dialogue, and environmental status. Based on this multimodal data, the target audio track is predicted, making the target audio track more consistent with the user's actual preferences.

[0086] In some embodiments, see Figure 6Visual modal data is processed with a 3D CNN (3D Convolutional Neural Network) to extract spatiotemporal features through sliding convolution; auditory modal data is processed with a 2D CNN (2D Convolutional Neural Network) to extract sound features through sliding convolution; and motion modal data is processed with a 1D CNN (1D Convolutional Neural Network) to extract motion features through sliding convolution.

[0087] In some embodiments, see Figure 6 The spatiotemporal features, acoustic features, and motion features are aligned as vectors and then normalized to obtain a multimodal feature vector. This allows the multimodal data to be fused and transformed into a feature vector that can be used by computers to predict target audio tracks based on context-based prediction models.

[0088] Step S53: Input the multimodal feature vector into the context prediction model, and the context prediction model predicts the target audio track that matches the multimodal data.

[0089] Figure 7 Prediction process of the context prediction model provided for some embodiments of this application Figure 1 .

[0090] In some embodiments, see Figure 7 The electronic device 200 can store and update a context memory pool in its memory. The context memory pool stores playback history information, including audio track information related to different user habits and environmental states. The context memory pool can be continuously accumulated and updated based on historical audio tracks and playback data. It acts as a "knowledge base" or "experience base," providing prior knowledge for predicting target audio tracks, improving the accuracy of target audio track prediction, and consequently improving the accuracy of automatic audio track matching and switching.

[0091] In some embodiments, see Figure 7 The context prediction model can access the context memory pool. Based on information such as user behavior and environmental status involved in the multimodal data, it retrieves the playback history information (hereinafter referred to as: target playback history information) associated with the multimodal data in the context memory pool, extracts the features of the target playback history information and performs vectorization processing to obtain the context feature vector.

[0092] In some embodiments, see Figure 7 The context prediction model concatenates the multimodal feature vector obtained in step S52 with the context feature vector to obtain the concatenated vector.

[0093] In some embodiments, see Figure 7 The context prediction model utilizes Graph Convolutional Neural Networks (GNNs) and Transformers to process the concatenated vector, obtaining spatiotemporal joint features. An LSTM (Long Short-Term Memory) encoder and regression network are then used to process these spatiotemporal joint features, providing a prediction result that includes the category of the target audio track to be switched to. This utilizes the target playback history information associated with the context memory pool as prior knowledge for subsequent target audio track predictions, improving the accuracy of target audio track prediction and consequently enhancing the accuracy of automatic audio track matching and switching. The Transformer, used in natural language processing, is a deep learning model based on attention mechanisms for processing sequential data. A Transformer can include an encoder and a decoder; the encoder converts the input data into vectors, and the decoder generates and outputs sentences based on these vectors, while retaining important information from the input sequence.

[0094] Figure 8 Prediction process of the context prediction model provided for some embodiments of this application Figure 2 .

[0095] In some embodiments, see Figure 8 The context prediction model can include a spatiotemporal fusion module. After obtaining the spliced ​​vector, the spatiotemporal fusion module can obtain the spatial and temporal features in the spliced ​​vector.

[0096] In some embodiments, see Figure 8 The spatiotemporal fusion module can perform spatial graph convolution operations on spatial features and graph convolutional neural networks (GNNs) to obtain spatial modal features. These spatial modal features are used to characterize the spatial relationships between objects, including distances and relative positions between objects.

[0097] In some embodiments, see Figure 8 The spatiotemporal fusion module can perform temporal transformation operations on temporal features and recurrent neural networks (RNNs) to obtain temporal modal features. Temporal modal features are used to characterize the dependencies between objects in a time series.

[0098] In some embodiments, see Figure 8 The spatiotemporal fusion module can perform cross-modal fusion of spatial modal features and temporal modal features to obtain and output spatiotemporal joint features based on multimodality.

[0099] In some embodiments, see Figure 8 The context prediction model can include a prediction output layer, which is used to predict spatiotemporal joint features and output the prediction results. The context prediction model can be configured to perform tasks on the prediction output layer based on multimodal spatiotemporal joint features, including predicting user behavior, predicting user trajectory sequences, and confidence prediction.

[0100] In some embodiments, see Figure 8 The prediction output layer can utilize fully connected layers and softmax (normalized exponential function) to predict user behavior based on spatiotemporal joint features, obtaining the probability distribution of user behavior categories. Here, softmax is a type of logistic function. The softmax function "compresses" a K-dimensional vector z containing arbitrary real numbers into another K-dimensional real vector σ(z), such that each element's value is between (0,1), and the sum of all elements equals 1.

[0101] In some embodiments, the prediction output layer predicts the probability distribution P(Behavior|Observed History) of the most likely behavior category that the target entity will perform within a future time period T_obs, specifically configured as follows:

[0102] (A1) Obtain the observation history S_obs=[s_{t-T_obs+1},...,s_t], where T_obs is the prediction time (e.g., 3s).

[0103] (A2) Convert S_obs into a fixed-length vector X_obs using a convolutional neural network (CNN).

[0104] (A3) Input the encoded feature vector X_obs into a K-dimensional fully connected layer, where the input of the first layer is h^{(0)}=X_obs.

[0105] (A4) Obtain the input h^{(l)} of the first layer of the fully connected layer, h^{(l)} = f(W^{(l)} * h^{(l-1)} + b^{(l)}). Where W^{(l)} is the weight matrix, b^{(l)} is the bias vector, f is the transformer, and l is a natural number from 1 to K, thus obtaining the output dimension z(l) of the fully connected layer.

[0106] (A5) The Softmax algorithm is used to convert z(l) into a probability distribution. In the specific implementation, P(Behavior=k|X_obs)=softmax(z)_k=exp(z_k) / Σ_{j=1}^K exp(z_j), and the probability distribution of the K-dimensional vector is denoted as P=[p_1,p_2,...,p_K]. Here, p_k represents the probability of the predicted entity performing the k-th action, and Σp_k=1.

[0107] (A6) Obtain the category with the highest probability argmax_k(p_k) and use argmax_k(p_k) as the final predicted behavior category.

[0108] In some embodiments, see Figure 8 The prediction output layer uses an LSTM decoder to predict the user trajectory sequence within a preset time period in the future. The user trajectory sequence is the sequence of location coordinates of the user's movement trajectory.

[0109] In some embodiments, the prediction output layer predicts the sequence of position coordinates of the target entity over the next T_pred time steps. The implementation steps are as follows:

[0110] (B1) Set the initial hidden state h_0 and cell state c_0 of the LSTM decoder.

[0111] (B2) Set the first prediction input_0.

[0112] (B3) Concatenate the predicted output input_{τ-1} from step (B2) with the context feature vector C to obtain decoder_input_τ, i.e. decoder_input_τ = concat([input_{τ-1},C]).

[0113] (B4) Input decoder_input_τ into the LSTM decoder, and combine the current hidden state h_{τ-1} and cell state c_{τ-1} to calculate the new hidden state h_τ and cell state c_τ, i.e. (h_τ,c_τ)=LSTM_Cell(decoder_input_τ,(h_{τ-1},c_{τ-1})).

[0114] (B5) The new hidden state h_τ of LSTM is processed by a fully connected layer to obtain output_τ, that is, output_τ=W_out*h_τ+b_out.

[0115] (B6) Set prediction coordinates

[0116] (B7) Predict the coordinates Add to Set the input for the next step, input_τ

[0117] (B8) After completing the loop of time step T_pred, the complete predicted trajectory sequence is obtained.

[0118] In some embodiments, see Figure 8The prediction output layer can determine the first candidate audio track based on the probability distribution of user behavior categories and user trajectory sequences. The first candidate audio track consists of several audio track categories that are most likely to occur and conform to user behavior trends and preferences, as predicted by context. That is, the first candidate audio track includes at least one audio track category, and the target audio track of the final prediction output is one of the audio track categories in the first candidate audio track.

[0119] In some embodiments, see Figure 8 The prediction output layer calculates the confidence score of the first candidate audio track, and normalizes the confidence score based on the regression network, outputting the percentage confidence score Confidence(i). Here, i represents the index of the first candidate audio track, 1≤i≤M, and M is the total number of first candidate audio tracks.

[0120] In some embodiments, the prediction output layer can obtain a second candidate audio track based on a confidence threshold. The confidence of the second candidate audio track is greater than the confidence threshold, that is, candidate audio tracks with higher confidence are selected, and unreliable candidate audio tracks with low confidence are filtered out. The confidence threshold is not limited, and can be set to 0.85 for example.

[0121] In some embodiments, if the prediction output layer fails to acquire a second candidate audio track, i.e., all Confidence(i) values ​​are less than the confidence threshold, then multimodal data can be reacquired and context prediction can be performed.

[0122] In some embodiments, if the prediction output layer obtains at least one second candidate audio track, it obtains the maximum confidence among the second candidate audio tracks, i.e., Confidence(max) = MAX{Confidence(j)|1≤j≤N}, where j represents the index of the second candidate audio track and N is the total number of second candidate audio tracks. In this way, the prediction output layer can determine the second candidate audio track corresponding to Confidence(max) as the target audio track for the final prediction output.

[0123] In some embodiments, for example, if an image sensor detects a user walking from the living room sofa to the kitchen, and a sound sensor captures the sound of running water after the kitchen faucet is turned on, then the context prediction model predicts a target audio track example as a recipe-based voice guide. This allows for better guidance of the user to replicate the cooking process using the recipe by switching to the recipe-based voice guide.

[0124] In some embodiments, for example, if an image sensor detects that a user is moving from the living room to the bedroom, and an environmental sensor detects a decrease in bedroom lighting, then the context prediction model predicts a target audio track as a sleep mode, such as white noise or soothing soft music. In this way, by switching to sleep mode, a suitable sound environment for rest and sleep is provided to the user.

[0125] In some embodiments, such as when an image sensor detects a vehicle entering a tunnel, an environmental sensor detects a sudden drop in ambient light, and a sound sensor detects increased ambient noise, the context prediction model predicts a target audio track example in noise reduction mode. Thus, by switching to noise reduction mode, the volume and clarity of the in-vehicle voice can be improved, ensuring that users can clearly hear the in-vehicle voice content even in a tunnel environment.

[0126] The context prediction model can call upon a context memory pool, which is a prior knowledge pool used to store audio track information related to user habits and environmental conditions during historical usage. This provides a knowledge reserve for subsequent prediction of target audio tracks. By performing spatiotemporal fusion based on multimodal feature vectors and the context memory pool, spatiotemporal joint features are obtained. Then, user behavior prediction can be performed, outputting the probability distribution of user behavior categories and predicting the user trajectory sequence within a preset time period. Through user behavior and action trajectory, multimodal and multidimensional analysis is achieved, thereby accurately predicting which audio playback trends / possibilities may exist based on the user behavior and action trajectory. This provides at least one first candidate audio track, calculates the confidence of the first candidate audio track, and if there is a second candidate audio track with a confidence value greater than the threshold, the second candidate audio track with the highest confidence value is selected as the target audio track. This can filter out the audio track category with the highest probability / trend, improve the accuracy of target audio track prediction, and thus improve the accuracy of automatic audio track matching and switching, and reduce the audio track mis-switching rate. In addition, by automatically predicting and matching target audio tracks through multimodal data, users do not need to manually switch audio tracks, thus improving the efficiency of audio track switching.

[0127] In some embodiments, after predicting the target audio track, a second audio stream corresponding to the target audio track can be preloaded. The controller can preload this second audio stream from an audio buffer pool.

[0128] In some embodiments, an audio buffer pool can be configured on the server side to store audio data corresponding to each audio resource. Thus, after the prediction output layer outputs a prediction result containing the target audio track, the electronic device can send an audio data retrieval request to the server based on the audio ID corresponding to the target audio track. In response to the audio data retrieval request, the server retrieves the audio data matching the audio ID from the audio buffer pool and pushes the audio data to the electronic device as a second audio stream, thereby completing the preloading of the second audio stream before switching tracks.

[0129] In some embodiments, the audio buffer pool can be configured on the electronic device. After the prediction output layer outputs a prediction result containing the target audio track, the electronic device can send an audio data retrieval request to the server based on the audio ID corresponding to the target audio track. In response to the audio data retrieval request, the server retrieves the audio data matching the audio ID and pushes the audio data to the electronic device as a second audio stream. Upon receiving the second audio stream, the electronic device caches it in its local audio buffer pool, thereby preloading the second audio stream before switching audio tracks. In this way, when switching audio tracks subsequently, the player can quickly retrieve and load the second audio stream from the local audio buffer pool, thereby improving the start-up speed of the second video stream corresponding to the target audio track and increasing the efficiency of audio track switching.

[0130] Step S54: Detect the silence segment, zero-crossing point, energy stability segment, and harmonic stability segment of the first audio stream.

[0131] After determining the category of the target audio track, it is also necessary to further determine the timing of switching to the target audio track. Before switching tracks, the electronic device 200 is playing the first audio stream corresponding to the first audio track. At this time, the audio attributes of the first audio stream can be analyzed to match the switching point of the target audio track.

[0132] In some embodiments, the controller 240 of the electronic device can preprocess and extract features from the first audio stream to obtain relevant feature information of the first audio stream, which can be used to analyze audio attributes. The relevant feature information includes, but is not limited to, sampling rate, frame length, frame shift, spectrum, harmonic noise ratio, etc. Audio attributes include, but are not limited to, silence segment length, zero-crossing rate, energy fluctuation range, and fundamental frequency change, etc. The silence segment length measures the duration the first audio stream remains silent; the zero-crossing rate characterizes the number of times the audio signal crosses zero points per unit time, measuring the frequency of audio signal changes; the energy fluctuation range measures the energy stability of the first audio stream; and the fundamental frequency change rate measures the harmonic stability of the first audio stream. Thus, based on the audio attributes of the first audio stream, silence segments, zero-crossing rate points, energy stability segments, and harmonic stability segments of the first audio stream can be detected, providing a reference for subsequently obtaining a candidate point set.

[0133] Step S55: Obtain the candidate point set.

[0134] In some embodiments, the candidate point set includes at least one of a first candidate point, a second candidate point, a third candidate point, and a fourth candidate point. The first candidate point is selected from silent segments, the second candidate point is selected from zero-crossing points of the first audio stream, the third candidate point is selected from energy-stable segments of the first audio stream, and the fourth candidate point is selected from harmonic-stable segments of the first audio stream. By detecting silent segments of the first video stream under the current first audio track and calculating audio attributes such as the zero-crossing rate, energy fluctuation range, and fundamental frequency change rate of the audio frames, a basis for obtaining the candidate point set is provided. This enables audio track switching at silent frames, continuous consonant frames, energy-stable frames, or harmonic-stable frames, avoiding audio track switching in vocal frequency bands, strong rhythmic intervals, vowel formant intervals, and emotionally intense areas, thereby improving the audio quality of audio transitions and connections during track switching.

[0135] Step S56: Select a target candidate point from the candidate point set. The target candidate point is the switching point when switching from the first audio track to the target audio track.

[0136] Figure 9 This is a flowchart illustrating the process of obtaining a set of candidate points according to some embodiments of this application.

[0137] In some embodiments, see Figure 9 The system identifies silent segments in the first audio stream and detects their length. If the length of the silent segment exceeds a first threshold, it indicates that the first audio stream contains a relatively long, continuous silent region, where the midpoint of the silent segment is a relatively stable silent frame (both before and after it are silent). The midpoint of the silent segment is then used as the first candidate point, thus using the silent frame as the switching point to avoid switching audio tracks in non-continuously silent regions such as the vocal frequency band. The first threshold is not limited; for example, it can be 100ms.

[0138] In some embodiments, see Figure 9 If the length of the silent segment is not greater than the first threshold, it indicates that the first audio stream does not have a continuous silent area. Based on the waveform of the first audio stream, the number of zero-crossing points per unit time is detected. If the zero-crossing rate of a preset number of consecutive audio frames is less than the second threshold, it indicates that the preset number of audio frames are consonant frames, and the first audio stream exhibits a discontinuous and unclear playback state. The waveform of the first audio stream may include multiple zero-crossing points, which have peak and valley points. The valley points among the zero-crossing rate points are natural pause points. Therefore, the valley points of the detected zero-crossing rate points are used as the second candidate points, thereby avoiding switching audio tracks in the vowel formant interval. The preset number and the second threshold are not limited. To ensure the accuracy of the second candidate point selection and reduce the computational load on the controller, the preset number should not be too large or too small; for example, a preset number of 3 frames.

[0139] In some embodiments, see Figure 9If the zero-crossing rate of a preset number of consecutive audio frames is not less than a second threshold, then it is determined whether the energy fluctuation range of the first audio stream is within the threshold range. If the energy fluctuation range is within the threshold range, it indicates that the energy fluctuation of the first audio stream is small and the energy is relatively stable. Then, the starting point of the stable energy segment of the first audio stream is taken as the third candidate point, thereby avoiding switching audio tracks in strong rhythm intervals. The energy fluctuation range is not limited, for example, it is [-3dB, +3dB].

[0140] In some embodiments, see Figure 9 If the energy fluctuation range exceeds the threshold range, it can be determined whether the fundamental frequency change rate of the first audio stream is less than the third threshold. If the fundamental frequency change rate is not less than the third threshold, it indicates that there are currently no candidate points that meet the conditions, and the first audio stream can be re-analyzed to select candidate points again. If the fundamental frequency change rate is less than the third threshold, it indicates that the harmonics of the first audio stream are relatively stable, and the first audio stream presents a relatively soothing or calm emotional state rather than a relatively strong emotional state. In this case, the midpoint of the harmonic stability segment of the first audio stream is taken as the fourth candidate point. The third threshold is not limited, for example, it can be 2%.

[0141] The candidate point set obtained through the above selection mechanism includes at least one of the first, second, third, and fourth candidate points. This allows for audio track switching at audio silence frames, continuous consonant frames, energy-stable frames, or harmonic-stable frames, avoiding switching audio tracks in vocal frequency bands, strong rhythm ranges, vowel formant ranges, and emotionally intense areas. By selecting a target candidate point from the candidate point set and determining the timing of switching to the target audio track based on that target candidate point, the audio quality of audio connection and transition during track switching is improved.

[0142] Figure 10 This is a flowchart illustrating the selection of audio track switching points based on a set of candidate points, provided for some embodiments of this application.

[0143] In some embodiments, see Figure 10 After obtaining the candidate point set, assuming the candidate point set is {D1, D2, D3, D4}, where D1 is the first candidate point, D2 is the second candidate point, D3 is the third candidate point, and D4 is the fourth candidate point. For each candidate point in the candidate point set (referred to as: preliminary candidate point), audio features of a preset length (not limited, for example, 500ms) before and after each preliminary candidate point are extracted to obtain the candidate point audio feature set {D1′, D2′, D3′, D4′}.

[0144] In some embodiments, see Figure 10The Transformer encoder is used to process the audio feature set of candidate points {D1′,D2′,D3′,D4′}, calculate the quality score corresponding to each candidate point in the candidate point set, and normalize the quality score to obtain the quality score set {E1,E2,E3,E4}.

[0145] In some embodiments, see Figure 10 In the quality score set, the first quality score that is less than the fourth threshold (not limited, for example, 0.7) is selected. The preliminary candidate point corresponding to the first quality score is the unqualified candidate point. In this way, unqualified candidate points can be filtered out from the candidate point set, and the remaining candidate points are the qualified candidate points.

[0146] In some embodiments, see Figure 10 If only one candidate point A remains in the filtered candidate point set, then candidate point A is taken as the target candidate point. For example, in {E1,E2,E3,E4}, E2, E3, and E4 are all less than the fourth threshold, and E1 is greater than the fourth threshold. E1 corresponds to candidate point D1, that is, the filtered candidate point set is {D1}, then candidate point D1 is determined as the target candidate point.

[0147] In some embodiments, see Figure 10 If multiple candidate points remain after filtering, the distance between any two adjacent candidate points can be calculated. Adjacent candidate points with a distance less than a fifth threshold (not limited, for example, 200ms) are then merged to obtain the target candidate point. For example, if the filtered candidate point set is {D1, D3}, and the distance between D1 and D3 is less than the fifth threshold, then D1 and D3 are merged to obtain the merged target candidate point D.

[0148] In some embodiments, when merging D1 and D3, D3 can be merged into D1, that is, candidate point D1 can be used as the target candidate point D. Alternatively, D1 can be merged into D3, that is, candidate point D3 can be used as the target candidate point D. Alternatively, the median value between D1 and D3 can be calculated, and the node corresponding to the median value can be used as the target candidate point D.

[0149] It should be noted that, due to the different attributes of different audio, the candidate point set is a subset of {D1,D2,D3,D4}. That is, the candidate point set may only include some candidate points from D1,D2,D3,D4. Threshold filtering based on quality score can further reduce the number of available candidate points. Finally, after merging adjacent points, it is basically possible to match only one target candidate point.

[0150] For each candidate point in the candidate point set, audio features of a preset length before and after the candidate point are extracted. The Transformer encoder is used to process the audio features of the preset length before and after the candidate point to predict the quality score of the candidate point. The quality score of the candidate point is compared with a fourth threshold. If the quality score of the candidate point is less than the fourth threshold, it indicates that the candidate point is of unqualified quality. The unqualified candidate point is then filtered out from the candidate point set, so that the candidate point set only contains qualified candidate points. If there are multiple candidate points remaining in the candidate point set, adjacent candidate points with a spacing of less than a fifth threshold can be merged. In this way, by using the audio attributes of the first audio stream (such as silence, filtering rate, harmonic stability, energy stability), candidate point quality screening, and merging of adjacent candidate points, a target candidate point is finally obtained. This target candidate point is the optimal audio track switching point, thereby improving the audio track switching effect, realizing a natural transition and seamless audio connection, eliminating the auditory blockage problem during audio switching, reducing audio track switching latency, and improving the speed and accuracy of audio track switching.

[0151] Step S57: Phase alignment is performed on the first audio track and the target audio track, and the target audio track is switched at the target candidate point to play the second audio stream corresponding to the target audio track.

[0152] Due to differences in recording conditions (such as microphone position and device latency) or processing methods (such as mixing plugin delays) between different audio tracks, sound waves of the same frequency may exhibit phase differences. When the phase difference approaches 180° (i.e., out of phase), the sound waves will cancel each other out, resulting in a loss of energy in specific frequency bands (especially low frequencies), causing the audio sound to sound weak. Conversely, when in phase, certain frequencies may be over-amplified, disrupting the audio playback balance. To address this, phase alignment can be performed on the first and target audio tracks to improve audio transitions, avoid audio attenuation or distortion caused by phase cancellation, eliminate auditory gaps during audio transitions, reduce track switching latency, and achieve seamless, lossless, and smooth transitions between tracks, thereby improving the speed and accuracy of track switching.

[0153] In some embodiments, the phase alignment process may include: time-domain alignment, frequency-domain correction, and phase compensation. Time-domain alignment includes: performing cross-correlation and time delay compensation on the first and target audio tracks, and then aligning the waveforms of the first and target audio tracks. Frequency-domain correction includes: after aligning the waveforms of the tracks, performing a Fast Fourier Transform (FFT) on the waveforms of the first and second audio tracks to obtain the phase spectra corresponding to the two tracks, then analyzing the phase spectra to calculate the phase difference matrix between the two tracks, and then performing phase rotation based on the phase difference matrix. Time delay compensation includes: after phase rotation, performing a Hilbert Transform on the frequency domain signals of the two tracks to obtain an analytic signal, performing phase offset correction based on the analytic signal, and finally reconstructing the waveform of the target audio track. Specific implementation methods for phase alignment can be found in related technologies, and will not be elaborated further in this embodiment.

[0154] In some embodiments, when an electronic device plays news audio, assuming the image sensor detects a user walking from the living room sofa to the kitchen, and the sound sensor captures the sound of running water after the kitchen faucet is turned on, the context prediction model predicts a target audio track example as a recipe voice guidance. This allows for better guidance of the user to replicate the cooking process by switching to the recipe voice guidance. By analyzing the audio attributes of the news audio, the target candidate point is determined to be the silent frame at the end of the news sentence (the final sound). Through phase alignment, the fundamental frequency (e.g., 120Hz) of the recipe voice guidance audio is aligned to the frequency (e.g., 115Hz) of the news sentence's final sound, thereby switching to the recipe voice guidance audio at the silent frame of the news sentence's final sound.

[0155] In some embodiments, to achieve precise control of audio track switching, the controller can generate an audio track switching command based on the category of the target audio track and the switching timing corresponding to the target candidate point, and send the audio track switching command to the switching actuator. In response to the audio track switching command, the switching actuator switches to the target audio track at the target candidate point to play the second audio stream corresponding to the target audio track.

[0156] In some embodiments, the switching actuator can be a software program or hardware for switching audio tracks. For example, the switching actuator can be a software program such as an audio player or media asset player, or it can be an audio output device (such as a built-in speaker or an external amplifier).

[0157] In some embodiments, when executing a track switching instruction, the switching actuator can transition from the first track to the target track at the target candidate point using a preset transition method. The preset transition method is not limited; for example, it can employ fade-in / fade-out, crossfade, or other methods to achieve seamless and smooth track switching, thus improving the transition effect during track switching.

[0158] In some embodiments, fade-in and fade-out refer to adjusting the keyframes of the audio level (i.e., volume) to gradually increase the audio volume from silence to normal (i.e., fade-in) or gradually decrease the audio volume from normal to silence (i.e., fade-out), thereby avoiding the abrupt auditory impact caused by the sudden appearance or disappearance of audio during abrupt transitions; controlling the emotional tension through gradual speed, such as creating suspense with a slow fade-in and creating an abrupt ending with a fast fade-out, thereby enhancing the emotional expression of the audio; the fade-in and fade-out of multiple audio tracks can form a stereo sound field and enhance the sense of content layering.

[0159] In some embodiments, crossfading refers to a technique that achieves a smooth transition by controlling the volume gradation of multiple audio tracks, eliminating abruptness and discontinuity during audio switching. For example, while the volume of the current audio track 1 is linearly attenuated, the volume of the new audio track 2 is linearly increased, and audio tracks 1 and 2 are played overlapping within a set time period to create a smooth transition effect.

[0160] Figure 9 The provided audio track switching solution captures user behavior, environmental conditions, and user voice through detection devices, thereby dividing multimodal data into visual modal data (such as user body movements and visual direction collected by visual sensors / cameras), auditory modal data (such as rhythm, emotion, and semantic coherence involved in user dialogue), and motion modal data (such as changes in environmental scenes (including changes in environmental conditions such as noise, temperature, humidity, and lighting), changes in the number of people in the space, changes in user movement trajectory, and changes in user heart rate during exercise). This collection of multimodal data can provide multiple dimensions and factors for subsequent prediction of target audio tracks, making the automatically matched target audio tracks more in line with the user's actual habits and preferences. Context-based prediction models can perform multi-dimensional analysis of multimodal feature vectors, thereby automatically predicting and matching target audio tracks without requiring manual user intervention in track switching, thus improving track switching efficiency. By detecting silent segments in the first video stream under the current first audio track, audio attributes such as zero-crossing rate, energy fluctuation range, and fundamental frequency change rate of audio frames are calculated, providing a basis for obtaining a candidate point set. This enables audio track switching at silent frames, continuous consonant frames, energy-stable frames, or harmonic-stable frames, avoiding switching in the human voice frequency band, strong rhythm range, vowel formant range, and... The system switches audio tracks based on areas of strong emotion. This involves selecting a target candidate point from a set of candidate points and determining the timing of switching to the target audio track based on that target candidate point. This eliminates auditory gaps during audio switching, reduces track switching latency, and improves the audio quality of audio transitions and connections during track switching. By aligning the phase of the first and target audio tracks, the system achieves better audio transitions during track switching, avoiding audio attenuation or distortion caused by phase cancellation. This results in seamless, lossless, and smooth switching between audio tracks, improving the efficiency and accuracy of track switching.

[0161] In some embodiments, before executing step S57, the controller can generate a first audio track switching instruction based on the target audio track predicted by the currently acquired multimodal data and the target candidate point used when switching to the target audio track. In response to the first audio track switching instruction, the switching actuator transitions from the first audio track to the target audio track at the target candidate point using a preset transition method, thereby playing the second audio stream corresponding to the target audio track.

[0162] Since user behavior and environmental state are variables, multimodal data may be updated as user behavior, trajectory and environmental state change. When multimodal data is updated, the controller can send new audio track switching instructions to the switching actuator. Thus, the switching actuator may receive one or more audio track switching instructions in a timing sequence.

[0163] In some embodiments, before sending the first audio track switching instruction (audio track switching instruction A) to the switching actuator, the controller may detect whether the switching actuator currently has a second audio track switching instruction (audio track switching instruction B) being executed.

[0164] In some embodiments, if the switching actuator is not currently executing any audio track switching instruction, i.e., there is no second audio track switching instruction, the controller can directly send the first audio track switching instruction to the switching actuator.

[0165] In some embodiments, if the switching actuator is currently executing a second audio track switching instruction, the controller needs to obtain the instruction priority between the first and second audio track switching instructions, and control the switching actuator to execute the higher-priority audio track switching instruction according to the instruction priority. This provides a reference for the switching actuator to execute audio track switching instructions, improving the accuracy of multi-instruction execution in dynamic multimodal scenarios.

[0166] Figure 11 This is a UI diagram illustrating the acquisition instruction priority provided in some embodiments of this application.

[0167] In some embodiments, assuming the electronic device has a UI display capability, the priority of user-defined instructions can be obtained through a UI query mechanism. See also Figure 11 When the controller generates track switching instruction A, and the switching actuator is currently executing track switching instruction B, the controller can control the display to show the instruction priority setting page 110.

[0168] In some embodiments, see Figure 11The instruction priority setting page 110 includes a prompt message 111, a switch button 112, and a no-switch button 113. The prompt message 111 prompts the user to set the priority of audio track switching instruction A and audio track switching instruction B. For example, a sample prompt message 111 could be: "Switching to audio B. The system analyzes that you may be interested in audio track A. Switch to audio track A?"

[0169] In some embodiments, see Figure 11 After viewing prompt 111, if a user wants to switch to audio track A, indicating a preference for track A, they can trigger switch button 112 to set track A to have a higher priority than track B. In response to the triggering of switch button 112, the controller sets the track switching instruction A corresponding to track A to a high priority and the track switching instruction B corresponding to track B to a low priority. That is, the instruction priority is: track switching instruction A > track switching instruction B. In this way, the controller can write the priorities into the track switching instructions respectively, marking high priority in the newly generated track switching instruction A and low priority in the old track switching instruction B.

[0170] In some embodiments, see Figure 11 After viewing prompt 111, if a user does not wish to switch to audio track A, indicating a preference for audio track B and an unwillingness to interrupt the switching and playback of track B, they can trigger the "Do Not Switch" button 113. In response to triggering the "Do Not Switch" button 113, the controller sets the audio track switching instruction A corresponding to audio track A to a low priority and the audio track switching instruction B corresponding to audio track B to a high priority, i.e., the instruction priority is: audio track switching instruction B > audio track switching instruction A. In this way, the controller can write the priorities into the audio track switching instructions respectively, marking low priority in newly generated audio track switching instruction A and high priority in older audio track switching instruction B. By soliciting instruction priorities from the user, this provides a reference for the switching executor to execute audio track switching instructions, improving the accuracy of multi-instruction execution in dynamic multimodal scenarios and making audio track switching more in line with user habits and preferences.

[0171] In some embodiments, the controller can also control the audio output device to broadcast the text content corresponding to the prompt message 111 in voice form, and obtain the priority voice command input by the user, such as the user saying "Do not switch to audio track A" or "Do not switch to audio track B". In this way, the controller obtains the command priority by performing semantic parsing on the priority voice command.

[0172] In some embodiments, the controller can preset priority rules. These priority rules are configured, for example, to assign higher priority to later-arriving track switching instructions and lower priority to earlier-arriving instructions, based on the order in which the track switching instructions are generated. This ensures that the switching executor always executes the latest track switching instruction. The method for obtaining instruction priority is not limited to the preceding embodiments.

[0173] In some embodiments, assuming that the first audio track switching instruction corresponds to a first instruction priority and the second audio track switching instruction corresponds to a second instruction priority, the first instruction priority and the second instruction priority can be compared. If the first instruction priority is higher than the second instruction priority, the switching executor stops executing the second audio track switching instruction and switches to executing the first audio track switching instruction, that is, interrupting the previously executed low-priority second audio track switching instruction and instead executing the high-priority first audio track switching instruction.

[0174] In some embodiments, if the priority of the first instruction is higher than the priority of the second instruction, the switching executor, after stopping the execution of the second audio track switching instruction and executing the first audio track switching instruction, may write the first playback history information to the context memory pool. The first playback history information is used to indicate that, under the user behavior and environmental conditions corresponding to the currently collected multimodal data, the user tends to switch to the target audio track (the currently predicted output audio track A).

[0175] In some embodiments, if the priority of the first instruction is lower than that of the second instruction, the switching executor continues to execute the first high-priority second track switching instruction, and does not execute the later low-priority first track switching instruction, so as to avoid the low-priority instruction interfering with and interrupting the execution process of the high-priority instruction.

[0176] In some embodiments, if the priority of the first instruction is lower than that of the second instruction, the switching executor may write second playback history information to the context memory pool while continuing to execute the second audio track switching instruction and not executing the first audio track switching instruction. The second playback history information indicates that, given the user behavior and environmental state corresponding to the currently collected multimodal data, the user tends to switch to the audio track corresponding to the second audio track switching instruction (the previously predicted output audio track B).

[0177] In other words, if the switching actuator is currently executing audio track switching instruction B and receives a new audio track switching instruction A, it compares the priorities of audio track switching instructions A and B. If the priority of audio track switching instruction B is higher than that of audio track switching instruction A (i.e., the currently executing instruction takes precedence over the instruction to be executed), the switching actuator continues to execute audio track switching instruction B and does not execute audio track switching instruction A. This prioritizes the playback of audio stream B corresponding to audio track B indicated by audio track switching instruction B, thus not interrupting the playback of audio stream B. If the priority of audio track switching instruction A is higher than that of audio track switching instruction B (i.e., the instruction to be executed takes precedence over the currently executing instruction), then the execution of audio track switching instruction B is stopped, and audio track switching instruction A is executed. This interrupts the playback of audio stream B and switches to audio stream A corresponding to audio track A indicated by audio track switching instruction A. When the switching actuator controls the response and execution of instructions based on instruction priority, it synchronously writes the execution status into the context memory pool in the form of playback history information, thereby updating the context memory pool. This allows the context memory pool to continuously accumulate and correct prior knowledge based on the usage process, thereby improving the accuracy of target audio track matching and making the target audio track more in line with the user's habits and preferences.

[0178] Figure 12 This is a timing interaction diagram of the audio track switching method provided in some embodiments of this application.

[0179] In some embodiments, see Figure 12 The interaction objects involved in this time-series interaction process include the detection device, the context prediction engine, the double buffer engine, the switching actuator, and the audio buffer pool.

[0180] In some embodiments, see Figure 12 The detection device collects multimodal data and sends it to the context prediction engine. The context prediction engine can then invoke the context prediction model to process the multimodal data (see [reference]). Figures 6-8 It predicts and outputs the target audio track, and sends the target audio track's identification information (such as the track ID) to the double buffer engine.

[0181] In some embodiments, see Figure 12 The double buffering engine sends a preload request to the audio buffer pool, which contains the target audio track information. The audio buffer pool can be configured on the server side. In response to the preload request, the server sends a second audio stream that matches the target audio track information to the double buffering engine, which then caches the second audio stream locally.

[0182] In some embodiments, see Figure 12 The double-buffering engine can analyze the first audio stream and obtain a set of candidate points (see...). Figure 9), and select the target candidate point from the candidate point set (refer to Figure 10 The candidate point of the target is used as the timing point for switching to the target audio track.

[0183] In some embodiments, see Figure 12 The dual-buffering engine can perform phase alignment processing on the first audio track and the target audio track, and generate an audio track switching instruction A based on the target audio track information, target candidate points, and instruction priority. This instruction A is then sent to the switching executor. The audio track switching instruction A can also include information such as user behavior, trajectory, and environmental status recognized by multimodal data recognition, so that the switching executor can subsequently feed back the corresponding playback history information to the context memory pool after executing the audio track switching instruction A.

[0184] In some embodiments, see Figure 12 When the switching actuator receives the track switching instruction A, it can identify the instruction priority of the track switching instruction A. If the track switching instruction A has a high priority, it will execute the track switching instruction A and stop executing the previously executed track switching instruction B.

[0185] In some embodiments, see Figure 12 When the switching actuator executes the audio track switching instruction A, it can request the second audio stream from the double buffer engine. The double buffer engine then transmits the second audio stream to the switching actuator. The switching actuator then transitions from the first audio track to the target audio track from the target candidate point using a preset transition method (such as fade-in, fade-out, crossfade, etc.) and plays the second audio stream corresponding to the target audio track.

[0186] In some embodiments, see Figure 12 The switching actuator can feed back the first playback history information to the context prediction engine. This first playback history information is used to indicate the target audio track information mapped by the user behavior, trajectory and environmental status information identified by the currently collected multimodal data.

[0187] In some embodiments, see Figure 12 When the context prediction engine receives the first playback history information, it can write the first playback history information into the context memory pool. The context memory pool continuously accumulates and corrects prior knowledge based on the usage process, thereby improving the accuracy of target audio track matching and making the target audio track more in line with the user's habits and preferences.

[0188] This application utilizes variables such as different user behaviors, movement trajectories, and environmental states to adaptively predict potential target audio tracks and match appropriate audio track switching times, thus providing the switching executor with precise audio track switching criteria. If the target audio track predicted based on dynamically collected real-time multimodal data changes, or the audio track switching point changes, a new audio track switching command can be issued. When the switching executor's task queue contains multiple audio track switching commands, the highest priority command is executed to ensure the audio track switching matches user preferences. The timing interaction logic of the audio track switching method is not limited to... Figure 12 Examples of this technology can be adaptively adjusted and expanded based on different conditions such as dynamically acquired multimodal data, predicted target audio track categories, and instruction priorities. Furthermore, interactive objects can be configured (e.g., deleted, added, or replaced) based on different electronic devices to adapt to different audio playback scenarios.

[0189] In some embodiments, a computer storage medium is also provided, which may store a program. When the computer storage medium is configured in an electronic device, the program, when executed, may include the program steps involved in the audio track switching method in the above embodiments. The computer storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0191] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the foregoing exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be made based on the foregoing teachings. The selection and description of the above embodiments are for the purpose of better explaining the contents of this disclosure, thereby enabling those skilled in the art to better utilize the described embodiments.

Claims

1. A method for switching audio tracks, characterized in that, include: When playing the first audio stream corresponding to the first audio track, multimodal data is collected by a detection device. The multimodal data includes visual modal data, auditory modal data and motion modal data. The multimodal data is normalized using a convolutional neural network model to obtain multimodal feature vectors; The multimodal feature vector is input into the context prediction model, which then predicts the target audio track that matches the multimodal data. Detect the silence segment, zero-crossing point, energy stability segment, and harmonic stability segment of the first audio stream; Obtain a candidate point set, the candidate point set including at least one of a first candidate point, a second candidate point, a third candidate point, and a fourth candidate point; wherein, the first candidate point is selected from the silent segment, the second candidate point is selected from the zero-crossing rate point, the third candidate point is selected from the energy stable segment, and the fourth candidate point is selected from the harmonic stable segment; Select a target candidate point from the candidate point set, wherein the target candidate point is the switching point from the first audio track to the target audio track; The first audio track and the target audio track are phase aligned, and the target audio track is switched at the target candidate point to play the second audio stream corresponding to the target audio track.

2. The method according to claim 1, characterized in that, The step of obtaining the candidate point set includes: Calculate the zero-crossing rate, energy fluctuation range, and fundamental frequency change rate of the audio frames in the first audio stream; When the length of the silent segment is greater than the first threshold, the midpoint of the silent segment is taken as the first candidate point; And / or, when the zero-crossing rate of a consecutive preset number of audio frames is less than a second threshold, the valley point among the zero-crossing rate points is used as the second candidate point; And / or, when the energy fluctuation range is within the threshold range, the starting point of the energy steady segment is taken as the third candidate point; And / or, when the fundamental frequency change rate is less than the third threshold, the midpoint of the harmonic stability segment is taken as the fourth candidate point.

3. The method according to claim 2, characterized in that, The step of selecting a target candidate point from the candidate point set includes: Audio features of preset lengths before and after the candidate points are extracted sequentially; wherein, the candidate points are the candidate points included in the candidate point set. The encoder of the converter model processes the audio features of the candidate points at preset lengths before and after the candidate points to calculate the quality score of the candidate points. Unqualified candidate points are filtered out from the candidate point set, wherein the quality score of the unqualified candidate points is less than the fourth threshold. The target candidate points are obtained by merging adjacent candidate points in the filtered candidate point set whose spacing is less than the fifth threshold.

4. The method according to claim 1, characterized in that, The step of inputting the multimodal feature vector into a context prediction model, and having the context prediction model predict a target audio track matching the multimodal data, includes: The context prediction model, based on the multimodal feature vector and the context memory pool, performs joint modeling from spatial and temporal dimensions to obtain spatiotemporal joint features; wherein, the context memory pool is used to store playback history information, which includes audio track information related to different user habits and environmental states; The context prediction model uses a fully connected layer and a normalized exponential function to predict user behavior based on the spatiotemporal joint features, thereby obtaining the probability distribution of user behavior categories. The context prediction model uses a long short-term memory network decoder to predict user trajectory sequences within a preset time period in the future; The context prediction model determines the first candidate audio track based on the probability distribution of the user behavior category and the user trajectory sequence; The context prediction model outputs the confidence level of the first candidate audio track based on a regression network; Obtain a second candidate audio track, wherein the confidence level of the second candidate audio track is greater than the confidence level threshold; The second candidate audio track with the highest confidence level is determined as the target audio track.

5. The method according to claim 4, characterized in that, The step of obtaining spatiotemporal joint features by jointly modeling from spatial and temporal dimensions based on the multimodal feature vector and the context memory pool using the context prediction model includes: The context prediction model retrieves target playback history information associated with the multimodal data from the context memory pool, and generates a context feature vector based on the target playback history information; The context prediction model concatenates the multimodal feature vector and the context feature vector to obtain a concatenated vector; The context prediction model obtains the spatial and temporal features in the concatenated vector; The context prediction model performs spatial graph convolution operations on the spatial features and the graph convolutional neural network to obtain spatial modal features; The context prediction model performs a time-series transformation operation on the time features and the recurrent neural network to obtain time modal features; The context prediction model fuses the spatial modal features and the temporal modal features to obtain the spatiotemporal joint features.

6. The method according to claim 1, characterized in that, After inputting the multimodal feature vector into a context prediction model, and having the context prediction model predict a target audio track matching the multimodal data, the method further includes: The second audio stream corresponding to the target audio track is preloaded from the audio buffer pool, which is used to cache audio streams corresponding to different audio tracks.

7. The method according to claim 4, characterized in that, The step of switching to the target audio track at the target candidate point and playing the second audio stream corresponding to the target audio track includes: Based on the target audio track and the target candidate point, a first audio track switching instruction to be executed is generated; Send the first audio track switching command to the switching executor; In response to the first audio track switching command, the switching actuator transitions from the first audio track to the target audio track at the target candidate point using a preset transition method, so as to play the second audio stream corresponding to the target audio track.

8. The method according to claim 7, characterized in that, Before sending the first audio track switching command to the switching actuator, the method further includes: If the second audio track switching command currently being executed by the switching executor is detected, the first command priority corresponding to the first audio track switching command set by the user and the second command priority corresponding to the second audio track switching command are obtained; Write the first instruction priority into the first audio track switching instruction, and mark the second audio track switching instruction as corresponding to the second instruction priority.

9. The method according to claim 8, characterized in that, The step of having the switching actuator respond to the first audio track switching command, at the target candidate point, transition from the first audio track to the target audio track using a preset transition method, and play the second audio stream corresponding to the target audio track, includes: The switching executor obtains the priority of the first instruction from the first audio track switching instruction to be executed; The switching executor obtains the priority of the second instruction corresponding to the currently executing second track switching instruction; If the priority of the first instruction is higher than the priority of the second instruction, the switching executor stops executing the second audio track switching instruction and executes the first audio track switching instruction, writing the first playback history information to the context memory pool; wherein, the first playback history information is used to indicate that the user tends to switch to the target audio track under the user behavior and environmental state corresponding to the multimodal data; If the priority of the first instruction is lower than the priority of the second instruction, the switching executor continues to execute the second audio track switching instruction, does not execute the first audio track switching instruction, and writes the second playback history information to the context memory pool; wherein, the second playback history information is used to indicate that the user tends to switch to the audio track corresponding to the second audio track switching instruction under the user behavior and environmental state corresponding to the multimodal data.

10. An electronic device, characterized in that, include: The detection device is configured to acquire multimodal data, including visual modal data, auditory modal data, and motion modal data; The audio output device is configured to output the audio stream corresponding to the audio track. The controller, coupled to the detection device and the audio output device, is configured as follows: When the audio output device outputs the first audio stream corresponding to the first audio track, the multimodal data collected by the detection device is acquired; The multimodal data is normalized using a convolutional neural network model to obtain multimodal feature vectors; The multimodal feature vector is input into the context prediction model, which then predicts the target audio track that matches the multimodal data. Detect the silence segment, zero-crossing point, energy stability segment, and harmonic stability segment of the first audio stream; Obtain a candidate point set, the candidate point set including at least one of a first candidate point, a second candidate point, a third candidate point, and a fourth candidate point; wherein, the first candidate point is selected from the silent segment, the second candidate point is selected from the zero-crossing rate point, the third candidate point is selected from the energy stable segment, and the fourth candidate point is selected from the harmonic stable segment; Select a target candidate point from the candidate point set, wherein the target candidate point is the switching point from the first audio track to the target audio track; The first audio track and the target audio track are phase aligned, and the target audio track is switched at the target candidate point so that the audio output device outputs the second audio stream corresponding to the target audio track.