Audio processing method and device, electronic equipment, storage medium and program product

By acquiring user task information and audio file characteristics, the system automatically filters audio file combinations that match the task type and duration, solving the tedious problem of users manually selecting audio files and enabling users to have an immersive listening experience during tasks.

CN121884751APending Publication Date: 2026-04-17VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-01-19
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, users need to manually select the next audio file to play after the podcast audio has finished playing, which makes the operation cumbersome and difficult to meet the complex needs of a task-oriented listening experience.

Method used

By acquiring user task information, including task type and duration, features are extracted from the audio file list to filter out combinations of audio files that match the task type and duration, enabling automatic combination playback.

Benefits of technology

Users can immerse themselves in the task without having to manually select audio files, which enhances the flexibility of audio processing on electronic devices and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121884751A_ABST
    Figure CN121884751A_ABST
Patent Text Reader

Abstract

The invention discloses an audio processing method and device, electronic equipment, a storage medium and a program product, and belongs to the technical field of computers, and the method comprises the steps: obtaining user task information of a first user task; the user task information comprises a task type and a task duration; performing feature extraction processing on each audio file in the audio file list to obtain audio file features corresponding to each audio file; the audio file features comprise duration dimension features, content dimension features and interest dimension features; determining a first audio file combination according to the user task information and the audio file features corresponding to the audio files; the absolute value of the difference value between the audio duration of the first audio file combination and the task duration of the first user task is smaller than or equal to a first preset threshold value, and the audio content of the first audio file combination is matched with the task type of the first user task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer technology, and specifically relates to an audio processing method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] With the increasing richness of the podcast content ecosystem and the widespread adoption of smart devices, users can listen to podcast audio content anytime, anywhere. Typically, electronic devices can recommend and play podcast audio content for users based on collaborative filtering, content tag matching, or semantic retrieval mechanisms, following a "content-first" or "interest-first" logic.

[0003] However, in the above method, when electronic devices recommend podcasts based on the user's potential interests, the user needs to manually select the next podcast to play after the audio of a podcast has finished playing, which makes the user's operation rather cumbersome. Summary of the Invention

[0004] The purpose of this application is to provide an audio processing method, apparatus, electronic device, storage medium, and program product that can avoid the tedious operation of users manually selecting audio for playback.

[0005] In a first aspect, embodiments of this application provide an audio processing method, which includes: obtaining user task information of a first user task; the user task information including task type and task duration; performing feature extraction processing on each audio file in an audio file list to obtain audio file features corresponding to each audio file; the audio file features including duration dimension features, content dimension features, and interest dimension features; determining a first audio file combination based on the user task information and the audio file features corresponding to each audio file; the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task.

[0006] Secondly, embodiments of this application provide an audio processing apparatus, comprising: an acquisition module, a processing module, and a determination module; the acquisition module is used to acquire user task information of a first user task; the user task information includes task type and task duration; the processing module is used to perform feature extraction processing on each audio file in the audio file list to obtain audio file features corresponding to each audio file; the audio file features include duration dimension features, content dimension features, and preference dimension features; the determination module is used to determine a first audio file combination based on the user task information acquired by the acquisition module and the audio file features corresponding to each audio file obtained by the processing module; the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0011] In this embodiment, user task information for a first user task is obtained; this user task information includes task type and task duration; feature extraction processing is performed on each audio file in the audio file list to obtain audio file features corresponding to each audio file; these audio file features include duration dimension features, content dimension features, and preference dimension features; based on the user task information and the audio file features corresponding to each audio file, a first audio file combination is determined; the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task. In this solution, the electronic device can automatically combine audio files that match the task type and task duration from the audio file list according to the task type and task duration corresponding to the user task. That is, the combined audio files can match the user's listening preferences during the execution of a certain task, thus satisfying the user's actual listening needs for audio files during the execution of that task, allowing the user to immerse themselves in the task, thereby avoiding the tedious operation of manually selecting audio for playback. Attached Figure Description

[0012] Figure 1 This is one of the flowcharts of an audio processing method provided in the embodiments of this application;

[0013] Figure 2 This is a flowchart illustrating a method for determining task information provided in an embodiment of this application;

[0014] Figure 3 This is a second flowchart of an audio processing method provided in an embodiment of this application;

[0015] Figure 4 This is a flowchart of a method for parsing audio files provided in an embodiment of this application;

[0016] Figure 5 This is a flowchart of a method for determining an audio set provided in an embodiment of this application;

[0017] Figure 6 This is the third flowchart of an audio processing method provided in the embodiments of this application;

[0018] Figure 7 This is a flowchart of a method for merging audio files provided in an embodiment of this application;

[0019] Figure 8 This is a flowchart of a method for adjusting an audio segment provided in an embodiment of this application;

[0020] Figure 9 This is a schematic diagram of the structure of an audio processing device provided in an embodiment of this application;

[0021] Figure 10 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;

[0022] Figure 11 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects. For example, a first object can be one or more, where "more" means at least two. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0025] The terms "at least one," "at least one," etc., used in this application's specification refer to any one, any two, or a combination of two or more of the included objects. For example, "at least one of a, b, and c" can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple, and multiple means at least two. Similarly, "at least two" refers to two or more, and its meaning is similar to that of "at least one."

[0026] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.

[0027] A podcast is a series of digital audio files that combines the audio content of a radio broadcast with the subscription and distribution mechanism of a blog. Users can access it through subscription, download, or online playback, allowing them to listen to or watch non-live broadcast programs anytime, anywhere via portable electronic devices such as mobile phones and smart speakers.

[0028] Dynamic programming is an optimization algorithm that decomposes a complex problem into overlapping subproblems and uses the solutions to the subproblems to solve the original problem. It is widely used in combinatorial optimization, sequence decision-making, and resource allocation, such as the knapsack problem, shortest path problem, and longest common subsequence problem.

[0029] Semantic similarity is a core task in Natural Language Processing (NLP), aiming to measure the semantic similarity between two texts or semantic vectors. Unlike traditional methods based on superficial forms such as word overlap, semantic similarity can achieve more accurate matching by capturing the deeper meaning and contextual relationships of the text.

[0030] Natural Language Processing (NLP) is an artificial intelligence technology that enables computers to understand, interpret, and generate human language, achieving natural human-computer interaction. It integrates knowledge from multiple disciplines such as linguistics, computer science, mathematics, and statistics, and is widely used in translation, search, dialogue systems, text analysis, and other scenarios.

[0031] With the increasing richness of the podcast content ecosystem and the widespread adoption of smart terminals, users' consumption of audio content is gradually shifting from "passive reception" to "scene adaptation and personalized control".

[0032] Typically, when users listen to podcasts, electronic devices can recommend and play podcast content based on collaborative filtering, content tag matching, or semantic retrieval mechanisms, following a "content-first" or "interest-first" logic. However, when users have specific time windows and content preferences, such as during commutes, workouts, or household chores, the audio recommendation and playback logic in these technologies, while providing some personalization, often ignores the constraints of the user's task time, making it difficult to meet the complex needs of a "task-oriented listening experience." Consequently, after one audio clip finishes playing, users need to manually select the next podcast, leading to cumbersome operations and poor flexibility in audio processing by electronic devices.

[0033] The audio processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0034] The audio processing method provided in this application can be applied to scenarios where users need to listen to audio files. For example, while a user is driving, exercising, or doing housework, the user can trigger an electronic device to play an audio file.

[0035] The audio processing method provided in this application embodiment will be illustrated below with examples of specific scenarios.

[0036] Scenario 1: An electronic device can obtain task information corresponding to the user's task, such as the user being in a driving scenario and needing to drive for 2 hours. The electronic device can obtain a list of audio files, such as all podcast programs included in a podcast channel followed by the user, and perform feature extraction processing on each audio file in the list to obtain features corresponding to each audio file in dimensions such as duration, content, and interests. Then, the electronic device can filter the audio file list according to the driving scenario, such as filtering out audio files related to driving safety, and determine a combination of audio files with a duration of approximately 2 hours based on the filtered audio files according to the driving duration. Thus, the electronic device can continuously play audio files for the user during the 2 hours of driving.

[0037] Scenario 2: The electronic device can obtain task information corresponding to the user's task, such as the user being in a sports activity and needing to exercise for 1 hour. The electronic device can obtain a list of audio files, such as the user's favorite podcasts, and perform feature extraction processing on each audio file in the list to obtain features such as duration, content, and interests for each audio file. Then, the electronic device can filter the audio file list according to the sports activity, such as filtering out rhythmic music recommendations, and determine a combination of audio files with a duration of approximately 1 hour based on the filtered audio files and the duration of the exercise. Thus, the electronic device can continuously play audio files for the user during the 1 hour of exercise.

[0038] It should be noted that the above scenarios 1 and 2 are merely exemplary examples of some scenarios that may be applied to the embodiments of this application. In actual implementation, the embodiments of this application can also be applied to more possible scenarios, such as providing audio or video files according to the user's needs in terms of time, theme, etc. The embodiments of this application are not limited here.

[0039] Based on the above-mentioned scenario applied in the embodiments of this application, the audio processing method provided in the embodiments of this application obtains user task information of a first user task; the user task information includes task type and task duration; performs feature extraction processing on each audio file in the audio file list to obtain audio file features corresponding to each audio file; the audio file features include duration dimension features, content dimension features, and preference dimension features; determines a first audio file combination based on the user task information and the audio file features corresponding to each audio file; the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task. In this solution, the electronic device can automatically combine audio files from the audio file list that match the task type and duration based on the user's task. In other words, the combined audio files can match the user's listening preferences during the execution of a certain task, thus satisfying the user's actual listening needs for audio files during the execution of that task. This allows the user to immerse themselves in the task, avoiding the tedious operation of manually selecting audio to play during task execution and improving the flexibility of audio processing of the electronic device.

[0040] The audio processing method provided in this application is executed by an audio processing device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this method. The following will use an electronic device as an example to illustrate the audio processing method provided in this application.

[0041] This application provides an audio processing method. Figure 1 A flowchart of an audio processing method provided in an embodiment of this application is shown. Figure 1 As shown, the audio processing method provided in this application embodiment may include the following steps 201 to 203.

[0042] Step 201: The electronic device obtains the user task information of the first user task.

[0043] In some embodiments of this application, the first user task mentioned above may be a task that the user is currently performing or that the user is preparing to perform.

[0044] In some embodiments of this application, the aforementioned user task information may include task type and task duration.

[0045] In some embodiments of this application, the above-mentioned task types may include, but are not limited to, at least one of the following: driving type, sports type, housework type, work type, and sleep type.

[0046] In some embodiments of this application, the first task type described above may include multiple levels of types.

[0047] In some embodiments of this application, the above-mentioned task type can be a refined task type obtained by combining the time of task execution, that is, the second-level task type included in the first-level task type.

[0048] For example, if the first-level task type is driving, the corresponding second-level task types can include, but are not limited to, weekend driving, weekday commuting driving, and weekday commuting driving. Thus, the electronic device can subsequently provide different themed audio files based on the user's driving needs at different times.

[0049] For example, when the first-level task type is exercise, the corresponding second-level task type can include, but is not limited to, morning exercise, midday exercise, evening exercise, and nighttime exercise. Thus, the electronic device can subsequently provide the user with audio files of different rhythms based on the user's exercise needs at different times of day.

[0050] Optionally, the matching relationship between task type and audio content can be set by the user or determined based on the user's historical listening habits. For example, users may prefer podcasts related to their personal hobbies while driving on weekends, and podcasts related to their work while driving from Monday to Friday. Or, users may prefer industry news during midday from Monday to Friday, and in-depth news during evenings from Monday to Friday.

[0051] In some embodiments of this application, the electronic device can determine the task type of the first user task based on the user's focus mode settings.

[0052] It should be noted that the above focus model can be the focus mode set by default on electronic devices, or it can be a focus mode actively created by the user.

[0053] For example, the aforementioned focus mode may include, but is not limited to, at least one of the following: driving mode, sports mode, housework mode, work mode, and sleep mode.

[0054] In some embodiments of this application, the task type and task duration included in the above-mentioned user task information can be manually set by the user or automatically determined by the electronic device based on the acquired information.

[0055] In some embodiments of the application, the task duration can be a duration set by the user or a duration automatically predicted and generated by the electronic device based on historical behavior data.

[0056] In some embodiments of this application, a user can actively trigger an electronic device to start a certain focus mode by inputting information. Then, the electronic device can determine the task type of the first user task based on the focus mode. The electronic device can then display a prompt window to prompt the user to set the focus duration corresponding to the focus mode, that is, the task duration.

[0057] In some embodiments of this application, users can set the focus mode of an electronic device so that the electronic device can automatically start or exit the focus mode at a specified time. Thus, the duration between the start time and the exit time set by the user is the focus duration corresponding to the focus mode, which is also the task duration mentioned above.

[0058] In some embodiments of this application, a user can trigger the electronic device to start a certain focus mode by setting the mode of the electronic device; or, a user can trigger the electronic device to run a certain special application, so that the electronic device will automatically start the focus mode corresponding to the special application while running the special application.

[0059] For example, when the electronic device detects that the user has triggered the electronic device to run a navigation application, the electronic device can automatically activate driving mode; when the electronic device detects that the user has triggered the electronic device to run a sports application, the electronic device can automatically activate sports mode.

[0060] It should be noted that users can pre-configure the association between the aforementioned special applications and their corresponding focus modes.

[0061] In some embodiments of this application, step 201 can be specifically implemented by the following steps 201a and 201b.

[0062] Step 201a: The electronic device determines the task type based on the first information.

[0063] In some embodiments of this application, the first information mentioned above may include, but is not limited to, at least one of the following: user activity information, time information, user status information, and environmental information.

[0064] In some embodiments of this application, the aforementioned user activity information may include, but is not limited to, at least one of the following: walking, running, swimming, cycling, weightlifting, and remaining stationary.

[0065] In some embodiments of this application, the aforementioned user status information may include, but is not limited to, at least one of the following: heart rate, blood oxygen saturation, body temperature, and emotional state.

[0066] In some embodiments of this application, the electronic device can obtain user status information through sensors such as gyroscopes and accelerometers installed inside the electronic device, or other connected electronic devices, and infer the aforementioned user activity information based on the user status information.

[0067] For example, when the electronic device is a mobile phone, the mobile phone can obtain user status information such as heart rate, blood oxygen saturation, and body temperature from connected wearable electronic devices such as smartwatches and Bluetooth headsets, and infer the aforementioned activity information based on the user status information.

[0068] In some embodiments of this application, the above-mentioned time information can be used to represent the time information of the current moment.

[0069] In some embodiments of this application, the aforementioned environmental information can be understood as information about the environment in which the electronic device is located, obtained by the electronic device through sensors.

[0070] For example, an electronic device can obtain sounds from the environment through a microphone. For instance, if the microphone detects the sounds of appliances such as a vacuum cleaner or washing machine at a certain moment, the electronic device can determine that the first task type is a housework type, meaning that the user is currently doing housework.

[0071] It is understood that each type of task can possess a distinct behavioral pattern and sound reception characteristics, thereby allowing the electronic device to obtain the behavioral pattern and sound reception characteristics based on the aforementioned first information, and thus determine the type of the first task. The behavioral pattern may include at least one of the user activity information and user state information from the aforementioned first information, and the sound reception characteristics may be the environmental information from the aforementioned first information.

[0072] In some embodiments of this application, the first information may further include the operating status of the electronic device.

[0073] In some embodiments of this application, the operating state of the electronic device may include, but is not limited to, at least one of the following: the focus mode of the electronic device, and the application running on the electronic device.

[0074] For example, when the electronic device is in driving mode, the electronic device can automatically determine the first task type as driving based on the first information; when the electronic device is in sport mode, the electronic device can automatically determine the first task type as sport based on the first information.

[0075] For example, when the electronic device is running a navigation application, the electronic device can automatically determine the first task type as driving based on the first information; when the electronic device is running a sports application, the electronic device can automatically determine the first task type as sports based on the first information.

[0076] Step 201b: The electronic device determines the task duration based on the task duration of historical user tasks corresponding to the task type.

[0077] In some embodiments of this application, the above-mentioned task information may also include other feasible task-related information, such as the listening duration and completion rate of audio files by the user during the historical user task.

[0078] It should be noted that the above-mentioned listening completion rate can be understood as: the proportion of the time a user completes listening to an audio file out of the total duration of the audio file.

[0079] For example, suppose the duration of audio file 1 is 40 minutes, and the user stops playing audio file 1 by inputting input after listening for 20 minutes, that is, the user has listened for 20 minutes. Therefore, the listening completion rate of audio file 1 can be 20 / 40=0.5, or it can be recorded as 50%.

[0080] In some embodiments of this application, after the electronic device determines the task type of the first user task, since the user has not manually set the task duration, the electronic device can obtain or call the task information of historical user tasks stored in the electronic device, so as to determine the task duration of the first user task based on the task information of the historical user tasks.

[0081] In some embodiments of this application, when an electronic device stores task information of a historical user task corresponding to a task type, the electronic device can directly determine the task duration based on the task duration of that historical task.

[0082] In some embodiments of this application, step 201b can be specifically implemented by the following steps 201b1 and 201b2.

[0083] Step 201b1: The electronic device obtains the average task duration, standard deviation of task duration, and audio listening completion time of historical user tasks.

[0084] In some embodiments of this application, the electronic device can obtain the task duration of each of at least two historical user tasks, so as to calculate the above-mentioned average task duration based on the durations of at least two tasks. Standard deviation of task duration .

[0085] In some embodiments of this application, the above-mentioned audio listening completion time is... It can be used to calculate the average listening time of audio files in historical user tasks corresponding to the task type.

[0086] In some embodiments of this application, the above-mentioned audio listening completion time is... This can be calculated based on the total duration of the audio file and the historical completion rate. For example, the audio listening completion time. =Total duration of audio files × Historical listening completion rate.

[0087] Step 201b2: The electronic device performs a weighted calculation on the average task duration, the standard deviation of the task duration, and the audio listening completion time to determine the task duration.

[0088] In some embodiments of this application, the electronic device can estimate the duration of the above-mentioned task based on the following formula (1).

[0089] Formula (1)

[0090] in, It can be used to predict the duration of this task, i.e., the task duration mentioned above; It can be used to represent the average task duration of a user's historical tasks under the same task type as the task type, i.e., the average task duration mentioned above. It can be used to represent the average listening completion time when a user's historical execution of the same task type as the first user's task, i.e., the audio listening completion time mentioned above. This can be used to represent the standard deviation of the duration of the above tasks; , , It can be a dynamically adjustable personalized weighting coefficient, and .

[0091] It is understandable that the prediction function corresponding to the above formula (1) can integrate user behavior stability, interest stickiness, and task duration volatility, thereby allowing for the adjustment of personalized weight parameters. , , Fine-tuning the values, i.e., fine-tuning the behavior patterns, is to optimize the overall prediction accuracy of the prediction function for the duration of the aforementioned tasks.

[0092] It should be noted that the above , , This can be a dynamically adjusted summary of user behavior analysis reports obtained from electronic devices, aggregated across task types. Specifically, this user behavior analysis report can be generated after an audio file playback has finished.

[0093] In this embodiment, the electronic device can estimate the duration of the user's current first user task by integrating data from multiple dimensions, including the average task duration, average listening completion time, and duration standard deviation under the same user task type, as well as the personalized weight coefficients corresponding to each dimension. In other words, the electronic device can perform quantitative analysis based on the user's past playback of audio files in the same task scenario to reduce the user's manual input burden, thereby improving the flexibility and ease of use of the electronic device in obtaining task information.

[0094] In this embodiment, the electronic device can perform quantitative analysis based on the user's past playback of audio files in the same task scenario, reducing the user's manual input burden. That is, even if the user does not manually input task information, the electronic device can still determine the user task information of the first user task based on the collected information related to at least one of the user, the electronic device, and the environment. This improves the flexibility of the electronic device in obtaining task information. Furthermore, the electronic device can subsequently provide the user with a combination of audio files that match the user task information based on the obtained user task information. In this way, the tedious operation of manually selecting audio for playback during task execution is avoided, and the flexibility of audio processing of the electronic device is improved.

[0095] In some embodiments of this application, when the electronic device obtains the task duration based on the user's manual input, the electronic device can calculate the predicted task duration in parallel. To predict the duration of the task The duration manually entered by the user is corrected. That is, Figure 2 The flowchart shown.

[0096] For example, if the user's input of the task duration is abnormal, such as the duration being too long or too short due to an editing error, the electronic device can predict the task duration based on this. This serves as a benchmark for intelligent recommendations or correction suggestions, aiming to reduce or avoid errors in the task duration manually entered by the user.

[0097] In this way, electronic devices can have stronger rollback mechanisms and recovery strategies in the event of fuzzy input or unexpected interruptions, enhancing the system's fault tolerance. This "predictive + manual" dual-mode duration initialization method not only reflects the electronic device's deep understanding of user behavior patterns but also flexibly adapts to the usage habits of different users, allowing users to quickly enter the appropriate listening state without repeatedly adjusting parameters, significantly improving the intelligence level and ease of operation of electronic devices.

[0098] In some embodiments of this application, the task duration manually entered by the user is compared with the predicted task duration mentioned above. If the time difference between the two inputs exceeds a preset time, the electronic device can determine that the task duration manually entered by the user is incorrect. For example, the preset time could be 5 minutes.

[0099] It should be noted that the above-mentioned preset duration can be determined according to actual needs, and this application embodiment does not limit it here.

[0100] In some embodiments of this application, if the task duration manually entered by the user is incorrect, the electronic device can directly predict the task duration. Alternatively, the electronic device can display a prompt message to remind the user to check the entered task duration or to re-enter the task duration.

[0101] Step 202: The electronic device performs feature extraction processing on each audio file in the audio file list to obtain the audio file features corresponding to each audio file.

[0102] In this embodiment of the application, the above-mentioned audio file list may include at least one audio file.

[0103] In some embodiments of this application, the aforementioned audio file list may include, but is not limited to, at least one of the following: a list of audio files collected by the user, or a list of audio files recommended by the electronic device to the user.

[0104] In some embodiments of this application, the aforementioned audio files may include, but are not limited to, any of the following: podcast programs, music files, and audiobooks.

[0105] In some embodiments of this application, when the audio files include podcast programs, the electronic device can obtain the podcast channels that the user follows or favorites through an audio processing application, so as to identify all podcast programs in the podcast channel as audio files in the above-mentioned audio file list.

[0106] In some embodiments of this application, when the audio files include audiobooks, the electronic device can use an audio processing application to obtain works that the user has followed or collected, so as to determine the audio files of each chapter corresponding to the work as audio files in the above-mentioned audio file list.

[0107] In some embodiments of this application, when a user has a need to view videos, the aforementioned audio file list can be replaced with a video file list according to the settings. Then, the electronic device can obtain the video file characteristics corresponding to each video file in the video file list, and determine a video combination in subsequent steps based on the user task information and the video file characteristics corresponding to each video file.

[0108] In some embodiments of this application, the aforementioned audio file features may include duration dimension features, content dimension features, and interest dimension features.

[0109] In some embodiments of this application, the aforementioned duration dimension feature can be understood as the audio duration corresponding to each audio file acquired by the electronic device.

[0110] In some embodiments of this application, the above-mentioned content dimension features can be understood as the topic tags corresponding to each audio file obtained by the electronic device.

[0111] In some embodiments of this application, for a user's historically listened-to audio file in the audio file list, the interest dimension feature of the listened-to audio file can be: a completion score calculated by the electronic device based on the historical average playback completion rate corresponding to the listened-to audio file. In other words, electronic devices can use this completion score to determine how much interest a user has in the audio files they have listened to.

[0112] In some embodiments of this application, the electronic device can construct a feature vector for each audio file based on the audio file features corresponding to each of the aforementioned audio files. For example, the feature vector of the i-th audio file. ,in, It can be used to represent the audio duration corresponding to the i-th audio file, i.e., the duration dimension feature mentioned above; It can be used to represent the completion score of the i-th audio file, i.e., the interest dimension feature mentioned above; It can be used to represent a one-dimensional or multi-dimensional semantic topic tag embedding vector corresponding to the i-th audio file, that is, an audio topic vector, which is the content dimension feature mentioned above.

[0113] It should be noted that the above feature vectors The construction method and the feature vector The methods for obtaining features in each dimension are described in detail in the subsequent steps, and will not be repeated here in the embodiments of this application.

[0114] In some embodiments of this application, for a previously listened audio file that the user has not listened to in the audio file list, the interest dimension feature of the previously listened audio file can be a feature determined by the electronic device based on the user's interest dimension features for audio files of the same type as the previously listened audio file. For example, the electronic device can assign a completion rating based on audio files of the same type as the previously listened audio file. The average completion score was further calculated.

[0115] Step 203: The electronic device determines the first audio file combination based on the user task information and the characteristics of each audio file.

[0116] In some embodiments of this application, the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task may be less than or equal to a first preset threshold.

[0117] For example, assuming the first preset threshold is 2 and the duration of the first user task is 20 minutes, the audio duration of the first audio file combination can be within the range of [18, 22].

[0118] In some embodiments of this application, the first preset threshold can be determined according to actual needs, and this application does not limit it here.

[0119] In some embodiments of this application, the audio content of the first audio file combination described above may be matched with the task type of the first user task.

[0120] It should be noted that, for the specific implementation method of how the electronic device determines the first audio file combination based on user task information and the characteristics of each audio file, please refer to the description of the steps below, which will not be repeated here in the embodiments of this application.

[0121] In some embodiments of this application, when the audio file list is a list of audio files collected by the user, if the sum of the audio durations of at least one audio file included in the audio file list is less than or equal to the task duration in the aforementioned user task information, the electronic device can obtain audio files that the user may be interested in based on the user's preferences or listening habits, so as to generate a combination of audio files with sufficient audio duration.

[0122] In the audio processing method provided in this application embodiment, the electronic device can automatically combine audio files that match the task type and duration from the audio file list according to the task type and duration corresponding to the user task. In other words, the combined audio files can meet the user's listening preferences during the execution of a certain task, thereby satisfying the user's actual listening needs for audio files during the execution of that task. This allows the user to immerse themselves in the task, thus avoiding the tedious operation of manually selecting audio for playback during the execution of a task and improving the flexibility of audio processing of the electronic device.

[0123] In some embodiments of this application, combined with Figure 1 ,like Figure 3 As shown, step 203 can be implemented through steps 203a to 203c.

[0124] Step 203a: The electronic device obtains at least one set of candidate audio files based on the task duration of the first user task and the duration dimension features corresponding to each audio file.

[0125] In some embodiments of this application, in the at least one candidate audio file set, the absolute value of the difference between the total audio duration of each candidate audio file set and the task duration of the first user task can be less than or equal to a first preset threshold.

[0126] In some embodiments of this application, when the electronic device obtains the user task information and audio file list of the first user task, it can input them into the duration matching model. Then, the electronic device can use this duration matching model, with the task duration as the target value, and based on the duration dimension features corresponding to each audio file, filter out a set of candidate audio files whose total duration deviates from the target value by less than or equal to a first preset threshold, i.e., the aforementioned at least one set of candidate audio files.

[0127] In some embodiments of this application, the electronic device can be set with a maximum tolerance deviation. That is, the first preset threshold mentioned above, combined with the task duration. Construct the time window corresponding to the above candidate audio file set. That is, select an audio file whose total duration is within a certain range. The set of audio files within the specified range is used as the candidate audio file set.

[0128] It is understandable that the total audio duration of each candidate audio file set in at least one of the above candidate audio file sets is within a duration range determined based on the task duration of the first user task.

[0129] For example, the above maximum tolerance deviation is It can be 120 seconds, or 2 minutes.

[0130] It should be noted that the above duration matching model can be a model built based on the dynamic programming algorithm.

[0131] In some embodiments of this application, in the at least one set of candidate audio files mentioned above, the audio content of each set of candidate audio files can be matched with the task type of the first user task.

[0132] In some embodiments of this application, before step 203a above, the audio processing method provided in the embodiments of this application may further include the following step A1, and the above step 203a may be specifically implemented by the following step A2.

[0133] Step A1: The electronic device removes noisy audio segments from each audio file to obtain the processed audio files.

[0134] In some embodiments of this application, the aforementioned noisy audio segments may include, but are not limited to, at least one of the following: opening introduction, closing summary, advertisement, and silent segment.

[0135] For example, if the audio file is a podcast, the opening introduction may include, but is not limited to: the host's self-introduction and the announcement of the broadcast date; the closing summary may include, but is not limited to: a preview of the next episode and an introduction calling for subscription or sharing.

[0136] It should be noted that some audio files contain advertisements in the dialogue sections, such as jokes with advertisements in a talk show podcast, or silent segments created by deliberate pauses in the middle of the audio file. Therefore, removing advertisements or silent segments in the middle of the audio file may affect the user's listening experience. To reduce or avoid such impacts on the user's listening experience, electronic devices can remove noisy audio segments contained in at least one of the intro and outro of the audio file.

[0137] In some embodiments of this application, when an electronic device acquires an audio file, the electronic device can mark and delete specified audio segments based on the timeline of the audio file; or it can detect silent segments based on parameters such as audio waveform diagrams or energy thresholds, and identify and delete the silent segments.

[0138] In some embodiments of this application, when an electronic device acquires an audio file, the electronic device can convert the audio content contained in the audio file into text, and perform audio calibration processing on the audio file based on preset keywords, sensitive words or repeated paragraphs, so as to eliminate noisy audio segments.

[0139] Understandably, electronic devices can preprocess audio files to minimize the impact of non-core audio content on the user's listening experience of the core audio content. This core content can also be referred to as the "net content," "essential content," "main body," or other suitable names.

[0140] In some embodiments of this application, the electronic device can control the error of the actual playback duration corresponding to each processed audio file to be within a preset duration range, such as within ±5 seconds, so as to ensure that the accuracy of the alignment stage in the subsequent calculation of the total audio duration of an audio file set is controllable.

[0141] In some embodiments of this application, the electronic device can identify and remove the opening and closing advertisement segments of the audio file through the audio signal feature extraction module, and can make a joint judgment by detecting silent segments, identifying background music and identifying keywords. In other words, in order to maintain the structural stability of the audio file, the audio segments corresponding to each audio file in the first audio file combination finally generated by the electronic device can all start playing from the core part containing the actual content.

[0142] Step A2: The electronic device obtains at least one set of candidate audio files based on the task duration and the duration dimension features corresponding to each processed audio file.

[0143] It is understandable that electronic devices can determine at least one set of candidate audio files based on the core audio content of the audio files. This allows the first audio file combination to be obtained based on the core audio content. In other words, when the user listens to the first audio file combination later, he / she can directly hear the core content of each audio file, thereby reducing or avoiding the influence of noisy audio segments such as advertisements and silent segments contained in the original audio file.

[0144] In this embodiment, the electronic device can eliminate noisy audio segments included in the beginning and end of the audio file through refined audio calibration processing, thereby reducing the impact of these noisy audio segments on the user's listening experience. At the same time, by reducing the redundancy of noisy audio segments, the inconsistency of the user's listening process can be reduced or avoided. In this way, the continuity, compactness and consistency of the user experience during the playback of the audio file are guaranteed, and the operation required for the user to manually skip noisy audio segments such as advertisements during the performance of tasks is avoided, thereby improving the flexibility of the electronic device's audio processing.

[0145] Step 203b: The electronic device determines the first audio file set from at least one candidate audio file set based on the content dimension features and interest dimension features corresponding to each candidate audio file set.

[0146] It is understandable that the electronic device can select the set of audio files that best matches the task type of the first user task from at least one set of candidate audio files, i.e., the first set of audio files, so that the first audio file combination obtained by the electronic device based on the first set of audio files can meet the user's needs in terms of both duration and content theme.

[0147] In some embodiments of this application, step 203b can be specifically implemented by the following steps 203b1 to 203b3.

[0148] Step 203b1: The electronic device calculates the semantic similarity between the content dimension features corresponding to each candidate audio file set and the semantic topic vector corresponding to the task type of the first user task, and obtains the semantic similarity corresponding to each candidate audio file set.

[0149] In some embodiments of this application, the above-mentioned content dimension features can be represented in the form of vectors, such as audio topic vectors.

[0150] It should be noted that each audio file can correspond to an audio topic vector.

[0151] In some embodiments of this application, for a candidate audio file set, the electronic device can calculate the semantic similarity of each audio file in the candidate audio file set, that is, calculate the semantic similarity value between the audio topic vector of each audio file and the semantic topic vector corresponding to the task type of the first user task, and add them together to obtain the semantic similarity of the candidate audio file set.

[0152] In some embodiments of this application, the semantic topic vector corresponding to the above-mentioned task type can be a semantic topic vector generated by the electronic device based on the keywords directly determined by the task type.

[0153] For example, when the task type is driving, the electronic device can obtain relevant keywords based on driving, such as driving safety and driving regulations, and generate corresponding semantic topic vectors. Similarly, when the task type is exercise, the electronic device can obtain relevant keywords based on exercise, such as aerobic exercise, anaerobic exercise, benefits of exercise, and introduction to exercise movements, and generate corresponding semantic topic vectors.

[0154] In some embodiments of this application, the electronic device can obtain a set of topics of audio files that the user has listened to while performing the same type of task in the past, so as to determine the topics that the user may want to listen to while performing the same type of task this time, thereby obtaining the semantic topic vector corresponding to the above task type.

[0155] For example, if the task type is weekday commuting and the user likes to listen to current events news during their commute, the semantic topic vector determined by the electronic device based on the user's preferences when performing the task can include information from current events news.

[0156] Similarly, if the task type is weekend driving and the user likes to listen to music, talk shows, or comedy sketches on weekends, the semantic topic vector determined by the electronic device based on the user's preferences when performing the task can include information about music, talk shows, comedy sketches, etc.

[0157] In some embodiments of this application, for a set of candidate audio files, the electronic device can use the cosine similarity function sim() to calculate the semantic similarity value between the audio topic vector corresponding to each audio file and the semantic topic vector corresponding to the task type of the first user task.

[0158] For example, electronic devices can be Let represent the semantic similarity corresponding to the i-th audio file. It can be used to represent the semantic topic vector of the i-th audio file. It can be used to represent semantic topic vectors corresponding to task types.

[0159] It should be noted that for the specific calculation method of the above sim() function, please refer to the calculation method in the related technology, and the embodiments of this application will not be repeated here.

[0160] In some embodiments of this application, before step 203b1 above, the audio processing method provided in the embodiments of this application may further include the following steps B1 to B3.

[0161] Step B1: The electronic device performs semantic expansion processing on the audio topic tags of each audio file in the audio file list to obtain the audio topic set corresponding to each audio file.

[0162] In some embodiments of this application, the electronic device can first obtain the embedded audio theme tags corresponding to each audio file in the audio processing application through an audio processing application.

[0163] In some embodiments of this application, for each audio file, the electronic device can perform semantic analysis to expand the embedded audio topic tags corresponding to the audio file with keywords to obtain a keyword set, so as to obtain the audio topic set corresponding to each of the above audio files.

[0164] Step B2: The electronic device normalizes the audio theme set of each audio file to obtain the audio theme vector corresponding to each audio file.

[0165] Understandably, for each set of audio topics, the electronic device can normalize the multiple audio topic tags contained in the set to obtain fixed-dimensional audio topic vectors. This allows the electronic device to perform corresponding processing based on the data contained in the vectors. For example, calculating the cosine similarity between audio topic vectors.

[0166] In some embodiments of this application, the electronic device can generate the aforementioned audio topic vectors using a language model.

[0167] For example, electronic devices can generate embedding vectors, i.e., the aforementioned audio topic vectors, through pre-trained language models such as BERT.

[0168] Step B3: The electronic device obtains the content dimension features corresponding to each candidate audio file set based on the audio topic vectors corresponding to each audio file.

[0169] Understandably, for a set of candidate audio files, an electronic device can concatenate or combine the audio topic vectors of each audio file included in the set of candidate audio files to obtain a feature vector matrix, that is, to obtain the audio topic vector of the set of candidate audio files. Similarly, the electronic device can perform similar operations on each set of candidate audio files in at least one of the above-mentioned sets of candidate audio files to obtain the content dimension features corresponding to each set of candidate audio files.

[0170] In this embodiment, the electronic device can perform structured parsing of audio files to extract the core features of each audio file and construct a unified multidimensional feature vector. This provides a complete input vector for subsequent duration matching and topic relevance analysis, facilitating data processing by the electronic device.

[0171] Step 203b2: The electronic device calculates the topic adaptation score for each candidate audio file set based on the semantic similarity and interest dimension features corresponding to each candidate audio file set.

[0172] In some embodiments of this application, the completion score described above can be used to represent the historical average playback completion rate of an audio file.

[0173] In some embodiments of this application, the electronic device can use the following formula (2) to combine the user's past listening behavior data for the i-th audio file, i.e., the aforementioned interest dimension features, and use the average completion rate of the i-th audio file as a preference weight factor to calculate the completion score of the i-th audio file. That is, quantification is performed through a scoring function.

[0174] Formula (2)

[0175] in, It can be used to represent the weighted completion score of the i-th audio file, and ; This can represent the actual playback duration corresponding to the user's j-th listening to the i-th audio file; N can be used to represent the theoretical total duration of the i-th audio file, that is, the original audio duration including noisy audio segments; N can be used to represent the number of times the user has listened to the i-th audio file in history; It can be used to represent a completion sensitivity factor and to determine the amplified weight of high completion behaviors in the score.

[0176] It is understandable that the function of the above formula (2) not only reflects the user's continued interest in the i-th audio file, but also improves the priority of high-stickiness audio files in subsequent matching by non-linear weighting of high completion behavior.

[0177] It should be noted that the above-mentioned completion sensitivity factors The value can be a value preset by the electronic device, or a value set by the electronic device based on the user's characteristic behavior. For example, this completion sensitivity factor. The corresponding numerical range can be from 0.5 to 2.0. Among these, the completion sensitivity factor... When the value is greater than 1, the scoring weight of high completion rate behavior can be amplified.

[0178] In some embodiments of this application, for audio files in the above audio file list that the user has not listened to, since the electronic device has not played the audio file, the electronic device cannot calculate the completion score corresponding to the unlistened audio file using the above formula (2). In other words, an unlistened audio file may not have a completion score, or the completion score of an unlistened audio file may be a default value, such as 0.

[0179] In some embodiments of this application, the electronic device can construct a feature vector for each audio file based on the information corresponding to each audio file in the audio file list. This refers to the audio file characteristics of an audio file. For example, the feature vector of the i-th audio file. ,in, It can be used to represent the actual playback duration of the i-th audio file, that is, the audio duration after removing the noise audio segments, which is also the duration dimension feature in the above audio file features; It can be used to represent the completion score of the i-th audio file, which is the interest dimension feature in the above audio file features; This can be used to represent a one-dimensional or multi-dimensional semantic topic tag embedding vector corresponding to the i-th audio file, i.e., an audio topic vector, which is also the content dimension feature in the audio file features mentioned above. For example... Figure 4 The flowchart shown.

[0180] Understandably, electronic devices can transform the audio content of discrete audio files into a structure that can be computed or used by the device through the process of constructing feature vectors or feature matrices. Thus, the electronic device can, based on the feature matrices corresponding to each audio file, combine the task duration set by the user or predicted by the electronic device. This provides basic information for the subsequent steps of combining the first audio file.

[0181] In some embodiments of this application, for a set of candidate audio files, the electronic device can calculate the topic adaptation score corresponding to each audio file in the set of candidate audio files, and obtain the topic adaptation score corresponding to the set of candidate audio files by adding them together.

[0182] In some embodiments of this application, the electronic device can calculate the topic adaptation score corresponding to the i-th audio file using the following formula (3). .

[0183] Formula (3)

[0184] in, It can be used to represent the semantic topic vector of the i-th audio file. It can be used to represent semantic topic vectors corresponding to task types; It can be used to represent the semantic similarity of the i-th audio file; It can be used to represent the historical completion score of the i-th audio file.

[0185] Understandably, electronic devices can reflect the matching priority of an audio file to the current task scenario by multiplying the subject relevance by the user's historical listening completion rate of the audio file.

[0186] Step 203b3: The electronic device determines the first audio file set based on the topic adaptation scores corresponding to each candidate audio file set.

[0187] In some embodiments of this application, the electronic device may determine the first audio file set as the candidate audio file set with the highest topic adaptation score from at least one candidate audio file set.

[0188] In some embodiments of this application, it is assumed that the duration of the first user task is... The above audio file list includes n audio files. The electronic device can determine at least one candidate audio file set based on a state transition function, i.e., the following formula (2). Wherein, the total audio duration t corresponding to each candidate audio file set is in the state transition function. Within the range.

[0189] Formula (4)

[0190] in, It can be used to represent the topic relevance score of a set of candidate audio files with a total audio duration of t, determined by an electronic device based on the first k audio files; It can be used to represent the audio duration of the k-th audio file; It can be used to represent the theme adaptation score corresponding to the kth audio file.

[0191] Understandably, in the process of an electronic device determining whether the k-th audio file can be included in a set of candidate audio files, the electronic device can first select a set of candidate audio files a with a total audio duration of t seconds from the first k-1 audio files, and calculate the topic relevance score of the set of candidate audio files a. Then, the electronic device can base its decision on the audio duration of the k-th audio file. From the first k-1 audio files, select one with a total audio duration of... A set of candidate audio files in seconds, so that the electronic device can select based on the total audio duration of the selected files. Given a set of candidate audio files for a given second and the k-th audio file, determine a candidate audio file set b, and calculate the topic relevance score of this candidate audio file set b, i.e. Therefore, the electronic device can compare the relevance scores of these two topics to determine whether the candidate audio file set b containing the k-th audio file is more suitable for the requirements. That is, as... Figure 5 The flowchart shown illustrates how an electronic device can progressively determine the set of candidate audio files with the highest topic fit scores, thereby identifying the first set of audio files.

[0192] In this embodiment, the electronic device can use task type as an initial constraint, fully considering the behavioral patterns and sound reception characteristics of different tasks to provide a clear direction for content selection. Then, the electronic device can use task duration as the target, filtering a set of candidate audio files from the audio file list whose total duration deviation meets the requirements. Simultaneously, the electronic device can combine the semantic similarity and interest dimension features corresponding to each candidate audio file set to calculate a topic fit score, prioritizing content strongly related to the task scenario based on this topic fit score. Thus, the matching mechanism under multi-dimensional constraints reduces or avoids the tedious manual selection by the user, ensuring that the final combination of audio files highly matches the task requirements in terms of duration and topic. This allows users to listen continuously to content that meets their needs in specific scenarios, reducing interruptions or information redundancy caused by content mismatch, and significantly improving listening efficiency and experience.

[0193] Step 203c: The electronic device merges the audio files in the first audio file set to obtain the first audio file combination.

[0194] In some embodiments of this application, the electronic device may refer to at least one of the following conditions when determining the order of audio files in the first audio file set: collection time, update time, popularity, number of times the user has listened to the audio file, and the name of the audio file.

[0195] In some embodiments of this application, the electronic device can determine the order of the audio files based on the correlation between the audio files in the first audio file set. For details, please refer to steps 301 and 302 below, and the description of their related steps.

[0196] In some embodiments of this application, combined with Figure 3 ,like Figure 6 As shown, before step 203c above, the audio processing method provided in this application embodiment may further include step 301 below, and step 203c above can be specifically implemented by step 302 below.

[0197] Step 301: The electronic device determines the arrangement order of at least two audio files based on the semantic vectors of at least two audio files included in the first audio file set.

[0198] In some embodiments of this application, the electronic device can calculate the similarity between the semantic vectors of each pair of audio files in the at least two audio files to determine the order of the at least two audio files.

[0199] In some embodiments of this application, step 301 can be implemented by steps 301a and 301b.

[0200] Step 301a: The electronic device calculates the semantic similarity between the semantic vectors of the first audio segments of every two audio files in the first audio file set.

[0201] In some embodiments of this application, the first audio segment described above may include a start audio segment and an end audio segment.

[0202] In some embodiments of this application, when an electronic device obtains an audio file, the electronic device can perform speech-to-text processing on the audio file to determine the start and end paragraphs corresponding to the audio file through semantic analysis, thereby obtaining the semantic vectors corresponding to the start and end paragraphs.

[0203] In some embodiments of this application, when an electronic device acquires an audio file, it can determine the start and end audio segments based on parameters such as the audio waveform or energy threshold of the audio file. The electronic device can then perform speech-to-text processing and semantic analysis on the start audio segment to obtain its semantic vector, and perform the same process on the end audio segment to obtain its semantic vector.

[0204] In some embodiments of this application, for audio file a and audio file b in the aforementioned at least two audio files, the electronic device can calculate the similarity between the semantic vectors of the ending audio segment of audio file a and the beginning audio segment of audio file b. Furthermore, the electronic device can calculate the similarity between the semantic vectors of the ending audio segment of audio file b and the beginning audio segment of audio file a. .

[0205] It should be noted that for the specific calculation method of the above sim() function, please refer to the calculation method in the related technology, and the embodiments of this application will not be repeated here.

[0206] Step 301b: The electronic device determines the arrangement order of at least two audio files based on the semantic similarity between the semantic vectors of the first audio segments of every two audio files.

[0207] In some embodiments of this application, to address the semantic breaks and rhythmic imbalances that may exist between at least two audio files, the electronic device can utilize a cross-segment transition processing mechanism, employing the semantic transfer metric function according to the following formula (5). Evaluate the degree of semantic jump between audio file i and audio file j at the junction.

[0208] Formula (5)

[0209] in, It can be used to represent the semantic vector of the end audio segment of audio file i; It can be used to represent the semantic vector of the audio segment starting at the beginning of audio file j; It can be used to represent the cosine similarity between the semantic vectors of the ending audio segment of audio file i and the beginning audio segment of audio file j.

[0210] It is understandable that, if the above The smaller the value, the more natural the semantic transition from audio file i to audio file j. Therefore, electronic devices use this indicator as the basis for adjusting the merging and sorting of segments, prioritizing the logical consistency of adjacent content in the combination to ensure the overall semantic coherence of the merged audio file combination.

[0211] In this embodiment, the electronic device can utilize a semantic transition metric function to evaluate semantic transitions between audio files, adjust the order of transitional segments, and reduce logical breaks. This effectively reduces or avoids semantic gaps that may occur in the final generated and played audio file combination, enabling the electronic device to generate a seamless, customized playback stream. This allows users to immerse themselves in the task, experiencing no noticeable content fragmentation during listening, resulting in a stronger sense of immersion and more efficient information acquisition. This avoids the tedious manual selection of audio for playback during task execution, enhancing the flexibility of the electronic device's audio processing.

[0212] Step 302: The electronic device merges at least two audio files in the order they are arranged to obtain the first audio file combination.

[0213] In some embodiments of this application, the electronic device can merge the audio files in the first audio file set after removing the noise audio segments at the beginning and end of the video, in the order described above, to obtain the first audio file combination.

[0214] In this embodiment, the electronic device can adjust the order of transition segments to generate a seamless, customized playback stream. This allows users to experience no noticeable content breaks during listening, resulting in a more immersive and efficient information acquisition experience. This enhances the flexibility of the electronic device's audio processing.

[0215] In this embodiment, since the electronic device can obtain user task information such as task type and duration, it can filter audio files in the audio file list based on the user task information to obtain at least one set of candidate audio files. Then, the electronic device can determine a first audio file set from the at least one set of candidate audio files according to the task type and duration, and merge them to obtain a first audio file combination. In other words, the first audio file combination conforms to the user's listening habits when performing a task of this type and the duration of the user's current first user task. That is, the first audio file combination can ensure a high degree of fit between the theme and duration and the user's task requirements, allowing the user to immerse themselves in the task. This avoids the tedious operation of manually selecting audio to play during task execution and improves the flexibility of audio processing of the electronic device.

[0216] In some embodiments of this application, after step 203 above, the audio processing method provided in the embodiments of this application may further include the following step 401.

[0217] Step 401: The electronic device adjusts the playback rate of the audio segments in the first audio file combination based on the task duration of the first user task and the completion score of the audio files corresponding to each audio segment in the first audio file combination.

[0218] In some embodiments of this application, the playback rate of the adjusted audio segment can be within the range corresponding to the speech rate carrying capacity.

[0219] Understandably, electronic devices can optimize the playback continuity and compactness of the first audio file combination by compressing the duration, while the dynamic window constraint of the task duration ensures that the total duration still meets the requirements.

[0220] In some embodiments of this application, in order to further compress the overall playback duration and meet the task duration requirements... Dynamic window constraints allow electronic devices to adjust their response based on speech rate and capacity. The playback speed of each audio segment in the first audio file combination. Adjustments are made. Specifically, the electronic device can calculate the playback multiplier corresponding to the i-th audio segment in the first audio file combination using the following formula (6). :

[0221] Formula (6)

[0222] in, It can be used to represent the playback multiplier of the i-th audio segment; It can be used to represent the completion score of the audio file corresponding to the i-th audio segment; It can be used to represent the adjustment coefficient; It can be used to represent the urgency of a task; electronic devices can adjust the task duration accordingly. The error in the total duration of the original audio corresponding to the first audio file combination is dynamically calculated.

[0223] It should be noted that the above-mentioned speech rate capacity This can be understood as the adjustable range of playback speed, such as the playback speed being any value between 0.5x and 3x.

[0224] It is understandable that the above formula (6) can reflect the inverse relationship between playback speed and user interest intensity. That is to say, for audio segments with low completion, electronic devices can speed up their playback under the condition of overall time constraints, while maintaining a stable voice rhythm for high completion content, thereby maximizing the balance between content acceptance and playback efficiency.

[0225] In some embodiments of this application, the audio file corresponding to the i-th audio segment is an audio file that the user has not listened to before, that is, the audio file does not have a corresponding completion score. In this case, electronic devices can Calculate the playback ratio of the unlistened audio file.

[0226] In some embodiments of this application, the electronic device can also adjust the playback rate of audio segments according to the speech rate corresponding to different audio segments, so that the speech rate of each audio segment after adjustment can be roughly the same, so as to balance the information density of the audio file combination and improve the user's listening experience.

[0227] In this embodiment, the electronic device can dynamically adjust the playback rate based on the audio file's completion level and the urgency of the task, appropriately accelerating content with low completion level and maintaining the rhythm of content with high completion level, thus balancing information density and listening comfort. This effectively reduces or eliminates potential rhythmic imbalances in the first audio file combination, generating a seamless, customized playback stream. Users experience no noticeable content breaks during listening, resulting in a stronger sense of immersion and more efficient information acquisition, thereby enhancing the flexibility of the electronic device in audio processing.

[0228] It should be noted that the aforementioned processing of audio file sets and combinations, such as removing noisy audio segments, adjusting the order, and adjusting playback speed, are all done without disrupting the content structure. This allows electronic devices to ultimately output a seamless, customized playback stream with features such as task duration alignment, consistent content style, and coordinated playback rhythm. That is, Figure 7The flowchart shown.

[0229] In some embodiments of this application, after step 203 above, the audio processing method provided in the embodiments of this application may further include the following steps 501 and 502.

[0230] Step 501: While playing the first audio file combination, the electronic device receives a user voice command.

[0231] In some embodiments of this application, the electronic device can receive real-time control commands from the user during the playback of the first audio file combination through the interface of the voice interaction module and through a speech recognition and natural language understanding engine deployed on the device side or in the cloud.

[0232] In some embodiments of this application, the above-mentioned voice commands may include, but are not limited to, at least one of the following: "Skip the current program", "Extend for 10 minutes", "Replace with technology content".

[0233] Step 502: The electronic device updates the first audio file combination according to the user's voice command.

[0234] In some embodiments of this application, an electronic device can convert voice commands into voice command text to determine the user intent information corresponding to the voice command text and execute the operation corresponding to the user intent information.

[0235] In some embodiments of this application, when the electronic device converts voice commands into voice command text, the electronic device can identify the user's intent category using an instruction intent classification model. Examples include operation types such as skipping content, adjusting duration, modifying the theme, and changing the playback order.

[0236] In some embodiments of this application, the electronic device can utilize the dynamic semantic matching function according to the following formula (7). Achieve accurate mapping from natural language to actions of electronic devices.

[0237] Formula (7)

[0238] Where u can be used to represent the above voice command text; I can be used to represent the semantic embedding vector corresponding to the voice command text; I can be used to represent a predefined set of commands that can be executed by an electronic device. It can be used to represent the semantic embedding vector corresponding to instruction I; It can be used to represent the current context suitability score of instruction I; It can be used to represent context-sensitive adjustment coefficients.

[0239] It is understandable that the above formula (7) integrates static semantic similarity and dynamic contextual fit, enabling electronic devices to accurately identify duration parameter adjustment requests such as "extend by 10 minutes" and topic tag reconstruction operations such as "replace with technology content," and can directly affect the target duration in the aforementioned dynamic programming. and topic matching dimension in the feature matrix .

[0240] In some embodiments of this application, all voice control operations can respond within 200 milliseconds, thus ensuring that the user's thought process and operational fluency are not interrupted during the task. Simultaneously, the electronic device can record structured instructions during voice interaction and store them synchronously in a behavior log, providing traceable evidence of user intent for subsequent real-time dynamic adjustment mechanisms.

[0241] In some embodiments of this application, since the electronic device needs to maintain a state index table for the playback stream, when the electronic device receives a "skip current program" instruction, it can quickly retrieve the nearest content without disrupting the playback structure, and then, based on the semantic transfer metric function in step 301b above... Reorder the segments to maintain a natural transition.

[0242] In some embodiments of this application, the above-mentioned user voice command is to switch the currently playing audio segment; the above-mentioned step 502 can be specifically implemented by the following steps 502a to 502c.

[0243] Step 502a: The electronic device stops playing the currently playing audio segment and starts playing the first audio file.

[0244] In some embodiments of this application, the first audio file mentioned above may be the audio file with the highest semantic similarity to the currently playing audio segment among the unplayed audio files included in the audio file list.

[0245] Understandably, since electronic devices store audio topic vectors corresponding to each audio file in the audio processing list, when the electronic device determines that the user's intention is to switch the currently playing audio segment, the electronic device can first determine the audio file with the highest semantic similarity to the currently playing audio segment from the audio files included in the audio processing list, so that the user can continue to listen to the audio content of that topic.

[0246] Step 502b: The electronic device determines the second audio file combination based on the user task information and the characteristics of the audio files corresponding to each unplayed audio file.

[0247] In some embodiments of this application, the sum of the audio duration of the second audio file combination and the audio duration of the first audio file may have an absolute value less than or equal to a first preset threshold.

[0248] In some embodiments of this application, the audio content of the aforementioned second audio file combination may be matched with the task type of the first user task.

[0249] Understandably, since the electronic device has already played a portion of the audio segments from the first audio file combination, it can obtain the remaining task duration of the first user task and the audio duration of the first audio file to determine the remaining audio duration the user needs to listen to in order to complete the first user task. Then, based on this remaining audio duration, the electronic device can re-determine the aforementioned second audio file combination according to the characteristics of the audio files corresponding to each unplayed audio file in the audio file list.

[0250] In some embodiments of this application, the method by which the electronic device determines the above-mentioned second audio file combination can be found in the relevant description of step 203 above, and will not be repeated here in the embodiments of this application.

[0251] Step 502c: After the first audio file finishes playing, the electronic device plays the second audio file combination.

[0252] In this embodiment, when a user wants the electronic device to skip the currently playing audio segment, the electronic device can replace the currently playing audio segment with another audio file that is semantically similar or identical. This reduces the potential for audio content fragmentation caused by directly switching to the next audio segment. Simultaneously, the electronic device can regenerate a corresponding audio file combination for subsequent task durations based on the replaced audio file. This avoids situations where the replaced audio file and subsequent audio segments in the original first audio file have poor content continuity. Thus, users can be more immersed in the task during listening, experiencing less noticeable content fragmentation, resulting in a stronger sense of immersion, more efficient information acquisition, and further avoiding the tedious manual selection of audio for playback during task execution, thereby improving the flexibility of the electronic device's audio processing.

[0253] In this embodiment, the electronic device can support real-time recognition and execution of natural language commands, so as to flexibly adjust the playback strategy according to the user's instructions. In this way, the user can immerse themselves in the task without interrupting the user's task rhythm, avoiding the tedious operation of manually selecting audio for playback during the task execution process, and improving the flexibility of audio processing of the electronic device.

[0254] In some embodiments of this application, after step 203 above, the audio processing method provided in the embodiments of this application may further include the following step 601.

[0255] Step 601: When playing the first audio file combination, if the absolute value of the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to a second preset threshold, the electronic device performs one of the following:

[0256] Adjust the unplayed audio segments of the first audio file group;

[0257] Based on the user task information and the characteristics of each audio file, update the unplayed audio segments of the first audio file combination.

[0258] Understandably, if the absolute value of the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is less than the second preset threshold, the electronic device may not adjust the audio file combination being played.

[0259] In some embodiments of this application, the electronic device continuously tracks the difference between the remaining time of the task and the duration of the current playback stream through a real-time playback status monitoring module, so as to dynamically determine whether to trigger the content adjustment mechanism.

[0260] In some embodiments of this application, the electronic device can determine the remaining time of the aforementioned task. And the remaining duration of the current playlist, that is, the remaining playback duration of the first audio file combination mentioned above. Define the deviation amount ,like That is, the second preset threshold mentioned above.

[0261] In some embodiments of this application, in the deviation amount Exceeding the preset threshold For example, in the case of 60 seconds, the electronic device can automatically enter dynamic content adjustment mode to recalibrate the duration by deleting or inserting audio segments.

[0262] It should be noted that the second preset threshold may be the same as or different from the first preset threshold. For example, the second preset threshold may be less than the first preset threshold.

[0263] In some embodiments of this application, the electronic device can measure the deviation in real time. The detection can be performed either automatically or at a preset interval. For example, an electronic device can be tested every 5 minutes.

[0264] In some embodiments of this application, the aforementioned deviation amount is caused Situations that may exceed the preset threshold may include, but are not limited to, at least one of the following:

[0265] User-inputted voice commands involving duration adjustments, such as "play the next segment" or "extend by 10 minutes";

[0266] The task duration may change. For example, the electronic device may predict a duration of 10 minutes based on a driving application, but during the journey, due to traffic jams or other reasons, the electronic device may adjust the predicted duration.

[0267] In some embodiments of this application, in order to ensure that the adjusted playback stream fits the task length without compromising semantic continuity and content integrity, the electronic device can construct a multidimensional revenue optimization function using the following formula (8) to calculate... The overall fit score of the corresponding segment combination, and in accordance with the requirements Under the constraints, search The largest combination of segments.

[0268] Formula (8)

[0269] in, It can be used to represent the current set of adjustable playback segments, that is, the set of playable segments after duration compression and semantic optimization; It can be used to represent the user completion score of the i-th audio file corresponding to the i-th audio segment; It can be used to represent the audio topic vector of the i-th audio segment; It can be used to represent the semantic center vector corresponding to the first user task; It can be used to represent the actual playback duration of the i-th audio segment; It can be used to represent the reference duration that the i-th audio segment is expected to be used in the combination, which can be the duration obtained by dynamic programming in the above embodiments; These can be used to represent completion weight, semantic relevance weight, and duration error penalty weight, respectively. .

[0270] It should be noted that the above The specific weight values ​​can be optimized through model training and iteration to obtain the optimal solution, so as to adapt to different task scenarios and user behavior characteristics.

[0271] In some embodiments of this application, the above-mentioned "adjusting the unplayed audio segments of the first audio file combination" can be specifically implemented through the following steps C1 or C2.

[0272] Step C1: If the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to the second preset threshold, the electronic device adjusts the playback rate of the unplayed audio segment.

[0273] In some embodiments of this application, the electronic device can shorten the playback duration of the unplayed audio segment by increasing the playback rate of the unplayed audio segment, thereby making the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task less than a second preset threshold.

[0274] In some embodiments of this application, the electronic device can adjust the playback rate of at least a portion of an unplayed audio segment.

[0275] In some embodiments of this application, when an electronic device adjusts the playback magnification of multiple audio segments, the adjustment range corresponding to each of the multiple audio segments may be the same or different.

[0276] In some embodiments of this application, the electronic device can determine the adjustment range of the playback multiplier of the unplayed audio segment based on at least one of the completion score and semantic density of the audio file corresponding to the unplayed audio segment.

[0277] Step C2: If the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to the second preset threshold, the electronic device deletes the first audio segment from the unplayed audio segments and adjusts the playback order of the remaining unplayed audio segments.

[0278] In some embodiments of this application, the first audio segment can be an audio segment determined based on at least one of the following: the position of the audio segment in the first audio file combination, the completion score of the audio file corresponding to the audio segment, and the semantic density.

[0279] Understandably, when the duration of unplayed audio segments is long, electronic devices may prioritize cutting audio segments with relatively low overall importance, such as the last segment with low semantic density and low completion score.

[0280] In this embodiment, the electronic device can continuously monitor the deviation between the remaining task duration and the playlist duration. If the deviation exceeds a threshold, a dynamic adjustment mechanism is automatically triggered. This mechanism uses a multi-dimensional benefit optimization function to select suitable segments for addition or deletion, calibrating the duration while maintaining semantic coherence. In other words, this real-time interaction and dynamic adjustment capability allows users to control the playback process according to their needs. Furthermore, the electronic device can adapt to changes in task progress, avoiding interruptions due to duration deviations, significantly improving the flexibility and stability of the user experience.

[0281] In some embodiments of this application, the above-mentioned "updating the unplayed audio segments of the first audio file combination according to the user task information and the audio file characteristics corresponding to each audio file" can be specifically implemented through the following steps D1 and D2.

[0282] Step D1: If the difference between the remaining duration of the first user task and the remaining duration of the first audio file combination is greater than or equal to the second preset threshold, the electronic device determines the second audio file based on the audio file feature information of other audio files in the audio file list besides the first audio file combination.

[0283] Step D2: The electronic device adds the second audio file as an audio segment to the unplayed audio segment and adjusts the playback order of the added unplayed audio segments.

[0284] In some embodiments of this application, the second audio file can be an audio file determined based on the semantic similarity between the audio topic vector of the audio file and the semantic topic vector corresponding to the task type of the first user task, as well as the completion score of the audio file.

[0285] Understandably, if the duration of the unplayed audio segment is insufficient, the electronic device can select a higher-quality audio file from the list of unused audio files. high Short content is used to quickly fill gaps without disrupting the overall tone of the content.

[0286] In some embodiments of this application, since the entire supplementation and deletion process described above needs to maintain the semantic order structure of the original combination, the electronic device can call the semantic transfer metric function in step 301b above. The edited audio segments are rearranged and recombined to ensure the user doesn't perceive any noticeable interruptions or logical jumps during the transition. Simultaneously, the electronic device continuously feeds the final dynamic adjustment result as a new playback stream to the voice control module and stores it in the behavior log, providing data for future model parameter adjustments. Specifically, this can be done as follows: Figure 8 The flowchart shown.

[0287] In this embodiment, the electronic device can continuously monitor the deviation between the remaining task duration and the playlist duration. If the deviation exceeds a threshold, a dynamic adjustment mechanism is automatically triggered. This mechanism uses a multi-dimensional benefit optimization function to select suitable segments for addition or deletion, calibrating the duration while maintaining semantic coherence. In other words, this real-time interaction and dynamic adjustment capability allows users to control the playback process according to their needs. Furthermore, the electronic device can adapt to changes in task progress, preventing premature endings due to duration deviations, significantly improving the flexibility and stability of the user experience.

[0288] In some embodiments of this application, after the first audio file combination playback ends, the electronic device can generate a user behavior analysis report to record indicators such as task type, actual usage time, and content preference deviation. The electronic device can then optimize the weight parameters of subsequent matching algorithms based on the recorded indicators.

[0289] For example, after the playback stream ends, the electronic device can activate the user behavior attribution module to perform full behavior retrospective and structured analysis on the playback process of this audio file combination, and automatically generate a behavior analysis report as an important input for subsequent content matching model parameter adjustment and user profile iteration.

[0290] In some embodiments of this application, the electronic device may first record the type of the current task. and actual usage time And compared with the task target duration set by the user or predicted by the electronic device in step 201 above. Perform a comparison and calculate the deviation rate. To assess how well the electronic device's design aligns with real-world user behavior, the device analyzes the user's listening behavior for each segment during playback, identifying actions such as skipping, speeding up, and pausing, and combines this with the semantic topic vector of each audio file. With the main task semantic center Construct a content preference deviation index as shown in the following formula (9). .

[0291] Formula (9)

[0292] in, It can be used to indicate the number of audio segments played this time. It can be used to represent the semantic similarity between the i-th audio segment and the semantics of the user task's main task. This can be used to represent the user behavior retention coefficient of a segment, with weights assigned based on behaviors such as complete playback, skipping, and interruption, reflecting the semantic consistency between the user's actual attention and the content. For example, the retention coefficient corresponding to completing playback... Skip the corresponding The interruption value can be between 0 and 1, depending on the percentage of playback time.

[0293] Understandable, A value close to zero indicates that the content actually received by the user is highly consistent with the task scenario; conversely, a value higher indicates that the theme adaptation strategy needs to be adjusted.

[0294] In some embodiments of this application, during multiple rounds of use, the electronic device performs cross-task type summary statistics on the above-mentioned indicators to optimize the core algorithm of the electronic device in reverse, such as optimizing the topic matching weight in dynamic programming based on content preference deviation, and adjusting the personalized weight coefficients of the prediction model by analyzing the task duration deviation rate. And in combination with user behavior retention coefficient, the playback rate strategy function is optimized in reverse. The settings allow electronic devices to gradually approach the optimal coupling point between user preferences and task rhythm, ultimately synchronizing the analysis report summary to the user end and using it as an initialization strategy to participate in the construction cycle of the next round of task matching model.

[0295] In this way, through a closed-loop iteration mechanism, electronic devices can continuously learn users' behavior patterns and interests, making task duration prediction more accurate, content matching more tailored to individual needs, and playback strategies more in line with user habits. With long-term use, the intelligence level and personalized service capabilities of electronic devices are continuously improved, gradually forming a personalized listening solution that is highly adapted to the user.

[0296] In view of the various scenarios in which the embodiments of this application can be applied, and in conjunction with the various implementation schemes of the embodiments of this application described above, specific examples are given below to illustrate the implementation process of the embodiments of this application in various scenarios.

[0297] When an electronic device is running a navigation application, it can automatically determine that the user's current task is a driving task based on the application. Therefore, the electronic device can use this navigation application to obtain a predicted arrival time at the destination, such as 40 minutes.

[0298] The electronic device can access podcast channels that a user has favorited or followed in an audio playback application, using the podcast interfaces within those channels as a list of audio files. Based on the task duration, it determines at least one candidate podcast set, such as candidate podcast set 1 to candidate podcast set 3. The electronic device can then remove noisy audio segments from the podcast programs included in these three candidate podcast sets, such as removing end credits or end-credit advertisements, or silent segments. After the removal process, the total audio duration corresponding to each of the three candidate podcast sets is between 40 ± 2 minutes.

[0299] For candidate podcast set 1, the electronic device can obtain the topic tag for each podcast program in this set, expand it into a topic set through semantic analysis, and obtain the corresponding topic vector. Then, based on the topic vector of candidate podcast set 1, the electronic device can calculate the semantic similarity between it and the semantic topic vector corresponding to the task type, and calculate the topic fit score for candidate podcast set 1. Similarly, the electronic device can perform similar operations on candidate podcast sets 2 and 3 respectively. Then, the electronic device can determine the candidate audio set with the highest topic fit score as the audio set to be used, such as candidate audio set 2.

[0300] Then, the electronic device can calculate the semantic similarity between different audio programs based on the start and end segments of each audio program in the candidate audio set 2, to determine the arrangement order of each audio program, and further perform merging processing based on this arrangement order to obtain the final audio combination to be played. At the same time, the electronic device can appropriately accelerate low-completion content and maintain the rhythm of high-completion content based on information such as the completion score and semantic density of the audio programs, so as to balance information density and listening comfort.

[0301] During the playback of audio sequences on an electronic device, the device can acquire the user's voice commands and the corresponding user intent information, and then execute the corresponding operation. Simultaneously, the device can continuously monitor the deviation between the remaining task duration and the playlist duration. If the deviation exceeds a threshold, a dynamic adjustment mechanism is automatically triggered. This mechanism uses a multi-dimensional benefit optimization function to select suitable segments for addition or deletion, calibrating the duration while ensuring semantic coherence.

[0302] Finally, after the audio combination playback ends, the electronic device can generate a user behavior analysis report to record indicators such as task type, actual usage time, and content preference deviation. The electronic device can then use the recorded indicators to optimize the weight parameters of subsequent matching algorithms.

[0303] Furthermore, this application embodiment uses "task duration" as a strong constraint, combined with the core technical framework of "semantic matching" and "dynamic reorganization", which has high plasticity and cross-domain empowerment potential. Its value goes far beyond optimizing the audio file listening experience, and can also give rise to a brand-new content consumption paradigm and business model.

[0304] First, the application scenarios can be expanded from "listening to content" to "using content." Because this application can transform content without a fixed duration into a "functional content stream" precisely matched to specific task scenarios, this idea can be widely applied in scenarios including, but not limited to, any of the following:

[0305] Personalized education and training: In online learning or corporate training scenarios, a customized audio and video course with precise duration, coherent knowledge points, and suitability to individual learning progress can be dynamically generated from a massive course slice library based on the "30-minute learning time" set by the learner.

[0306] Smart Assistant and News Service: Users can give a command to a smart speaker or in-car assistant to "give me a 15-minute tech and financial morning news report". The electronic device can capture news clips from multiple sources in real time, automatically deduplicate, sort, and seamlessly splice them to generate a completely personalized news broadcast.

[0307] Interactive tour guides: In smart scenic spots or museums, a dynamic audio tour route that matches the length of a visitor's "1-hour tour time" and interests, such as "preferring ceramics," can be generated. The audio tour route can also be adjusted in real time based on the time visitors spend on a particular exhibit.

[0308] Health and sleep aid applications: In fitness scenarios, music, coaching guidance, and motivational voices can be dynamically arranged according to the user's set 45-minute running plan; in sleep aid scenarios, a guided audio with precise duration, a tempo that gradually slows down, and content that transitions from a story to white noise can be generated.

[0309] Secondly, regarding the cross-domain empowerment of the core technology, the solution provided in this application can construct a "time-aware" content service platform. It is understood that the core algorithms in this application, such as the dynamic programming duration matching model, semantic coherence evaluation, and user behavior feedback loop, can be abstracted into a general "time-aware content service platform" to serve as technological infrastructure empowering a wider range of digital content domains. For example,

[0310] Video streaming: Applicable to short video platforms, electronic devices can automatically generate a seamless, coherent collection of short videos based on the user's "20-minute lunch break" needs, thereby increasing user engagement.

[0311] Digital reading: Electronic devices can generate audio summaries of articles that allow users to "read the core content in 10 minutes" or reorganize long reports into "30-minute essential versions" suitable for listening to during commutes.

[0312] Thirdly, innovation in business models and ecosystem building. Based on the ideas in this application, new possibilities can be opened up for the commercialization of content platforms.

[0313] B2B technology licensing: The core algorithm engine can be licensed to third-party application developers, such as fitness applications, online education platforms, and car manufacturers, in the form of APIs, so that they can quickly integrate the "task duration matching" function and build a content service ecosystem.

[0314] Dynamic audio ad delivery: This technology changes the "one-size-fits-all" in-app advertising model. Electronic devices can dynamically insert highly relevant, user-friendly short audio ads into the generated content stream, based on the user's task scenario, content theme, and remaining time, thereby enhancing both advertising value and user experience.

[0315] Deep integration with Internet of Things (IoT) devices: It can be linked with smart home electronic devices. When the user's smart coffee machine starts working, that is, when the task begins, the electronic device can automatically play a morning audio file with a duration exactly equal to the coffee making time, truly achieving seamless integration of service with the user's lifestyle.

[0316] It should be noted that the above-described method embodiments, or the various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0317] It should be noted that the audio processing method provided in this application embodiment can be executed by an audio processing device. This application embodiment uses an audio processing device executing the audio processing method as an example to illustrate the audio processing device provided in this application embodiment.

[0318] Figure 9 A schematic diagram of a possible structure of the audio processing apparatus involved in an embodiment of this application is shown. For example... Figure 9 As shown, the audio processing device 80 may include: an acquisition module 81, a processing module 82, and a playback module 83;

[0319] The module includes an acquisition module 81, a processing module 82, and a determination module 83.

[0320] The acquisition module 81 is used to acquire user task information for the first user task; the user task information includes task type and task duration.

[0321] Processing module 82 is used to perform feature extraction processing on each audio file in the audio file list to obtain the audio file features corresponding to each audio file; the audio file features include duration dimension features, content dimension features and interest dimension features;

[0322] The determining module 83 is used to determine a first audio file combination based on the user task information obtained by the obtaining module 81 and the audio file characteristics corresponding to each audio file obtained by the processing module 82; the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task.

[0323] In one possible implementation, the aforementioned determining module 83 is specifically used to determine the task type based on first information, which includes at least one of the following: user activity information, time information, user status information, and environment information; and to determine the task duration based on the task duration of historical user tasks corresponding to the task type.

[0324] In one possible implementation, the acquisition module 81 is specifically used to acquire the average task duration, the standard deviation of the task duration, and the audio listening completion time of historical user tasks; the determination module 83 is specifically used to perform a weighted calculation on the average task duration, the standard deviation of the task duration, and the audio listening completion time to determine the task duration.

[0325] In one possible implementation, the audio processing apparatus 80 provided in this application embodiment may further include: a processing module 82; the aforementioned determining module 83 is specifically used to obtain at least one set of candidate audio files based on the task duration of the first user task and the duration dimension features corresponding to each audio file, wherein the absolute value of the difference between the total audio duration of each candidate audio file set and the task duration of the first user task is less than or equal to a first preset threshold; and to determine a first audio file set from the at least one set of candidate audio files based on the content dimension features and interest dimension features corresponding to each set of candidate audio files; the processing module 82 is used to merge the audio files in the first audio file set to obtain a first audio file combination.

[0326] In one possible implementation, the processing module 82 is further configured to, before obtaining at least one candidate audio file set based on the task duration of the first user task and the duration dimension features corresponding to each audio file, eliminate noisy audio segments in each audio file to obtain processed audio files; the determining module 83 is specifically configured to obtain at least one candidate audio file set based on the task duration and the duration dimension features corresponding to each processed audio file.

[0327] In one possible implementation, the audio processing apparatus 80 provided in this application embodiment may further include: a calculation module; the calculation module is used to calculate the semantic similarity between the content dimension features corresponding to each candidate audio file set and the semantic topic vector corresponding to the task type of the first user task, to obtain the semantic similarity corresponding to each candidate audio file set; and to calculate the topic adaptation score corresponding to each candidate audio file set based on the semantic similarity and interest dimension features corresponding to each candidate audio file set; the aforementioned determining module 83 is specifically used to determine the first audio file set based on the topic adaptation scores corresponding to each candidate audio file set.

[0328] In one possible implementation, the processing module 82 is further configured to perform semantic expansion processing on the audio topic tags of each audio file in the audio file list before calculating the semantic similarity between the content dimension features corresponding to each candidate audio file set and the semantic topic vector corresponding to the task type of the first user task, and obtaining the semantic similarity corresponding to each candidate audio file set; and to perform normalization processing on the audio topic sets of each audio file, to obtain the audio topic vector corresponding to each audio file; and to obtain the content dimension features corresponding to each candidate audio file set based on the audio topic vector corresponding to each audio file.

[0329] In one possible implementation, the processing module 82 is further configured to determine the arrangement order of at least two audio files based on the semantic vectors of at least two audio files included in the first audio file set before merging the audio files in the first audio file set to obtain the first audio file combination; specifically, the processing module 82 is configured to merge the at least two audio files according to the arrangement order to obtain the first audio file combination.

[0330] In one possible implementation, the audio processing apparatus 80 provided in this application embodiment may further include: a calculation module; the calculation module is used to calculate the semantic similarity between the semantic vectors of the first audio segments of every two audio files in the first audio file set, the first audio segment including a start audio segment and an end audio segment; the aforementioned determining module 83 is specifically used to determine the arrangement order of at least two audio files based on the semantic similarity between the semantic vectors of the first audio segments of every two audio files.

[0331] In one possible implementation, the audio processing apparatus 80 provided in this application embodiment may further include: a processing module 82; the processing module 82 is used to, after determining a first audio file combination based on user task information and the characteristics of the audio files corresponding to each audio file, adjust the playback multiplier of the audio segments in the first audio file combination based on the task duration of the first user task and the completion score of the audio files corresponding to each audio segment in the first audio file combination; wherein the adjusted playback multiplier of the audio segments is within the range corresponding to the speech rate carrying capacity.

[0332] In one possible implementation, the audio processing apparatus 80 provided in this application embodiment may further include: a receiving module and a processing module 82; the receiving module is used to receive user voice commands while playing the first audio file combination after determining the first audio file combination based on user task information and the audio file characteristics corresponding to each audio file; the processing module 82 is used to update the first audio file combination according to the user voice commands.

[0333] In one possible implementation, the aforementioned user voice command is to switch the currently playing audio segment; the audio processing device 80 provided in this application embodiment may further include: a playback module; the playback module is used to stop playing the currently playing audio segment and start playing a first audio file, the first audio file being the audio file with the highest semantic similarity to the currently playing audio segment among the unplayed audio files included in the audio file list; the aforementioned determining module 83 is specifically used to determine a second audio file combination based on user task information and the audio file characteristics corresponding to each unplayed audio file; the absolute value of the difference between the sum of the audio duration of the second audio file combination and the audio duration of the first audio file and the remaining duration of the first user task is less than or equal to a first preset threshold, and the audio content of the second audio file combination matches the task type of the first user task; the playback module is also used to play the second audio file combination after the first audio file has finished playing.

[0334] In one possible implementation, the processing module 82 is further configured to, after determining the first audio file combination based on the user task information and the audio file characteristics corresponding to each audio file, perform one of the following actions when playing the first audio file combination: if the absolute value of the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to a second preset threshold: adjust the unplayed audio segments of the first audio file combination; update the unplayed audio segments of the first audio file combination based on the user task information and the audio file characteristics corresponding to each audio file.

[0335] In one possible implementation, the processing module 82 is specifically used to adjust the playback rate of unplayed audio segments when the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to a second preset threshold; or, when the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to the second preset threshold, to delete the first audio segment in the unplayed audio segments and adjust the playback order of the remaining unplayed audio segments; wherein, the first audio segment is an audio segment determined based on at least one of the following: the position of the audio segment in the first audio file combination, the completion score of the audio file corresponding to the audio segment, and the semantic density.

[0336] In one possible implementation, the determining module 83 is specifically used to determine a second audio file based on the audio file feature information of other audio files in the audio file list besides the first audio file combination when the difference between the remaining duration of the first user task and the remaining duration of the first audio file combination is greater than or equal to a second preset threshold; the processing module 82 is specifically used to add the second audio file as an audio segment to the unplayed audio segments and adjust the playback order of the added unplayed audio segments; wherein, the second audio file is an audio file determined based on the semantic similarity between the audio topic vector of the audio file and the semantic topic vector corresponding to the task type of the first user task, as well as the completion score of the audio file.

[0337] In the audio processing device provided in this application embodiment, the audio processing device can automatically combine audio files that match the task type and task duration from the audio file list according to the task type and task duration corresponding to the user task. In other words, the combined audio files can meet the user's listening preferences during the execution of a certain task, thereby satisfying the user's actual listening needs for audio files during the execution of that task, allowing the user to immerse themselves in the task. This avoids the tedious operation of manually selecting audio for playback during the execution of a task, and improves the flexibility of audio processing by the audio processing device.

[0338] The audio processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0339] The audio processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0340] The audio processing apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0341] Optionally, such as Figure 10 As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described audio processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0342] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0343] Figure 11 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0344] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0345] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 11 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0346] The processor 110 is configured to acquire user task information for a first user task, including task type and task duration; and to perform feature extraction processing on each audio file in the audio file list to obtain audio file features corresponding to each audio file, including duration dimension features, content dimension features, and interest dimension features; and to determine a first audio file combination based on the user task information and the audio file features corresponding to each audio file, wherein the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task.

[0347] Optionally, the processor 110 is specifically configured to determine the task type based on first information, the first information including at least one of the following: user activity information, time information, user status information, and environment information; and to determine the task duration based on the task duration of historical user tasks corresponding to the task type.

[0348] Optionally, the processor 110 is specifically used to obtain the average task duration, the standard deviation of the task duration, and the audio listening completion time of historical user tasks; and to perform a weighted calculation on the average task duration, the standard deviation of the task duration, and the audio listening completion time to determine the task duration.

[0349] Optionally, the processor 110 is specifically configured to obtain at least one set of candidate audio files based on the task duration of the first user task and the duration dimension features corresponding to each audio file, wherein the absolute value of the difference between the total audio duration of each candidate audio file set and the task duration of the first user task is less than or equal to a first preset threshold; and to determine a first audio file set from the at least one set of candidate audio files based on the content dimension features and interest dimension features corresponding to each set of candidate audio files; and to merge the audio files in the first audio file set to obtain a first audio file combination.

[0350] Optionally, the processor 110 is further configured to, before obtaining at least one candidate audio file set based on the task duration of the first user task and the duration dimension features corresponding to each audio file, eliminate noisy audio segments in each audio file to obtain processed audio files; specifically, the processor 110 is configured to obtain at least one candidate audio file set based on the task duration and the duration dimension features corresponding to each processed audio file.

[0351] Optionally, the processor 110 is specifically configured to calculate the semantic similarity between the content dimension features corresponding to each candidate audio file set and the semantic topic vector corresponding to the task type of the first user task, thereby obtaining the semantic similarity of each candidate audio file set; and calculate the topic adaptation score of each candidate audio file set based on the semantic similarity and interest dimension features; and determine the first audio file set based on the topic adaptation scores of each candidate audio file set.

[0352] Optionally, the processor 110 is further configured to perform semantic expansion processing on the audio topic tags of each audio file in the audio file list before calculating the semantic similarity between the content dimension features corresponding to each candidate audio file set and the semantic topic vector corresponding to the task type of the first user task, to obtain the semantic similarity corresponding to each candidate audio file set; and to perform normalization processing on the audio topic sets of each audio file to obtain the audio topic vector corresponding to each audio file; and to obtain the content dimension features corresponding to each candidate audio file set based on the audio topic vector corresponding to each audio file.

[0353] Optionally, the processor 110 is further configured to determine the arrangement order of at least two audio files based on the semantic vectors of at least two audio files included in the first audio file set before merging the audio files in the first audio file set to obtain the first audio file combination; specifically, the processor 110 is configured to merge the at least two audio files according to the arrangement order to obtain the first audio file combination.

[0354] Optionally, the processor 110 is specifically configured to calculate the semantic similarity between the semantic vectors of the first audio segments of every two audio files in the first audio file set, the first audio segment including a start audio segment and an end audio segment; and to determine the arrangement order of at least two audio files based on the semantic similarity between the semantic vectors of the first audio segments of every two audio files.

[0355] Optionally, the processor 110 is further configured to, after determining the first audio file combination based on the user task information and the characteristics of the audio files corresponding to each audio file, adjust the playback rate of the audio segments in the first audio file combination based on the task duration of the first user task and the completion score of the audio files corresponding to each audio segment in the first audio file combination; wherein the playback rate of the adjusted audio segments is within the range corresponding to the speech rate carrying capacity.

[0356] Optionally, the user input unit 107 is used to receive user voice commands while playing the first audio file combination after determining the first audio file combination based on the user task information and the characteristics of the audio files corresponding to each audio file.

[0357] The processor 110 is also used to update the first audio file combination according to the user's voice command.

[0358] Optionally, the aforementioned user voice command is to switch the currently playing audio segment; the audio output unit 103 is used to stop playing the currently playing audio segment and start playing the first audio file, which is the audio file with the highest semantic similarity to the currently playing audio segment among the unplayed audio files included in the audio file list; the processor 110 is specifically used to determine the second audio file combination based on the user task information and the audio file characteristics corresponding to each unplayed audio file; the absolute value of the difference between the sum of the audio duration of the second audio file combination and the audio duration of the first audio file and the remaining duration of the first user task is less than or equal to a first preset threshold, and the audio content of the second audio file combination matches the task type of the first user task; the audio output unit 103 is also used to play the second audio file combination after the first audio file finishes playing.

[0359] Optionally, the processor 110 is further configured to, after determining the first audio file combination based on the user task information and the audio file characteristics corresponding to each audio file, when playing the first audio file combination, if the absolute value of the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to a second preset threshold, perform one of the following: adjust the unplayed audio segments of the first audio file combination; update the unplayed audio segments of the first audio file combination based on the user task information and the audio file characteristics corresponding to each audio file.

[0360] Optionally, the processor 110 is specifically configured to adjust the playback rate of unplayed audio segments when the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to a second preset threshold; or, when the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to the second preset threshold, delete the first audio segment in the unplayed audio segments and adjust the playback order of the remaining unplayed audio segments; wherein, the first audio segment is an audio segment determined based on at least one of the following: the position of the audio segment in the first audio file combination, the completion score of the audio file corresponding to the audio segment, and the semantic density.

[0361] Optionally, the processor 110 is specifically configured to, when the difference between the remaining duration of the first user task and the remaining duration of the first audio file combination is greater than or equal to a second preset threshold, determine a second audio file based on the audio file feature information of other audio files in the audio file list besides the first audio file combination; add the second audio file as an audio segment to the unplayed audio segments, and adjust the playback order of the added unplayed audio segments; wherein, the second audio file is an audio file determined based on the semantic similarity between the audio topic vector of the audio file and the semantic topic vector corresponding to the task type of the first user task, and the completion score of the audio file.

[0362] In the electronic device provided in this application embodiment, the electronic device can automatically combine audio files that match the task type and duration from the audio file list according to the task type and duration corresponding to the user task. In other words, the combined audio files can meet the user's listening preferences during the execution of a certain task, thereby satisfying the user's actual listening needs for audio files during the execution of that task, allowing the user to immerse themselves in the task. This avoids the tedious operation of manually selecting audio for playback during the execution of a task and improves the flexibility of audio processing of the electronic device.

[0363] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0364] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.

[0365] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0366] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0367] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0368] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0369] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0370] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0371] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0372] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, further details are omitted here.

[0373] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0374] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0375] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio processing method, characterized in that, The method includes: Obtain the user task information for the first user task; the user task information includes the task type and task duration. Feature extraction is performed on each audio file in the audio file list to obtain the audio file features corresponding to each audio file; the audio file features include duration dimension features, content dimension features, and interest dimension features; Based on the user task information and the characteristics of the audio files corresponding to each audio file, a first audio file combination is determined; the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task.

2. The method according to claim 1, characterized in that, The step of obtaining the user task information of the first user task includes: The task type is determined based on the first information, which includes at least one of the following: user activity information, time information, user status information, and environment information; The task duration is determined based on the task duration of historical user tasks corresponding to the task type.

3. The method according to claim 2, characterized in that, Determining the task duration based on the task duration of historical user tasks corresponding to the task type includes: Obtain the average task duration, standard deviation of task duration, and audio listening completion time of the historical user tasks; The average task duration, the standard deviation of the task duration, and the audio listening completion time are weighted and calculated to determine the task duration.

4. The method according to claim 1, characterized in that, The step of determining the first audio file combination based on the user task information and the characteristics of each audio file includes: Based on the task duration of the first user task and the duration dimension features corresponding to each audio file, at least one set of candidate audio files is obtained, and the absolute value of the difference between the total audio duration of each set of candidate audio files and the task duration of the first user task is less than or equal to the first preset threshold. Based on the content dimension features and interest dimension features corresponding to each of the candidate audio file sets, the first audio file set is determined from the at least one candidate audio file set; The audio files in the first audio file set are merged to obtain the first audio file combination.

5. The method according to claim 4, characterized in that, Before obtaining at least one set of candidate audio files based on the task duration of the first user task and the duration dimension features corresponding to each of the audio files, the method further includes: For each of the audio files, the noisy audio segments in each audio file are removed to obtain the processed audio files. Based on the task duration of the first user task and the duration dimension features corresponding to each of the audio files, at least one set of candidate audio files is obtained, including: Based on the task duration and the duration dimension features corresponding to each of the processed audio files, the at least one set of candidate audio files is obtained.

6. The method according to claim 4, characterized in that, The step of determining the first audio file set from the at least one candidate audio file set based on the content dimension features and interest dimension features corresponding to each of the candidate audio file sets includes: Calculate the semantic similarity between the content dimension features corresponding to each of the candidate audio file sets and the semantic topic vector corresponding to the task type of the first user task, to obtain the semantic similarity corresponding to each of the candidate audio file sets. Based on the semantic similarity and interest dimension features corresponding to each of the candidate audio file sets, the topic adaptation score corresponding to each of the candidate audio file sets is calculated. The first audio file set is determined based on the topic adaptation score corresponding to each of the candidate audio file sets.

7. The method according to claim 6, characterized in that, Before calculating the semantic similarity between the content dimension features corresponding to each of the candidate audio file sets and the semantic topic vector corresponding to the task type of the first user task, and obtaining the semantic similarity for each of the candidate audio file sets, the method further includes: Semantic expansion processing is performed on the audio topic tags of each audio file in the audio file list to obtain the audio topic set corresponding to each audio file; The audio topic sets of each audio file are normalized to obtain the audio topic vectors corresponding to each audio file. Based on the audio topic vectors corresponding to each of the audio files, the content dimension features corresponding to each of the candidate audio file sets are obtained.

8. The method according to claim 4, characterized in that, Before merging the audio files in the first audio file set to obtain the first audio file combination, the method further includes: Based on the semantic vectors of at least two audio files included in the first audio file set, determine the arrangement order of the at least two audio files; The step of merging the audio files in the first audio file set to obtain the first audio file combination includes: According to the stated arrangement order, the at least two audio files are merged to obtain the first audio file combination.

9. The method according to claim 8, characterized in that, Determining the order of the at least two audio files based on their semantic vectors from the first set of audio files includes: Calculate the semantic similarity between the semantic vectors of the first audio segments of every two audio files in the first audio file set, wherein the first audio segment includes a start audio segment and an end audio segment; The arrangement order of the at least two audio files is determined based on the semantic similarity between the semantic vectors of the first audio segments of each pair of audio files.

10. The method according to claim 1, characterized in that, After determining the first audio file combination based on the user task information and the characteristics of each audio file, the method further includes: Based on the task duration of the first user task and the completion score of the audio files corresponding to each audio segment in the first audio file combination, the playback multiplier of the audio segments in the first audio file combination is adjusted. The playback rate of the adjusted audio clips is within the range corresponding to the speech rate carrying capacity.

11. The method according to claim 1, characterized in that, After determining the first audio file combination based on the user task information and the characteristics of each audio file, the method further includes: While playing the first audio file combination, receive user voice commands; The first audio file combination is updated according to the user's voice command.

12. The method according to claim 11, characterized in that, The user's voice command is to switch the currently playing audio segment; Updating the first audio file combination according to the user's voice command includes: Stop playing the currently playing audio segment and start playing the first audio file, which is the audio file with the highest semantic similarity to the currently playing audio segment among the unplayed audio files included in the audio file list; Based on the user task information and the characteristics of the audio files corresponding to each of the unplayed audio files, a second audio file combination is determined; the absolute value of the difference between the sum of the audio duration of the second audio file combination and the audio duration of the first audio file and the remaining duration of the first user task is less than or equal to the first preset threshold, and the audio content of the second audio file combination matches the task type of the first user task. After the first audio file finishes playing, the second audio file combination is played.

13. The method according to claim 1, characterized in that, After determining the first audio file combination based on the user task information and the characteristics of each audio file, the method further includes: When playing the first audio file combination, if the absolute value of the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to a second preset threshold, one of the following actions is performed: Adjust the unplayed audio segments of the first audio file combination; Based on the user task information and the characteristics of the audio files corresponding to each of the audio files, update the unplayed audio segments of the first audio file combination.

14. The method according to claim 13, characterized in that, The adjustment of the unplayed audio segments of the first audio file combination includes: If the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to the second preset threshold, the playback rate of the unplayed audio segment is adjusted. or, If the difference between the remaining duration of the first audio file combination and the remaining duration of the first user task is greater than or equal to the second preset threshold, delete the first audio segment in the unplayed audio segments and adjust the playback order of the remaining unplayed audio segments. The first audio segment is an audio segment determined based on at least one of the following: the position of the audio segment in the first audio file combination, the completion score of the audio file corresponding to the audio segment, and the semantic density.

15. The method according to claim 13, characterized in that, The step of updating the unplayed audio segments of the first audio file combination based on the user task information and the audio file characteristics corresponding to each of the audio files includes: If the difference between the remaining duration of the first user task and the remaining duration of the first audio file combination is greater than or equal to the second preset threshold, the second audio file is determined based on the audio file feature information of other audio files in the audio file list besides the first audio file combination. Add the second audio file as an audio segment to the unplayed audio segment, and adjust the playback order of the added unplayed audio segments; The second audio file is determined based on the semantic similarity between the audio topic vector of the audio file and the semantic topic vector corresponding to the task type of the first user task, as well as the completion score of the audio file.

16. An audio processing apparatus, characterized in that, The audio processing device includes: an acquisition module, a processing module, and a determination module; The acquisition module is used to acquire user task information of the first user task; the user task information includes task type and task duration; The processing module is used to perform feature extraction processing on each audio file in the audio file list to obtain the audio file features corresponding to each audio file; the audio file features include duration dimension features, content dimension features and interest dimension features; The determining module is used to determine a first audio file combination based on the user task information obtained by the acquiring module and the audio file features corresponding to each audio file obtained by the processing module; the absolute value of the difference between the audio duration of the first audio file combination and the task duration of the first user task is less than or equal to a first preset threshold, and the audio content of the first audio file combination matches the task type of the first user task.

17. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the audio processing method as described in any one of claims 1 to 15.

18. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio processing method as described in any one of claims 1 to 15.

19. A computer program product, characterized in that, The computer program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the audio processing method as described in any one of claims 1 to 15.