Method, device, intelligent equipment and system for determining audio features

By building an audio feature library that is connected to the intelligent device, using the server to obtain audio features, the problem of delay in vehicle local audio feature extraction is solved, real-time audio feature acquisition is realized, and user experience is optimized.

CN120279935APending Publication Date: 2025-07-08NIO TECH ANHUI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510336792.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-20
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

传统的音频特征提取在车辆本地进行时由于算法性能限制和延迟问题,无法实时为车载应用提供音频特征,影响用户体验。

Method used

Build an audio feature library that communicates with smart devices, obtain audio features matching audio identifiers through the server, reduce local computing resources and delay, and realize real-time audio feature acquisition.

Benefits of technology

The audio features required by the target function can be obtained without the need for local audio feature algorithm calculations, reducing the calculation process and delay, and optimizing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279935A_ABST
    Figure CN120279935A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for determining audio features, intelligent equipment and a system, and belongs to the technical field of automobiles, and the method comprises the following steps: determining a first feature category which is a feature category required when an intelligent application realizes a target function, and obtaining an audio identifier of a target audio currently played in the intelligent equipment; sending a feature request message to a server based on the audio identifier; and under the condition that a first target audio feature matched with the audio identifier and sent by the server is received, allocating an audio feature matched with the first feature category in the first target audio feature to the intelligent application, so that the intelligent application realizes a target function based on the audio feature. Therefore, the audio features of the target audio can be obtained without calculation of a local audio feature algorithm, the calculation process of the audio features is reduced, the time delay of obtaining the audio features is reduced, the application obtains the needed audio features in real time, and the user experience is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the patent application titled "Method, Device, Intelligent Device and System for Determining Audio Features" with the application number 202411902565.9 and filed with the China Patent Office on December 20, 2024, the entire content of which is incorporated herein by reference. Technical Field

[0002] This application belongs to the field of automotive technology, and particularly relates to a method, device, intelligent device and system for determining audio features. Background Art

[0003] With the increasing level of vehicle intelligence, various audio-based innovative applications in in-vehicle systems emerge in an endless stream. For example, action control applications of in-vehicle artificial intelligence robots, in-vehicle ambient light rhythm applications, in-vehicle microphone-free singing applications, etc. When the above innovative applications are running, they often need the audio features of the audio to formulate corresponding control strategies. For example, in the in-vehicle ambient light rhythm application, the vehicle obtains the audio features of the music currently played by the in-vehicle multimedia device, formulates a corresponding light display strategy based on the audio features, and then controls the state of the vehicle lights based on the light display strategy.

[0004] In the related art, in order to obtain audio features, an audio processing algorithm is installed locally in the vehicle. When it is necessary to obtain the audio features of a certain piece of audio, the audio features of the audio are obtained based on the audio processing algorithm installed locally in the vehicle. However, due to technical reasons such as the limitation of the algorithm performance locally in the vehicle and the inherent delay of the music processing algorithm itself, it is impossible to provide audio features for various in-vehicle applications in real time, thereby affecting the user experience. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, intelligent device and system for determining audio features, aiming to solve the problem that the traditional audio feature extraction has a large delay and affects the user experience.

[0006] The first aspect of the embodiment of this application provides a method for determining audio features, and the method includes:

[0007] Determine a first feature category, where the first feature category is the feature category required for the intelligent application to achieve the target function, and the intelligent application is an application running in the intelligent device;

[0008] Obtain the audio identifier of the target audio currently played in the intelligent device;

[0009] Based on the audio identifier, send a feature request message to the server, where the feature request message requests the server to obtain the first target audio feature matching the audio identifier in the audio feature library;

[0010] Upon receiving the first target audio feature, allocate the audio features in the first target audio feature that match the first feature category to the intelligent application, so that the intelligent application can implement the target function based on the audio features in the first target audio feature that match the first feature category.

[0011] In the embodiments of the present application, an audio feature library communicatively connected to the intelligent device is constructed. The audio feature library stores the audio features of multiple audios, where each audio feature is associated with its corresponding audio identifier. In this way, when the intelligent application needs to obtain the target audio feature of the music function, it only needs to query the audio feature of the target audio required for the intelligent application to implement the target function from the audio feature library, and feedback the first target audio feature of the queried target audio to the intelligent device. After receiving the first target audio feature from the cloud, the intelligent device determines whether the first target audio feature contains the audio features of the required first feature category, and allocates the audio features in the first target audio feature that match the first feature category to the intelligent application, so that the intelligent application can implement the target function based on the audio features that match the first feature category. In this way, the audio features required for the target function can be obtained without local audio feature algorithm calculation, avoiding occupying the computing resources of the intelligent device. Moreover, the calculation process of the audio features is reduced, and the delay in obtaining the audio features is reduced, enabling the target function to obtain the required audio features in real time, thereby optimizing the user experience.

[0012] In some embodiments, the method further includes:

[0013] When the first target audio feature does not contain all the audio features of the first feature category, determine the second target audio feature of the target audio based on the first audio feature algorithm. The second target audio feature contains the audio features of the first feature category that are missing in the first target audio feature. The first audio feature algorithm is an audio feature algorithm deployed on the intelligent device;

[0014] Allocate the second target audio feature to the intelligent application, so that the intelligent application can implement the target function based on the second target audio feature.

[0015] In this implementation, when the feature categories included in the first target audio feature do not include all the first feature categories, calculate the audio features corresponding to the target feature categories required for the music function, and thus control the music function based on the audio features. In this way, the intelligent device can calculate the audio features required for the music function, ensuring that when the audio feature library does not store the audio features that meet the application requirements, the intelligent device can still normally implement the application requirements and ensure the normal operation of the music function.

[0016] In some embodiments, determining the second target audio feature of the target audio based on the first audio feature algorithm includes:

[0017] Determine a second feature category based on the first feature category and the feature categories included in the first target audio feature, where the second feature category is the first feature category not included in the first target audio feature;

[0018] Determine the first audio feature algorithm for calculating the second feature category;

[0019] Based on the first audio feature algorithm, determine the second target audio feature of the target audio.

[0020] In this implementation, when the at least one first feature category is not included in the first target audio feature, the intelligent device can determine the second feature category missing in the first target audio feature, and thus determine the corresponding second target audio feature through the first audio feature algorithm for determining the second feature category. In this way, the intelligent device does not need to calculate the first feature categories included in the first target audio feature, reducing the computing pressure of the intelligent device, and further reducing the delay in obtaining audio features, enabling the target function to obtain the required audio features in real time, and thus optimizing the user experience.

[0021] In some embodiments, the method further includes:

[0022] Send the second target audio feature, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature to the server, where the first source identifier is used to indicate that the second target audio feature is calculated by the first audio feature algorithm.

[0023] In this implementation, after calculating the audio feature, the intelligent device uploads the calculated audio feature to the server for storage in the audio feature library, so that the audio features stored in the audio feature library are more abundant, so that when the server does not upload audio features, other intelligent devices can still obtain the audio features of the target audio through the audio feature library.

[0024] In some embodiments, allocating the audio features in the first target audio feature that match the first feature category to the intelligent application includes:

[0025] Use the audio features in the first target audio feature that match the first feature category as execution parameters to generate an execution policy for implementing the target function, where the execution policy includes the execution parameters of components in the intelligent device;

[0026] Control the target component based on the execution policy to implement the target function with the execution parameters.

[0027] In this implementation manner, when the first target audio feature is received, the intelligent device controls the music function based on the audio feature of the target feature category, ensuring the accurate control of the music function by the intelligent device.

[0028] In some embodiments, the audio feature library has audio features of one or more audios, and each audio feature is associated with a source identifier, an audio identifier, and a feature category of the audio feature. The source identifier is a first source identifier or a second source identifier. The first source identifier is used to indicate that the audio feature is calculated by a first audio feature algorithm, and the second source identifier is used to indicate that the audio feature is calculated by a second audio feature algorithm. The algorithm performance of the second audio feature algorithm is higher than that of the first audio feature algorithm.

[0029] In this implementation manner, the audio feature library stores audio features from multiple sources, ensuring that the audio feature library can provide audio features for the intelligent device in a timely manner, thereby reducing latency and optimizing the user experience.

[0030] In a second aspect of the embodiments of the present application, a method for determining an audio feature is provided, which is applied to a server. The server is communicatively connected to the intelligent device. The method includes:

[0031] Construct an audio feature library, where the audio feature library includes audio features of one or more audios, and each audio feature has an associated audio identifier and a category identifier;

[0032] Receive a feature request message from the intelligent device. The feature request message includes an audio identifier of a target audio, and the audio identifier is used to represent the target audio currently played in the intelligent device;

[0033] Based on the feature request message, request to query whether the first target audio feature associated with the audio identifier exists in the audio feature library;

[0034] When the first target audio feature associated with the audio identifier is included in the audio feature library, send the first target audio feature to the intelligent device.

[0035] In some embodiments, when the first target audio feature associated with the audio identifier does not exist in the audio feature library, send a query response message to the intelligent device. The query response message is used to indicate that the first target audio feature associated with the audio identifier does not exist in the audio feature library.

[0036] In some embodiments, the audio feature library further includes a source identifier associated with each audio feature, where the source identifier is a first source identifier or a second source identifier. The first source identifier is used to indicate that the audio feature is calculated by a first audio feature algorithm, and the second source identifier is used to indicate that the audio feature is calculated by a second audio feature algorithm. Herein, the calculation performance of the second audio feature algorithm is higher than that of the first audio feature algorithm.

[0037] In some embodiments, the method further includes:

[0038] Receiving a third target audio feature sent by another device, the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature;

[0039] When the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, maintaining the audio features stored in the current audio feature library;

[0040] When the audio feature corresponding to the audio identifier is not stored in the audio feature library, storing the third target audio feature, the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature in the audio feature library.

[0041] In some embodiments, when the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the first source identifier, the method further includes:

[0042] Determining the feature category of the stored audio feature;

[0043] When the feature category of the stored audio feature is different from the feature category of the third target audio feature, associatively storing the stored audio feature and the third target audio feature;

[0044] When the feature category of the stored audio feature is the same as the feature category of the third target audio feature, maintaining the audio features stored in the current audio feature library.

[0045] In some embodiments, building the audio feature library includes:

[0046] Obtaining a plurality of audios;

[0047] Determining the audio features of any one of the audios through one or more second audio feature algorithms;

[0048] Store the audio features of one or more of the obtained audios and the audio identifiers associated with each audio feature in the audio feature library.

[0049] The third aspect of the embodiments of the present application provides a system for determining audio features. The system includes a server and one or more intelligent devices; an audio feature library is deployed in the server, and the audio feature library includes audio features of one or more audios.

[0050] The server is communicatively connected to one or more of the intelligent devices.

[0051] The intelligent device is used to implement the method according to any one of the first aspects of the embodiments of the present application.

[0052] The server is used to query in the audio feature library whether there is an audio feature that matches the audio identifier carried in the feature request message according to the feature request message; in the case of querying the first target audio feature that matches the audio identifier, send the first target audio feature to the intelligent device.

[0053] In some embodiments, the server is further used to receive the third target audio feature sent by other devices, as well as the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature.

[0054] In the case that the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, keep the audio features stored in the current audio feature library.

[0055] In the case that the audio feature corresponding to the audio identifier is not stored in the audio feature library, store the third target audio feature, as well as the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature in the audio feature library.

[0056] In some embodiments, in the case that the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the first source identifier, the server is further used to determine the feature category of the stored audio feature.

[0057] In the case that the feature category of the stored audio feature is different from the feature category of the third target audio feature, store the stored audio feature and the third target audio feature in an associated manner.

[0058] In the case that the feature category of the stored audio feature is the same as the feature category of the third target audio feature, keep the audio features stored in the current audio feature library.

[0059] A fourth aspect of the embodiments of the present application provides an intelligent device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for determining audio features as described above is implemented.

[0060] A fifth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method for determining audio features as described in the first aspect or any possible implementation manner of the first aspect.

[0061] A sixth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method for determining audio features as described in the second aspect or any possible implementation manner of the second aspect.

[0062] A seventh aspect of the embodiments of the present application provides a computer program product, which when running on an intelligent device, enables the intelligent device to implement the method for determining audio features as described in the first aspect or any possible implementation manner of the first aspect.

[0063] An eighth aspect of the embodiments of the present application provides a computer program product, which when running on a server, enables the intelligent device to implement the method for determining audio features as described in the second aspect or any possible implementation manner of the second aspect.

[0064] A ninth aspect of the embodiments of the present application provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for determining audio features as described in the second aspect or any possible implementation manner of the second aspect is implemented.

[0065] A tenth aspect of the embodiments of the present application provides a device for determining audio features, including:

[0066] A first determination module, configured to determine a first feature category, where the first feature category is a feature category required for an intelligent application to implement a target function, and the intelligent application is an application running in the intelligent device;

[0067] An acquisition module, configured to acquire an audio identifier of a target audio currently played in the intelligent device;

[0068] A request module, configured to send a feature request message to a server based on the audio identifier, where the feature request message requests the server to obtain a first target audio feature matching the audio identifier from an audio feature library;

[0069] A control module, configured to, upon receiving the first target audio feature, allocate the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application implements the target function based on the audio feature in the first target audio feature that matches the first feature category.

[0070] In some embodiments, the apparatus further includes:

[0071] A second determination module, configured to, when the first target audio feature does not include all audio features of the first feature category, determine a second target audio feature of the target audio based on a first audio feature algorithm, where the second target audio feature includes the audio features of the first feature category that are missing in the first target audio feature, and the first audio feature algorithm is an audio feature algorithm deployed on the intelligent device;

[0072] The control module is configured to allocate the second target audio feature to the intelligent application, so that the intelligent application implements the target function based on the second target audio feature.

[0073] In some embodiments, the second determination module is configured to determine a second feature category based on the first feature category and the feature categories included in the first target audio feature, where the second feature category is the first feature category not included in the first target audio feature; determine a first audio feature algorithm for calculating the second feature category; and determine a second target audio feature of the target audio based on the first audio feature algorithm.

[0074] In some embodiments, the apparatus further includes:

[0075] A sending module, configured to send the second target audio feature, a first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature to the server, where the first source identifier is used to indicate that the second target audio feature is calculated by the first audio feature algorithm.

[0076] In some embodiments, the control module is configured to use the audio feature in the first target audio feature that matches the first feature category as an execution parameter to generate an execution policy for implementing the target function, where the execution policy includes execution parameters of components in the intelligent device; and control a target component based on the execution policy to implement the target function with the execution parameters.

[0077] In some embodiments, the audio feature library has audio features of one or more audios. Each audio feature is associated with a source identifier, an audio identifier, and a feature category of the audio feature. The source identifier is a first source identifier or a second source identifier. The first source identifier is used to indicate that the audio feature is calculated by a first audio feature algorithm, and the second source identifier is used to indicate that the audio feature is calculated by a second audio feature algorithm. The algorithm performance of the second audio feature algorithm is higher than that of the first audio feature algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 FIG. shows a schematic diagram of an audio feature acquisition system involved in a method for determining audio features provided by an exemplary embodiment;

[0079] Figure 2 FIG. shows a schematic flowchart of a method for determining audio features provided by an exemplary embodiment;

[0080] Figure 3 FIG. shows a schematic flowchart of a method for determining audio features provided by an exemplary embodiment;

[0081] Figure 4 FIG. shows a schematic flowchart of a method for determining audio features provided by an exemplary embodiment;

[0082] Figure 5 FIG. shows a schematic flowchart of a method for determining audio features provided by an exemplary embodiment;

[0083] Figure 6 FIG. shows a schematic flowchart of a method for determining audio features provided by an exemplary embodiment;

[0084] Figure 7 FIG. shows a block diagram of the structure of a device for determining audio features provided by an exemplary embodiment;

[0085] Figure 8 is a schematic diagram of the structure of an intelligent device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] In order to make the technical problems, technical solutions, and beneficial effects to be solved by the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0087] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality of" means two or more, unless otherwise specifically defined.

[0088] As the degree of vehicle intelligence increases, various audio-based innovative applications in in-vehicle systems emerge in an endless stream. For example, action control applications of in-vehicle artificial intelligence robots, in-vehicle ambient light rhythm applications, in-vehicle microphone-free singing applications, etc.

[0089] The above innovative applications often require audio features of audio to formulate corresponding control strategies during operation. For example, in the in-vehicle ambient light rhythm application, the vehicle obtains the audio features of the music currently played by the in-vehicle multimedia device, formulates a corresponding light display strategy based on the audio features, and then controls the state of the ambient light based on the light display strategy.

[0090] In some embodiments, in order to obtain audio features, an audio processing algorithm is installed locally in the vehicle. When audio features need to be obtained, the audio features are obtained based on the audio processing algorithm installed locally in the vehicle. However, due to technical reasons such as the limitation of the algorithm performance locally in the vehicle and the inherent delay of the music processing algorithm itself, it is impossible to provide audio features for each in-vehicle application in real time, thus affecting the user experience.

[0091] This application provides an audio feature acquisition method, system, vehicle, and storage medium, constructs an audio feature library communicatively connected to an intelligent device, and stores audio features of multiple audios in the audio feature library. Among them, each audio feature is associated with its corresponding audio identifier and feature category. In this way, when an intelligent application needs to obtain the target audio feature of a music function, it only needs to query the category identifier of the feature category of the audio feature required by the intelligent application to implement the target function from the audio feature library, and feedback the target audio feature of the target audio corresponding to the queried category identifier to the intelligent device. Thus, the intelligent device can implement the target function of the intelligent application based on the target audio feature. In this way, the audio feature required for the target function can be obtained without calculating through the local audio feature algorithm, avoiding occupying the computing resources of the intelligent device. Moreover, the calculation process of the audio feature is reduced, the delay in obtaining the audio feature is reduced, so that the target function can obtain the required audio feature in real time, and thus the user experience is optimized.

[0092] The following describes this application with specific embodiments. Refer to Figure 1 , which shows a schematic diagram of a system for determining audio features involved in a method for determining audio features provided by an exemplary embodiment. Refer to Figure 1, the system includes: a server and one or more intelligent devices. An audio feature library is deployed in the server, and the audio feature library includes audio features of one or more audios. Among them, the server is communicatively connected to one or more intelligent devices respectively.

[0093] For example, the server can be communicatively connected to one or more intelligent devices through a bench local area network.

[0094] An audio feature library is deployed in the server, and the audio feature library includes audio features of one or more audios. Among them, the audio feature of a certain audio can be determined by the server based on the second audio feature algorithm in one or more second audio feature algorithms, or the audio feature of a certain audio can also be determined by the intelligent device based on the first audio feature algorithm in one or more first audio feature algorithms.

[0095] As an example, the computing performance of the first audio feature algorithm is lower than that of the second audio feature algorithm.

[0096] It can be understood that the first audio feature algorithms deployed in different intelligent devices may be different or the same, and the embodiments of the present application do not limit this.

[0097] The audio feature library can be stored in the server or can be an independently deployed database accessible by a server, and the embodiments of the present application do not limit this.

[0098] Optionally, the audio feature of each audio in the audio feature library has identification information, and the identification information is used to indicate which audio the audio feature corresponds to. Optionally, the identification information of any audio feature can include the audio identification of the audio, where the audio identification is determined by the identification of the audio associated with the audio feature. For example, the identification information of any audio feature can be the name of the audio, etc.

[0099] Optionally, the identification information of any audio feature can also include a source identifier, which is used to indicate the source of the audio feature, that is, which audio feature algorithm the audio feature is determined by. For example, the source identifier can include a first source identifier and a second source identifier, where the first source identifier is used to indicate that the audio feature is calculated by the first audio feature algorithm. The second source identifier is used to indicate that the audio feature is calculated by the second audio feature algorithm.

[0100] The server can be a server with audio feature recognition function. The server can communicate with multiple intelligent devices and the audio feature library through a bench local area network constructed by the vehicle enterprise developer.

[0101] In some embodiments, one or more second audio feature algorithms are provided in the server. Different second audio feature algorithms are used to determine different categories of audio features in the audio. For example, the one or more second audio feature algorithms may include an audio beat extraction algorithm, a chord extraction algorithm, an emotion recognition algorithm, a music genre recognition algorithm, a sound quality enhancement algorithm, a voice accompaniment separation algorithm, and the like.

[0102] It should be noted that both the number and type of the second audio feature algorithms can be set as needed, and no specific limitations are imposed in the embodiments of the present application. In practical applications, more or fewer second audio feature algorithms can be provided in the server.

[0103] Among them, the server is used to obtain multiple audio; determine the audio features of any audio through one or more audio feature algorithms; and store the audio features of the obtained one or more audio and the audio identifiers associated with each audio feature in the audio feature library.

[0104] In some embodiments, the server can obtain multiple audio by obtaining an audio list.

[0105] One way is: the audio list can be an audio list uploaded by the user through a terminal device such as a mobile terminal or a vehicle-mounted terminal. Correspondingly, the user accesses the server through a mobile terminal or a vehicle-mounted terminal and uploads the audio list. The audio list includes information of one or more audio.

[0106] Another way is: the audio list can also be an audio list obtained by the server through an associated application. For example, the audio list can be an audio list generated according to various lists in the application, or the audio list can be an audio list created by the user in the application, etc. In the embodiments of the present application, no specific limitations are imposed on the way of obtaining the audio list.

[0107] It should be noted that the server can also communicate with the server of the application used to generate the audio list. When obtaining the audio list of the associated application, the server sends a list acquisition request to the server of other applications. After receiving the list acquisition request sent by the server, if the list is a list created by the user, the server of the application sends a prompt message to the user, and the prompt message is used to ask the user whether to agree to share the audio list. If the operation of sharing the audio list triggered by the user is received, the server of the application feeds back the audio list.

[0108] The server further includes an audio playback module, which is used to play the audio in the audio list. Thus, the server sequentially identifies the audio features of the currently playing audio through a second audio feature algorithm, and caches the identified audio features in a preset format. Among them, the preset format can be set as needed. In the embodiments of the present application, the format of the audio features is not specifically limited.

[0109] After the server calculates the audio features of a certain audio, the server can also set the source identifier of the audio features as a second source identifier, and then the server stores the audio features marked with the source identifier in the audio feature library.

[0110] Specifically, the server can set the attribute information of the audio features, and the attribute information includes a second source identifier; the second source identifier is used to identify that the audio features are calculated by the second audio feature algorithm.

[0111] Among them, the position and form of the second source identifier in the audio features can be set as needed. In the embodiments of the present application, the position and form of the second source identifier in the audio features are not specifically limited. For example, the second source identifier can be set to represent the source of the audio features by different numerical values. Among them, the numerical values of the source identifiers corresponding to different sources can be set as needed. In the embodiments of the present application, the numerical values of the source identifiers are not specifically limited. For example, when the audio features are the audio features determined by the second audio feature algorithm in the server, the second source identifier can be set to 1, and when the audio features are the audio features determined by the intelligent device, the first source identifier of the audio features can be set to 0.

[0112] The audio features may have more other sources. In the embodiments of the present application, this is not specifically limited.

[0113] Another point to note is that since the server can obtain multiple audio lists, there may be duplicate audios in the multiple audio lists. For duplicate audios, the server can repeatedly determine the audio features of the audio and upload them to the audio feature library; the server can also record the audio identifiers of the audios for which the audio features have been determined. When determining the audio features based on the audio list, based on the stored audio identifiers, it is determined whether the audio in the audio list has been determined for audio features. If the audio identifier of the audio is the stored audio identifier, the audio features of the audio are not repeatedly determined, thereby reducing the calculation of duplicate audio features and saving computing resources.

[0114] The audio identifiers stored in the server can be updated periodically. That is, every once in a while, the server can delete the stored audio identifiers to re-determine the audio characteristics of the audio. Among them, the update period can be set as needed. In the embodiments of this application, no specific limitation is imposed on the update period. For example, the update period can be one week, one month, etc.

[0115] The intelligent device can be an intelligent driving device (such as a vehicle), an in-vehicle terminal, a mobile terminal, a wearable device, etc. In the embodiments of this application, no specific limitation is imposed on the intelligent device.

[0116] One or more application programs are installed in the intelligent device. Among them, there is a certain application program among the one or more application programs that needs to use the audio characteristics of the audio during operation.

[0117] For example, taking the intelligent device as a vehicle, the one or more application programs include: an in-vehicle atmosphere light control application. For another example, applications such as the swaying together application of the in-vehicle artificial intelligence robot, the rhythm application of the in-vehicle atmosphere light, and the in-vehicle karaoke without a microphone.

[0118] When the user turns on the function that requires the opening of relevant audio characteristics in the vehicle, the vehicle can control the parameters of the components in the intelligent device based on the audio characteristics required by the application to control the components.

[0119] It can be understood that the categories of audio characteristics required by different application programs in the vehicle during operation may be different. For example, the swaying together application of the in-vehicle artificial intelligence robot requires music characteristics such as beats, music styles, structures, and emotions, the rhythm application of the in-vehicle atmosphere light requires music characteristics such as beats and rhythms, and applications such as the in-vehicle karaoke without a microphone require music characteristics such as real-time pitch correction and vocal separation.

[0120] Specifically, when the user turns on the function that requires the opening of relevant audio characteristics for the atmosphere light control application in the vehicle, the vehicle controls the atmosphere light in the vehicle through the atmosphere light control application in combination with the audio characteristics of the currently playing target audio.

[0121] It should be noted that the target audio is the audio played in the intelligent device. Among them, the intelligent device can play the target audio through the application for playing music in the intelligent device.

[0122] In the case of detecting the audio feature requirements of an intelligent application running in an intelligent device, the intelligent device determines the feature category of the audio features required when the intelligent application implements the target function, and obtains the audio identifier of the target audio currently played in the intelligent device. Wherein, the audio identifier is used to identify the target audio, and the audio identifier is an identifier that can uniquely represent the target audio. In some embodiments, the audio identifier may be an identifier that combines relevant information of multiple audios such as the audio name, album name, and artist name. In the embodiments of the present application, the content included in the audio identifier and the form of the audio identifier are not specifically limited.

[0123] In some embodiments, the intelligent terminal determines the target function that the intelligent application currently needs to implement; based on the target function, determines the feature category of the audio features required to implement the target function. Each function of the intelligent application corresponds to at least one feature category, and the feature categories corresponding to different functions of the intelligent application may be the same, different, or partially the same. In the embodiments of the present application, no specific limitation is made thereto.

[0124] The intelligent application may be an application for implementing one or more specific functions. When the intelligent application is an application for implementing a specific function, when the intelligent device detects the startup of the application, it determines that the intelligent application triggers an audio feature requirement, where the specific function is the target function. When the intelligent application is an application for implementing multiple specific functions, when the intelligent device determines that a certain specific function is currently started in the intelligent application, it determines that the intelligent application triggers an audio feature requirement and determines the specific function as the target function.

[0125] It should be noted that the audio feature requirement may also be triggered by the target audio currently played in the intelligent device. For example, in a state where the target function has been enabled, when the intelligent device detects a change in the target audio currently played, it determines that the intelligent application triggers an audio feature requirement.

[0126] It should be noted that another point is that the application for playing audio in the intelligent device may be the intelligent application or other applications with audio playback functions. In the embodiments of the present application, no specific limitation is made thereto. When the audio playback function is an application with other audio playback functions, the intelligent application may obtain the currently played audio in the application through an interface or other communication methods.

[0127] After the intelligent device obtains the audio identifier of the target audio and the feature category of the required audio features, based on the feature category and the audio identifier, it sends a feature request message to the server. Correspondingly, the server receives the feature request message and queries whether the target audio corresponding to the audio identifier has the audio features of the feature category from the audio feature library based on the audio identifier and the feature category in the feature request message. Among them, the audio feature library includes one or more audio features, each audio feature has an associated audio identifier and feature category, and the audio identifier associated with any audio feature is the audio identifier of the audio corresponding to the audio feature.

[0128] In some embodiments, the intelligent device generates a feature request message based on the audio identifier and sends the feature request message carrying the audio identifier to the server. Correspondingly, after the server receives the feature request message, it parses the feature request message to obtain the audio identifier of the target audio; it queries the first target audio feature matching the audio identifier from the audio feature library and sends the first target audio feature to the intelligent device. If the server queries the first target audio feature matching the audio identifier from the audio feature library, it sends the first target audio feature to the intelligent device.

[0129] Among them, the feature category is a category divided based on the characteristics or attributes of the audio described by the audio features. The categories of audio features required by different applications may be different. For example, for the in-vehicle ambient light rhythm application, the feature category of the required audio features may include audio beat features, etc. For the in-vehicle microphone-free singing application, the feature category of the required audio features may include the accompaniment features of the audio after the voice and accompaniment are separated, etc. Among them, different feature categories can be distinguished by different category identifiers. Among them, different feature categories can be represented by category identifiers with different flag values.

[0130] In some embodiments, the audio feature library stores multiple audio features, and each audio feature stores an associated audio identifier. Correspondingly, when the server receives the feature request message, it parses the feature request message to obtain the audio identifier carried by the feature request message, traverses the audio features stored in the audio feature library based on the audio identifier, and if it queries the first target audio feature matching the audio identifier, it sends the first target audio feature to the intelligent device.

[0131] It should be noted that the audio feature library can also be queried in combination with the required feature categories. Accordingly, multiple audio features are stored in the audio feature library, and each audio feature is correspondingly stored with an audio identifier and a feature category. When the server receives a feature request message, it parses the feature request message to obtain the audio identifier and feature category carried in the feature request message, traverses the audio features stored in the audio feature library based on the audio identifier. If the audio feature corresponding to the audio identifier is stored in the audio feature library, the server continues to determine whether the first target audio feature corresponding to the feature category is included in the audio feature corresponding to the audio identifier.

[0132] In the above process, if the server queries the first target audio feature, the server sends the first target audio feature to the intelligent device. If the first target audio feature does not exist in the audio feature library, the server does not respond to the feature request message, or the server sends a query response message to the intelligent device, and the query response message is used to indicate that there is no matching audio feature in the audio feature library.

[0133] Among them, the situation where the first target audio feature does not exist in the audio feature library includes that the audio feature library does not include the audio feature corresponding to the audio identifier carried in the feature request message, or the audio feature library includes the audio feature corresponding to the audio identifier carried in the feature request message, but the audio feature corresponding to the audio identifier does not include the audio feature corresponding to the category identifier.

[0134] The intelligent device is also used to, in the case of querying the first audio feature from the audio feature library, implement the application requirements corresponding to the target function based on the first audio feature. Accordingly, when the intelligent device receives the first target audio feature, it allocates the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application implements the target function based on the audio feature that matches the first feature category. Among them, the intelligent device uses the audio feature in the first target audio feature that matches the first feature category as an execution parameter to generate an execution policy for implementing the target function, and the execution policy includes the execution parameters of the components in the intelligent device; controls the target component to implement the target function with the execution parameter based on the execution policy.

[0135] When the intelligent device does not receive the first target audio feature sent by the audio feature library within a preset duration, or when the first target audio feature library received by the intelligent device does not contain all the first feature categories required to implement the target function, or when the intelligent device receives a query response message sent by the audio feature library, the intelligent device determines the required audio feature category through a locally deployed first audio feature algorithm. Among them, the preset duration can be set as needed. In the embodiments of the present application, no specific limitation is imposed on the preset duration. For example, the preset duration can be 100 milliseconds, 50 milliseconds, etc. In the embodiments of the present application, for the sake of distinction, the audio feature of a certain audio obtained from the server can be called the first target audio feature. The audio feature of a certain audio calculated locally by the intelligent device using the first audio feature algorithm is called the second target audio feature. Of course, it can be understood that the first target audio feature of a certain audio may be calculated by the intelligent device or may be calculated by the server. The embodiments of the present application do not make any limitation in this regard.

[0136] It should be noted that in order to ensure that the audio feature can meet the application requirements, after receiving the audio feature sent by the server, the intelligent device also combines the feature category of the audio feature of the first target audio and the first feature category required by the target function to determine whether to implement the corresponding target function through the first target audio feature.

[0137] When the first target audio feature does not contain all the audio features of the first feature category, the second target audio feature of the target audio is determined based on the first audio feature algorithm. The second target audio feature contains the audio features of the first feature category missing in the first target audio feature. The first audio feature algorithm is an audio feature algorithm deployed locally; the second target audio feature is assigned to the intelligent application so that the intelligent application can implement the target function based on the audio feature matching the first feature category.

[0138] Among them, the fact that the first target audio feature does not contain all the audio features of the first feature category means that the first target audio feature contains some audio features of the first feature category, or the first target audio feature does not contain any audio features of the first feature category.

[0139] It should be noted that the audio feature library has the audio features of one or more audios. Each audio feature is associated with a source identifier, an audio identifier, and the feature category of the audio feature. The source identifier is the first source identifier or the second source identifier. The first source identifier is used to identify that the audio feature is calculated by the first audio feature algorithm, and the second source identifier is used to identify that the audio feature is calculated by the second audio feature algorithm. The algorithm performance of the second audio feature algorithm is higher than that of the first audio feature algorithm.

[0140] In some embodiments, the intelligent device sends the second target audio feature, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature to the server, where the first source identifier is used to indicate that the second target audio feature is calculated by the first audio feature algorithm.

[0141] Correspondingly, the server is further configured to receive the second target audio feature sent by the intelligent device, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature.

[0142] When the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, the audio features stored in the current audio feature library are maintained; when the audio feature corresponding to the audio identifier is not stored in the audio feature library, the second target audio feature, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature are stored in the audio feature library.

[0143] The intelligent device sets the relevant information of the second target audio feature, where the relevant information includes the first source identifier; the first source identifier is used to identify that the audio feature is calculated by the first audio feature algorithm; the relevant information of the second target audio feature and the second target audio feature are sent, where the relevant information of the second target audio feature includes the first source identifier, the audio identifier of the target audio, and the feature category of the second target audio feature, and the relevant information of the second target audio feature is used to determine whether to add the second target audio feature to the audio feature library.

[0144] Wherein, the position and form of the first source identifier in the audio feature can be set as needed. In the embodiments of the present application, the position and form of the first source identifier in the audio feature are not specifically limited. For example, the first source identifier can be set to represent the source of the audio feature by different values. Among them, the values of the source identifiers corresponding to different sources can be set as needed. In the embodiments of the present application, the values of the source identifiers are not specifically limited. For example, when the audio feature is the audio feature determined by the second audio feature algorithm in the server, the second source identifier can be set to 1, and when the audio feature is the audio feature determined by the intelligent device, the first source identifier of the audio feature can be set to 0.

[0145] The server is also used to receive the third target audio feature sent by the other device, as well as the first source identifier, audio identifier, and feature category of the third target audio feature associated therewith; when there is an audio feature corresponding to the audio identifier stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, maintain the audio features stored in the current audio feature library; when there is no audio feature corresponding to the audio identifier stored in the audio feature library, store the third target audio feature, the first source identifier, audio identifier, and feature category of the third target audio feature associated therewith in the audio feature library.

[0146] In practical applications, the audio features received by the server may be calculated by the server through the second audio feature algorithm or by the intelligent device through the first audio feature algorithm. To prevent duplicate storage of audio features, after receiving the audio features, the server can also determine whether to update the audio features in the audio feature library based on the source of the audio features and the audio features stored in the current audio feature library.

[0147] Among them, the other device may be an intelligent device or a server, etc. Correspondingly, the third target audio feature may be calculated by the server through the second audio feature algorithm or by the intelligent device through the first audio feature algorithm. For example, the server may determine the received second target audio feature as the third target audio feature.

[0148] After the server receives the third target audio feature uploaded by the other device, determine the audio identifier corresponding to the third target audio feature, and query whether there is an audio feature corresponding to the audio identifier stored in the audio feature library. If there is no audio feature corresponding to the audio identifier in the audio feature library, store the third target audio feature and the relevant information of the third target audio feature correspondingly.

[0149] For the case where there is an audio feature corresponding to the audio identifier stored in the audio feature library, the intelligent device can determine whether to update the audio features stored in the audio feature library by combining the stored audio features, the source, feature category, and other information of the audio features.

[0150] Correspondingly, the audio feature library also has a source identifier associated with each audio feature, and the source identifier is the first source identifier or the second source identifier. The first source identifier is used to identify that the audio feature comes from the intelligent device, and the second source identifier is used to identify that the audio feature comes from the server.

[0151] In some embodiments, the server is further configured to receive a third target audio feature sent by another device, as well as a source identifier, an audio identifier, and a feature category of the third target audio feature associated with the third target audio feature; when the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, keep the audio features stored in the current audio feature library; when the audio feature corresponding to the audio identifier is not stored in the audio feature library, store the third target audio feature, the source identifier associated with the third target audio feature, the audio identifier, and the feature category of the third target audio feature in the audio feature library.

[0152] It should be noted that the server can also receive audio features calculated by a second audio feature algorithm deployed in the server. When the server receives a third target audio feature from the server, the third target audio feature is associated with a second audio identifier and corresponds to the first source identifier; after the server parses the source identifier of the third audio feature, store the third target audio feature and the relevant information of the third target audio feature correspondingly. Among them, when the audio feature corresponding to the audio identifier has been stored in the server, the audio feature corresponding to the second audio identifier is overwritten by the third target audio feature. The following will illustrate this in combination with several specific situations.

[0153] Since the algorithm performance of the second audio feature algorithm deployed in the server is higher than that of the first audio feature algorithm deployed in other intelligent devices, the audio features uploaded by the server are more comprehensive and accurate. Therefore, when receiving audio features determined by the second audio feature algorithm deployed in the server, the newly received audio features and audio identifiers can be directly stored in the audio feature library correspondingly.

[0154] It should be noted that when the audio feature corresponding to the audio identifier has been stored in the audio feature library, the original audio feature is overwritten by the third audio feature corresponding to the newly received audio identifier to prevent the storage of the same audio feature.

[0155] In some other embodiments, when the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the first source identifier, the server is further configured to determine the feature category of the stored audio feature; when the feature category of the stored audio feature is different from the feature category of the third target audio feature, store the stored audio feature and the third target audio feature associatively; when the feature category of the stored audio feature is the same as the feature category of the third target audio feature, keep the audio features stored in the current audio feature library.

[0156] When the audio feature library has stored the audio features uploaded by the server, when the server receives the audio features calculated by the first audio feature algorithm in the intelligent device, no other processing is performed on the currently stored audio features. In some embodiments, when the audio feature library has stored the audio features uploaded by the server, when the server receives the audio features calculated by the first audio feature algorithm in the intelligent device, the server also generates and sends an error message to prompt the staff to maintain the server.

[0157] When both the audio features received by the server and the stored audio features are the audio features uploaded by the intelligent device, the server can directly store the audio features uploaded by multiple intelligent devices. In some other embodiments, the server can also compare whether the two audio features are of the same feature category. If the audio features received by the server and the stored audio features are of the same feature category, the currently stored audio features are maintained; if the feature categories of the audio features received by the server and the stored audio features are different, the relevant information of the audio features received by the server and the stored audio features is stored correspondingly. In this way, the audio feature library can obtain different audio features from multiple devices, enabling the audio feature library to store rich audio features, so that the intelligent device can obtain audio features, thereby reducing the audio feature calculation process, reducing the delay in obtaining audio features, enabling the application program to obtain the required audio features in real time, and thus optimizing the user experience.

[0158] In the embodiments of the present application, an audio feature library communicatively connected to the intelligent device is constructed. The audio feature library stores the audio features of multiple audios, where each audio feature is associated with its corresponding audio identifier. In this way, when the intelligent application needs to obtain the target audio features of the music function, it only needs to query the audio features of the target audio required by the intelligent application to achieve the target function from the audio feature library, and feedback the first target audio features of the queried target audio to the intelligent device. After receiving the first target audio features from the cloud, the intelligent device determines whether the first target audio features contain the audio features of the required first feature category, and allocates the audio features in the first target audio features that match the first feature category to the intelligent application, so that the intelligent application can implement the target function based on the audio features that match the first feature category. In this way, the audio features required for the target function can be obtained without calculating through the local audio feature algorithm, avoiding occupying the computing resources of the intelligent device. Moreover, the audio feature calculation process is reduced, and the delay in obtaining audio features is reduced, enabling the target function to obtain the required audio features in real time, and thus optimizing the user experience.

[0159] The present application will be described below in conjunction with a specific method flow. Refer to Figure 2 which shows a schematic flowchart of a method for determining audio features provided by an exemplary embodiment.

[0160] S201, the intelligent device detects whether it has received the audio feature requirement of the intelligent application.

[0161] The intelligent application can be an application for implementing one or more specific functions. When the intelligent application is an application for implementing a specific function, when the intelligent device detects the startup of the application, it determines that the intelligent application has triggered the audio feature requirement, where the specific function is the target function. When the intelligent application is an application for implementing multiple specific functions, when the intelligent device determines that a certain specific function is currently started in the intelligent application, it determines that the intelligent application has triggered the audio feature requirement.

[0162] It should be noted that the audio feature requirement can also be triggered by the target audio currently played in the intelligent device. For example, in a state where the target function has been enabled, when the intelligent device detects a change in the target audio currently played, it determines that the intelligent application has triggered the audio feature requirement.

[0163] It can be understood that there is one or more applications in the intelligent device, where the target function can refer to the function that requires the use of the audio features of the audio in multiple applications. For example, the target function is a function that can control the intelligent device or components of other associated devices based on audio features in combination with the intelligent device. For example, the target function can be an in-vehicle ambient light control application.

[0164] It can be understood that the target function can include a first function and a second function. The first function is a function that requires the use of the audio features of the audio, and the second function is a function that does not require the use of the audio features of the audio. Accordingly, the intelligent device can determine the required audio features based on the different functions started by the target function. As an example, the audio feature requirement can be the requirement triggered when the first function of the music function is enabled.

[0165] S202, in response to the audio feature requirement of the intelligent application running in the intelligent device, the intelligent device obtains the audio identifier of the target audio currently played in the intelligent device.

[0166] The audio identifier is used to represent the audio features of the target audio required by the target function, and is used to query whether the audio features associated with the audio identifier exist in the audio feature library. The target audio can be any type of audio. For example, audio such as songs and poetry recitations. Accordingly, the audio identifier is used to identify the target audio, and the audio identifier is an identifier that can uniquely represent the target audio. In some embodiments, the audio identifier can be an identifier that combines relevant information of multiple audios such as the audio name, album name, and artist name. In the embodiments of the present application, the content included in the audio identifier and the form of the audio identifier are not specifically limited.

[0167] When the beat features and rhythm features of the audio named "We All Have a Home" are required, the audio identifier can be "We All Have a Home", so that the server can know to obtain the audio features such as the beat and rhythm of the song "We All Have a Home".

[0168] In the case where the audio identifier is the song name of a certain song, since the smart device does not indicate the feature category of the specific audio features required, when the server receives the audio identifier, it will query whether the audio features of the song "We All Have a Home" exist in the music feature library, and then send all the audio features of the retrieved song "We All Have a Home" to the smart device.

[0169] In some embodiments, the smart device obtains the audio identifier corresponding to the audio features required for the target function, and this audio identifier is the audio identifier of the audio corresponding to the audio features. In other embodiments, the smart device can obtain the audio features of the upcoming played audio in advance. Correspondingly, the smart device obtains the playlist in the application for playing the audio, and based on this playlist, obtains the audio identifiers of multiple audios in the list.

[0170] It should be noted that this audio identifier can be the identifier of the target audio played in the smart terminal. The smart device can play the target audio through the application player installed in the smart device, or the target audio can also be the audio played in the terminal communicating with the smart device.

[0171] S203, the smart device sends a feature request message to the server based on this audio identifier.

[0172] This feature request message is used to request the server to query in the audio feature library whether the target audio has audio features matching this audio identifier.

[0173] Correspondingly, the smart device generates a feature request message based on this audio identifier and sends this feature request message to the server. For example, when the beat features and rhythm features of the audio named "We All Have a Home" are required, the audio identifier can be "We All Have a Home". Correspondingly, the generated feature request message includes "We All Have a Home", so that the server can know to obtain the audio features of the song "We All Have a Home".

[0174] S204, the server receives the feature request message from this smart device.

[0175] The server has an audio feature library, or the server can access this audio feature library, and this audio feature library is deployed in another server. The audio feature library has the audio features of one or more audios. Each of these audio features has an associated audio identifier.

[0176] S205. When the first target audio feature matching the target audio associated with the audio identifier is included in the audio feature library, the server sends the first target audio feature to the intelligent device.

[0177] It can be understood that when the server receives the feature request message, the server can query in the audio feature library to determine whether there is a first target audio feature matching the audio identifier in the audio feature library. If there is, the first target audio feature is sent to the intelligent device. If there is no first target audio feature, the server does not respond to the feature request message, or sends a query response message to the intelligent device to indicate that the first target audio feature does not exist.

[0178] In some embodiments, the intelligent device can also send the required first feature category to the server. Correspondingly, the server queries the audio feature in combination with the audio identifier and the first feature category. When the audio feature associated with the audio identifier is included in the audio feature library and the feature category of the audio feature in the audio feature library is different from the first feature category, the server does not respond to the feature request message, or sends a query response message to the intelligent device, and the query response message is used to indicate that the first target audio feature does not exist in the audio feature library.

[0179] Multiple audio features are stored in the audio feature library, and each audio feature is correspondingly stored with an audio identifier. After the server receives the feature request message, it parses the feature request message to obtain the audio identifier carried by the feature request message, and queries the audio features stored in the audio feature library one by one based on the audio identifier. If the audio feature corresponding to the audio identifier is stored in the audio feature library, continue to query whether there is an audio feature corresponding to the required feature category in the audio feature.

[0180] S206. When the first target audio feature is received, the first target audio feature is allocated to the intelligent application so that the intelligent application can implement the target function based on the first target audio feature.

[0181] Specifically, after the intelligent device receives the first target audio feature, it uses the first target audio feature as an execution parameter to generate an execution policy for implementing the target function. The execution policy includes the execution parameters of the components in the intelligent device; based on the execution policy, it controls the target component to implement the target function with the execution parameter.

[0182] For example, the intelligent device is a vehicle-mounted terminal, and the target function is the vehicle-mounted ambient light control function. Correspondingly, the vehicle-mounted terminal, through the ambient light control function, combines the beat characteristics of the target audio currently being played, determines the brightness parameter of the in-vehicle lights based on the beat characteristics, generates an execution policy based on the brightness parameter to control the brightness of the in-vehicle ambient lights, so as to achieve the control of the in-vehicle ambient lights to flash with the audio.

[0183] In a possible embodiment of the present application, since the audio feature library may or may not have the audio features associated with the audio identifier. In the case where the audio features associated with the audio identifier are not available, the method provided by the embodiments of the present application may further include:

[0184] In the case where the first target audio feature is not received, determine the second target audio feature of the target audio based on the first audio feature algorithm; allocate the second target audio feature to the intelligent application, so that the intelligent application implements the target function based on the second target audio feature.

[0185] When the intelligent device does not receive the audio feature sent by the server within a preset duration, or when the intelligent device receives a query response message including the first indication information sent by the server, it is determined that there is no audio feature in the audio feature library that matches the feature category of the target audio. Wherein, the preset duration can be set as needed, and in the embodiments of the present application, the preset duration is not specifically limited. For example, the preset duration can be 100 milliseconds, 50 milliseconds, etc.

[0186] When the intelligent device determines that there is no audio feature of the currently required target audio in the audio feature library, determine the audio feature of the target audio through the first audio feature algorithm configured by the intelligent device, and control the music function based on the audio feature of the target audio.

[0187] It should be noted that the audio features stored in the audio feature library may include the audio features determined by the intelligent device through the first audio feature algorithm, may also include the audio features determined by other intelligent devices through the first audio feature algorithm, or may include the audio features determined by the server through the second audio feature algorithm. Among them, the algorithm performance (such as computing power) of the second audio feature algorithm is higher than that of the first audio feature algorithm.

[0188] Therefore, since the audio features required during the operation of the same application in different intelligent devices may be different. For example, for vehicle A, the user can set that the in-vehicle ambient light rhythm application in vehicle A requires music features such as beats and rhythms during operation. For vehicle B, the user may set that the in-vehicle ambient light rhythm application in vehicle B requires beat music features during operation. Or the audio features required by different applications in the same vehicle for the same audio are different, or the audio features required by different applications in different vehicles for the same audio are different. For the in-vehicle ambient light rhythm application, the feature categories of the required audio features may include audio beat features, etc. For the in-vehicle microphone-free singing application, the feature categories of the required audio features may include the accompaniment features of the audio after the voice accompaniment is separated, etc.

[0189] In some embodiments, the server determines the audio features only based on the audio identifier in the feature request message. However, there is a problem that the feature category of the audio features of a certain audio stored in the audio feature library is different from the target audio category currently required by the intelligent device. To ensure that the received audio features can meet the application requirements, after receiving the audio features associated with the audio identifier, the intelligent device can also combine the feature category of the audio features associated with the audio identifier and the feature category of the required audio features to determine whether to implement the corresponding target function through the audio features associated with the audio identifier.

[0190] Correspondingly, the intelligent device determines the feature category of the audio features from the server. If the feature category of the audio features from the server includes the required feature category, the target function is implemented based on the audio features belonging to the required feature category in the audio features. In other words, when the audio features from the server include the required audio features, the target function is implemented based on the required audio features.

[0191] For example, if the feature category of the audio features associated with the audio identifier stored in the audio feature library includes music features such as beats and rhythms, and the target feature category required by vehicle B is beat music features, since the feature category of the audio features associated with the audio identifier already includes the feature category of beat music features, the intelligent device can combine the audio features associated with the audio identifier to meet the corresponding application requirements.

[0192] Specifically, the intelligent device can combine the beat music features in the audio features associated with the audio identifier to meet the corresponding application requirements.

[0193] In a possible implementation of the present application, if the audio feature from the server does not include the audio feature of the target feature category, the intelligent device determines the audio feature of the audio corresponding to the target feature category based on the first audio feature algorithm; the target feature category is the feature category of the audio feature required by the music function; the target function is implemented based on the audio feature of the audio corresponding to the target feature category.

[0194] In a possible implementation of the present application, after the intelligent device determines the audio feature of the audio based on the first audio feature algorithm, the intelligent device can also upload the audio feature to the audio feature library.

[0195] For example, the intelligent device sends the second target audio feature, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature to the server, where the first source identifier is used to indicate that the second target audio feature is calculated by the first audio feature algorithm. After receiving the audio feature and the first source identifier associated with the audio feature, the server can determine whether to add the audio feature to the audio feature library. Alternatively, the intelligent device can send the audio feature, the first source identifier associated with the audio feature, and the feature category to a device that deploys or has a management audio feature library, so that the device that deploys or has a management audio feature library determines whether to add the audio feature to the audio feature library.

[0196] Specifically, the server can add the audio feature to the audio feature library when it determines that the audio feature associated with the audio identifier is not stored in the audio feature library according to the audio identifier.

[0197] Among them, the other device can be an intelligent device or a server, etc. Correspondingly, the third target audio feature can be calculated by the server through the second audio feature algorithm or by the intelligent device through the first audio feature algorithm. When the server receives the third target audio feature uploaded by other devices, it determines the audio identifier corresponding to the third target audio feature and queries whether the audio feature corresponding to the audio identifier is stored in the audio feature library. If the audio feature corresponding to the audio identifier does not exist in the audio feature library, the third target audio feature and the relevant information of the third target audio feature are stored correspondingly.

[0198] For the case where the audio feature corresponding to the audio identifier is already stored in the audio feature library, the intelligent device can determine whether to update the audio feature stored in the audio feature library by combining the stored audio feature, the source of the audio feature, the feature category, and other information.

[0199] Accordingly, the audio feature library also has a source identifier associated with each of the audio features. The source identifier is the first source identifier or the second source identifier. The first source identifier is used to identify that the audio feature comes from the intelligent device, and the second source identifier is used to identify that the audio feature comes from the server.

[0200] In some embodiments, the server is further configured to receive a third target audio feature sent by another device, as well as the source identifier, audio identifier, and feature category of the third target audio feature associated therewith; when the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, keep the audio features stored in the current audio feature library; when the audio feature corresponding to the audio identifier is not stored in the audio feature library, store the third target audio feature, the source identifier, audio identifier, and feature category of the third target audio feature associated therewith in the audio feature library.

[0201] It should be noted that the server can also receive audio features calculated by a second audio feature algorithm deployed in the server. When the server receives a third target audio feature from the server, the third target audio feature is associated with a second audio identifier and corresponds to the first source identifier; after the server parses the source identifier of the third audio feature, store the third target audio feature and the relevant information of the third target audio feature accordingly. Among them, when the audio feature corresponding to the audio identifier has been stored in the server, the audio feature corresponding to the second audio identifier is overwritten by the third target audio feature. The following describes this in combination with several specific situations.

[0202] Since the algorithm performance of the second audio feature algorithm deployed in the server is higher than that of the first audio feature algorithm deployed in other intelligent devices, the audio features uploaded by the server are more comprehensive and accurate. Therefore, when the received audio feature is an audio feature determined by the second audio feature algorithm deployed in the server, the newly received audio feature and audio identifier can be directly stored in the audio feature library in correspondence.

[0203] It should be noted that when the audio feature corresponding to the audio identifier has been stored in the audio feature library, the original audio feature is overwritten by the third audio feature corresponding to the newly received audio identifier to prevent the storage of the same audio feature.

[0204] In some other embodiments, the audio features corresponding to the audio identifier are stored in the audio feature library, and the source identifier of the stored audio features is the first source identifier. The server is further configured to determine the feature category of the stored audio features; in a case where the feature category of the stored audio features is different from the feature category of the third target audio features, the stored audio features and the third target audio features are stored in an associated manner; in a case where the feature category of the stored audio features is the same as the feature category of the third target audio features, the audio features stored in the current audio feature library are maintained.

[0205] When the audio features uploaded by the server are already stored in the audio feature library and the server then receives the audio features calculated by the first audio feature algorithm in the intelligent device, no other processing is performed on the currently stored audio features. In some embodiments, when the audio features uploaded by the server are already stored in the audio feature library and the server then receives the audio features calculated by the first audio feature algorithm in the intelligent device, the server also generates and sends an error message to prompt the staff to perform maintenance on the server.

[0206] When both the audio features received by the server and the stored audio features are the audio features uploaded by the intelligent device, the server can directly store the audio features uploaded by multiple intelligent devices. In some other embodiments, the server can also compare whether the two audio features are of the same feature category. If the audio features received by the server and the stored audio features are of the same feature category, the currently stored audio features are maintained; if the feature categories of the audio features received by the server and the stored audio features are different, the relevant information of the audio features received by the server and the stored audio features is stored correspondingly. In this way, the audio feature library can obtain different audio features from multiple devices, enabling the audio feature library to store rich audio features, so that the intelligent device can obtain audio features, thereby reducing the calculation process of audio features, reducing the delay in obtaining audio features, enabling the application program to obtain the required audio features in real time, and thus optimizing the user experience.

[0207] Specifically, the intelligent device sets the attribute information of the audio features, and the attribute information includes the first source identifier; the first source identifier is used to identify that the audio features are calculated by the first audio feature algorithm.

[0208] Among them, the position and form of the first source identifier in the audio feature can be set as needed. In the embodiments of the present application, the position and form of the first source identifier in the audio feature are not specifically limited. For example, the first source identifier can be set to represent the source of the audio feature by different numerical values. Among them, the numerical values of the source identifiers corresponding to different sources can be set as needed. In the embodiments of the present application, the numerical values of the source identifiers are not specifically limited. For example, when the audio feature is the audio feature determined by the second audio feature algorithm in the server, the second source identifier can be set to 1, and when the audio feature is the audio feature determined by the intelligent device, the first source identifier of the audio feature can be set to 0. In a possible implementation manner of the present application, after the intelligent device determines the target feature category required for the music function, the target feature category can also be carried when sending the feature request message, which is convenient for the server to query the target feature corresponding to the target feature category and send it to the intelligent device.

[0209] In another possible implementation manner of the present application, when the server receives the feature request message, if it determines that it does not have the audio feature associated with the audio identifier, or the feature category of the audio feature associated with the audio identifier is not the target feature category, or does not include the target feature category, the server can also determine the audio feature of the audio based on the second audio feature algorithm and then send it to the intelligent device.

[0210] It should be noted that for any intelligent device, when sending the audio feature of a certain audio to the server, the audio feature category corresponding to the audio feature and the identification information of the intelligent device can also be sent to the server, which is convenient for the server to determine which intelligent device specifically calculated the audio feature.

[0211] As an example, if the server receives the audio features of audio 1 sent by vehicle 1 and vehicle 2 respectively, and the audio features sent by vehicle 1 are different from those sent by vehicle 2. For example, the audio feature sent by vehicle 1 is audio feature 1, and the audio feature sent by vehicle 2 is audio feature 2. Both the audio feature 1 and the audio feature 2 have the first source identifier. Since the two audio features are different, the server can store the audio feature 1 and the audio feature 2 corresponding to the audio in the audio feature library, as well as the information of the intelligent devices corresponding to the audio feature 1 and the audio feature 2 respectively. In other words, there are multiple audio features corresponding to a certain audio in the audio feature library.

[0212] To prevent duplicate storage of audio features, after the server receives the audio features sent by a certain intelligent device, it can also determine whether to update the audio features in the audio feature library according to the source of the audio features and the audio features stored in the current audio feature library.

[0213] It should be noted that when the audio feature library has stored the audio features corresponding to the audio identifier, the third target audio features corresponding to the newly received audio identifier are used to overwrite the original audio features to prevent the storage of the same audio features.

[0214] When the audio feature library has stored the audio features uploaded by the server, when the server receives the audio features calculated by the first audio feature algorithm in the intelligent device, no other processing is performed on the currently stored audio features. In some embodiments, when the audio feature library has stored the audio features uploaded by the server, when the server receives the audio features calculated by the first audio feature algorithm in the intelligent device, the server also generates and sends an error message to prompt the staff to maintain the server.

[0215] When both the audio features received by the server and the stored audio features are the audio features uploaded by the intelligent device, the server can directly store the audio features uploaded by multiple intelligent devices. In other embodiments, the server can also compare whether the two audio features are of the same feature category. If the audio features received by the server and the stored audio features are of the same feature category, the currently stored audio features are maintained; if the audio features received by the server and the stored audio features are of different feature categories, the relevant information of the audio features received by the server and the stored audio features is stored correspondingly. In this way, the audio feature library can obtain different audio features from multiple devices, enabling the audio feature library to store rich audio features, so that the intelligent device can obtain audio features, thereby reducing the audio feature calculation process, reducing the delay in obtaining audio features, enabling the application to obtain the required audio features in real time, and thus optimizing the user experience.

[0216] In the embodiments of the present application, an audio feature library communicatively connected to the intelligent device is constructed. The audio feature library stores the audio features of multiple audios. Among them, each audio feature is associated with its corresponding audio identifier and feature category. In this way, when the intelligent application needs to obtain the target audio features of the music function, it only needs to query the category identifier of the feature category of the audio features required for the intelligent application to implement the target function from the audio feature library, and feedback the target audio features of the target audio corresponding to the queried category identifier to the intelligent device. Thus, the intelligent device can implement the target function of the intelligent application based on the target audio features. In this way, the audio features required for the target function can be obtained without going through the local audio feature algorithm calculation, avoiding occupying the computing resources of the intelligent device. Moreover, the audio feature calculation process is reduced, and the delay in obtaining audio features is reduced, enabling the target function to obtain the required audio features in real time, and thus optimizing the user experience.

[0217] It should be noted that the categories of audio features required for different target functions may be different. Therefore, after receiving the first target audio feature, it is also possible to determine whether the required feature category is included in the first target audio feature library according to the feature category required by the target function. To ensure that the intelligent device can obtain all the audio features of the required first feature categories to ensure the normal operation of the target function, after receiving the first target audio feature fed back by the server, the intelligent device can further check the first target audio feature to ensure that all the audio features of the required first feature categories of the target audio are received. Refer to Figure 3 , which shows a schematic flowchart of an audio feature acquisition method provided by an exemplary embodiment.

[0218] S301. The intelligent device determines a first feature category, which is the feature category required for the intelligent application to implement the target function, and the intelligent application is an application running on the intelligent device.

[0219] Among them, the feature category is a category divided based on the characteristics or attributes of the audio described by the audio feature. The categories of audio features required by different applications may be different. For example, for the in-vehicle ambient light rhythm application, the feature categories of the required audio features may include audio beat features, etc. For the in-vehicle microphone-free singing application, the feature categories of the required audio features may include the accompaniment features of the audio after the voice and accompaniment are separated, etc.

[0220] The feature category refers to the category of audio features required to implement the target function. In some embodiments, the intelligent terminal determines the target function that the intelligent application currently needs to implement; based on the target function, it determines the feature category of the audio features required to implement the target function. Among them, each function of the intelligent application corresponds to at least one feature category, and the feature categories corresponding to different functions of the intelligent application may be the same, different, or partially the same. In the embodiments of the present application, no specific limitation is made on this.

[0221] As an example, the feature category can be represented by a category identifier. For example, the in-vehicle ambient light rhythm application requires audio features such as beats and rhythms. If the feature category identifier for beats is 01 and the feature category identifier for rhythms is 00.

[0222] When the beat feature and rhythm feature of the audio with the name "We All Have a Home" are required, then the audio identifier can be "We All Have a Home", and the category identifier can be 01 + 00. In this way, the server can know to obtain the audio features such as the beats and rhythms of the song "We All Have a Home".

[0223] In the case where the audio identifier is the song name of a certain song, since the smart device does not indicate the feature category of the specific audio features required, for the server, after receiving the audio identifier, it will query whether the music feature library has the audio features of the song "We All Have a Home", and then send all the audio features of the song "We All Have a Home" retrieved to the smart device.

[0224] In some embodiments, the smart device obtains the audio identifier corresponding to the audio features required for the target function, and this audio identifier is the audio identifier of the audio corresponding to the audio features. In other embodiments, the smart device can obtain in advance the audio features of the audio to be played. Correspondingly, the smart device obtains the playlist in the application for playing the audio, and based on this playlist, obtains the audio identifiers of multiple audios in the list.

[0225] It should be noted that this audio identifier can be the identifier of the target audio played in the smart terminal. The smart device can play this target audio through the application player installed in the smart device, or this target audio can also be the audio played in the terminal communicating with the smart device.

[0226] S302, the smart device obtains the audio identifier of the target audio currently played in this smart device.

[0227] The principle of this step is the same as that of step S202, and will not be elaborated here.

[0228] S303, when the smart device receives the first target audio features, it allocates the audio features in the first target audio features that match the first feature category to this smart application, so that this smart application realizes the target function based on the audio features in the first target audio features that match the first feature category.

[0229] When the smart device receives the first target audio features, it judges whether the feature category included in the first target audio features includes the required first feature category. If it includes, it allocates the audio features in the first target audio features that match the first feature category to this smart application, so that this smart application realizes the target function based on the audio features that match the first feature category.

[0230] Among them, the smart device uses the audio features in the first target audio features that match the first feature category as execution parameters to generate an execution policy for realizing the target function. This execution policy includes the execution parameters of the components in the smart device; based on this execution policy, it controls the target component to realize the target function with these execution parameters, so that the principle that this smart application realizes the target function based on the first target audio features is the same as that of step S203, and will not be elaborated here.

[0231] In addition, the intelligent device can determine whether the audio feature contains all the first feature categories based on the source identifier of the first target audio feature. Correspondingly, when receiving the first target audio feature, the source identifier of the first target audio feature is determined. If the source identifier of the first target audio feature indicates that the first target audio feature is an audio feature calculated by the second audio feature algorithm, it is determined that the first audio feature contains the at least one first feature category. If the source identifier of the first target audio feature indicates that the first target audio feature is an audio feature calculated by the first audio feature algorithm, the at least one second feature category included in the first audio feature is determined. If the at least one second feature category is the same as the at least one first feature category, it is determined that the first audio feature contains the at least one first feature category. Among them, the algorithm performance of the second audio feature algorithm is higher than that of the first audio feature algorithm.

[0232] In this implementation manner, it is determined whether the first target audio feature contains the required first feature category according to the source identifier of the received first target audio feature. Since the algorithm performance of the second audio feature algorithm is high and the calculated audio features are relatively rich, when the source identifier of the first target audio feature indicates that the first target audio feature is calculated by the second audio feature algorithm, it can be determined that the first target audio feature contains the at least one first feature category. When the source identifier of the first target audio feature indicates that the first target audio feature is calculated by the first audio feature algorithm, since the algorithm performance of the first audio feature algorithm is low, only a single audio feature can be calculated, and it cannot be guaranteed that the first target audio feature contains the at least one first feature category. Therefore, the second feature category included in the first target audio feature is determined, and whether the first target audio feature includes the at least one first feature category is determined by comparing the second feature category and the first feature category. In this way, the source identifier of the first target audio feature is determined first and then the feature categories included in the first target audio feature are compared, which ensures the accurate judgment of the audio feature categories and improves the efficiency of judging the audio feature categories.

[0233] In some embodiments, the first target audio feature may include audio features of some required first feature categories. At this time, the intelligent device may determine the feature categories that need to be calculated according to the feature categories existing in the first target audio feature, and then calculate the audio features corresponding to the remaining feature categories. Correspondingly, when the first target audio feature does not include all the audio features of the first feature categories, the second target audio feature of the target audio is determined based on the first audio feature algorithm. The second target audio feature includes the audio features of the first feature categories missing from the first target audio feature. The first audio feature algorithm is an audio feature algorithm deployed on the intelligent device; the second target audio feature is allocated to the intelligent application so that the intelligent application can implement the target function based on the second target audio feature.

[0234] In this implementation, when the feature categories included in the first target audio feature do not include all the first feature categories, the audio features corresponding to the target feature categories required for the music function are calculated, and then the music function is controlled based on the audio features. Thus, the intelligent device can calculate the audio features required for the music function, ensuring that when the audio feature library does not store audio features that meet the application requirements, the intelligent device can still normally implement the application requirements and ensure the normal operation of the music function.

[0235] Among them, the intelligent device may select the audio feature algorithm corresponding to the missing feature category to determine the corresponding human feature category. Correspondingly, the intelligent device determines a second feature category based on the first feature category and the feature categories included in the first target audio feature. The second feature category is the first feature category not included in the first target audio feature; determines the first audio feature algorithm for calculating the second feature category; and determines the second target audio feature of the target audio based on the first audio feature algorithm.

[0236] In this implementation, when the first target audio feature does not include the at least one first feature category, the intelligent device may determine the second feature category missing from the first target audio feature, and then determine the corresponding second target audio feature through the first audio feature algorithm for determining the second feature category. In this way, the intelligent device does not need to calculate the first feature categories included in the first target audio feature, reducing the calculation pressure of the intelligent device, and further reducing the delay in obtaining audio features, enabling the target function to obtain the required audio features in real time, and thus optimizing the user experience.

[0237] In an embodiment of the present application, an audio feature library communicatively connected to an intelligent device is constructed. The audio feature library stores audio features of multiple audios, where each audio feature is associated with its corresponding audio identifier. In this way, when an intelligent application needs to obtain the target audio feature of the music function, it only needs to query the audio feature of the target audio required by the intelligent application to implement the target function from the audio feature library, and feedback the first target audio feature of the queried target audio to the intelligent device. After receiving the first target audio feature from the cloud, the intelligent device determines whether the first target audio feature contains the audio feature of the required first feature category, and allocates the audio feature matching the first feature category in the first target audio feature to the intelligent application, so that the intelligent application can implement the target function based on the audio feature matching the first feature category. In this way, the audio feature required for the target function can be obtained without calculating through the local audio feature algorithm, avoiding occupying the computing resources of the intelligent device. Moreover, the calculation process of the audio feature is reduced, and the delay in obtaining the audio feature is reduced, enabling the target function to obtain the required audio feature in real time, thereby optimizing the user experience.

[0238] To further illustrate the application scenario of the present application, the present application will be described below in conjunction with the interaction between an intelligent device, a server, and an audio feature library. Refer to Figure 4 , which shows a schematic flowchart of an audio feature acquisition method provided by an exemplary embodiment.

[0239] S401, the server obtains a plurality of audio lists.

[0240] In some embodiments, the server can obtain multiple audios by obtaining an audio list. Among them, the audio list can be an audio list uploaded by a user. Correspondingly, the server can also communicate with the users registered in the server. The user logs in to the server through a mobile terminal or a vehicle-mounted terminal, etc., and uploads the audio list to be recognized. The audio list can also be an audio list obtained by the server through an associated application. For example, the audio list can be an audio list generated according to various lists in the application, or the audio list can be an audio list created by the user in the application, etc. In the embodiment of the present application, the acquisition method of the audio list is not specifically limited.

[0241] It should be noted that this server can also communicate with the server of the application used to generate the audio list. When obtaining the audio list of the associated application, the server sends a list acquisition request to the servers of other applications. After receiving the list acquisition request sent by the server, if the list is a list created by the user, the server of the other application sends a prompt message to the user, and this prompt message is used to ask the user whether to agree to share this audio list. If the operation of sharing the same audio list triggered by the user is received, the server of the application feeds back this audio list.

[0242] S402. The server determines the audio features of the multiple audios based on the second audio feature algorithm.

[0243] This server further includes an audio playback module, and this audio playback module is used to play the audios in the audio list. Thus, the server sequentially identifies the audio features of the currently playing audio through the second audio feature algorithm, and caches the identified audio features in a preset format. Among them, this preset format can be set as needed. In the embodiments of the present application, the format of the audio features is not specifically limited.

[0244] It should be noted that another point is that since the audio lists obtained by the server can be multiple, there may be duplicate audios in these multiple audio lists. For the duplicate audios, the server can repeatedly determine the audio features of this audio and upload them to the audio feature library; the server can also record the audio identifiers of the audios whose audio features have been determined. When determining the audio features based on the audio list, based on the stored audio identifiers, it is determined whether the audio in the audio list has had its audio features determined. If the audio identifier of the audio is the stored audio identifier, the audio features of this audio are not repeatedly determined, thereby reducing the calculation of duplicate audio features and saving computing resources.

[0245] The audio identifiers stored in the server can be updated periodically, that is, every once in a while, the server can delete the stored audio identifiers so as to re-determine the audio features of the audios. Among them, this update period can be set as needed. In the embodiments of the present application, the update period is not specifically limited. For example, this update period can be 1 week, 1 month, etc.

[0246] S403. The server uploads the audio features to the audio feature library.

[0247] Please refer to Figure 5 , after the server obtains the audio features, based on the source identifier marked for the audio features by this server, it uploads the audio features marked with the source identifier to the audio feature library, and the audio feature library receives the audio features uploaded by the server and stores the audio identifier and the audio features correspondingly.

[0248] Accordingly, set the attribute information of the audio feature, where the attribute information includes a second source identifier; the second source identifier is used to identify that the audio feature is from the server; send the attribute information of the audio feature and the audio feature, and the related information of the audio feature includes the second source identifier and the audio identifier of the audio corresponding to the audio feature, and the related information of the audio feature is used to determine whether to add the audio feature to the audio feature library.

[0249] Among them, the position and form of the second source identifier in the audio feature can be set as needed. In the embodiments of the present application, the position and form of the second source identifier in the audio feature are not specifically limited. For example, the second source identifier can be set to represent the source of the audio feature by different numerical values. Among them, the numerical values of the source identifiers corresponding to different sources can be set as needed. In the embodiments of the present application, the numerical values of the source identifiers are not specifically limited. For example, when the audio feature is the audio feature determined by the second audio feature algorithm in the server, the second source identifier can be set to 1, and when the audio feature is the audio feature determined by the intelligent device, the first source identifier of the audio feature can be set to 0.

[0250] The audio feature can also have more other sources. In the embodiments of the present application, this is not specifically limited.

[0251] S404. In response to the audio feature requirement of the intelligent application running in the intelligent device, the intelligent device determines the feature category of the audio feature required when the intelligent application implements the target function, and obtains the audio identifier of the target audio currently played in the intelligent device.

[0252] S405. Based on the audio identifier, the intelligent device sends a feature request message to the server.

[0253] S406. According to the feature request message, the server queries in the audio feature library whether there is an audio feature matching the target audio.

[0254] In some embodiments, the server queries whether there is an audio feature matching the audio identifier in the audio feature library. In some embodiments, the feature request message sent by the intelligent device further includes the required first feature category. Accordingly, after the server queries the first audio feature corresponding to the audio identifier, it determines whether the feature category of the first audio feature includes the first feature category carried in the feature request message. If the feature category of the first audio feature includes the first feature category carried in the feature request message, the server sends the first audio feature to the intelligent device. If the feature category of the first audio feature does not include the feature category carried in the feature request message, the server does not send the first audio feature to the intelligent device.

[0255] After the server queries the first target audio feature corresponding to the audio identifier, it sends the first target audio feature to the intelligent device, and the intelligent device executes step S407.

[0256] It should be noted that when the audio feature library does not have the audio feature corresponding to the audio identifier, or when the feature category of the audio feature corresponding to the audio identifier stored in the audio feature library does not include the feature category corresponding to the audio identifier, the server may not send a message to the intelligent device; or the server may send a query response message to the intelligent device, and the query response message is used to indicate that the audio feature corresponding to the audio identifier does not exist in the audio feature library. Correspondingly, when the intelligent device does not receive the first audio feature sent by the server within a preset duration, or when the intelligent device receives the query response message sent by the server, it is determined that the first audio feature of the audio does not exist in the audio feature library, and step S408 is executed. Among them, the preset duration can be set as needed, and in the embodiments of the present application, the preset duration is not specifically limited. For example, the preset duration can be 100 milliseconds, 50 milliseconds, etc.

[0257] S407. In the case of receiving the first target audio feature, the intelligent device allocates the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application realizes the target function based on the audio feature that matches the first feature category.

[0258] S408. In the case of not receiving the first target audio feature, or when the first target audio feature does not include all the audio features of the first feature category, the intelligent device determines the second target audio feature of the target audio based on the first audio feature algorithm.

[0259] The first audio feature algorithm is an audio feature algorithm deployed in the intelligent device. Due to the limitation of the computing resources of the intelligent device, the algorithm performance of the first audio feature algorithm is weak. In the intelligent device, only the first audio feature algorithm for separately calculating the audio feature of a certain feature category is deployed, and the number of the first audio feature algorithms in the intelligent device can be one or more. In the embodiments of the present application, this is not specifically limited.

[0260] In this step, the intelligent device calls the corresponding first audio feature based on the feature category corresponding to the target audio, and determines the second audio feature of the target audio based on the first audio feature.

[0261] S409. The intelligent device allocates the second target audio feature to the intelligent application, so that the intelligent application realizes the target function based on the second target audio feature.

[0262] S410. The intelligent device sends the second target audio feature, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature to the server.

[0263] Please continue to refer to Figure 5 , after the intelligent device determines the second target audio feature based on the locally deployed first audio feature algorithm, it uploads the second target audio feature to the server. Among them, the intelligent device can label the source identifier and feature category of the audio feature based on the first audio feature algorithm, and upload the audio feature labeled with the source identifier and feature category to the audio feature library.

[0264] S411. The server receives the second target audio feature from the intelligent device.

[0265] S412. The server receives the third target audio feature sent by other devices, as well as the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature.

[0266] The third target audio feature is the audio feature uploaded by other devices. Among them, the other device can be an intelligent device or a server. Correspondingly, the server can determine the received second target audio feature as the third target audio feature.

[0267] S413. When the audio feature corresponding to the audio identifier is not stored in the audio feature library, the server stores the third target audio feature, the source identifier associated with the third target audio feature, the audio identifier, and the feature category of the third target audio feature in the audio feature library.

[0268] When it is determined that the second audio feature is not stored in the audio feature library according to the first audio identifier, or when the source identifier of the audio feature corresponding to the first audio identifier stored in the audio feature library is the second source identifier, the server adds the first audio feature to the audio feature library.

[0269] S414. When the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, the server keeps the audio feature stored in the current audio feature library.

[0270] Refer to Figure 6, when the audio feature library has stored the audio features corresponding to the second source identifier, when the server receives the audio features again, it does not perform other processing on the currently stored audio features. In some embodiments, when the audio feature library has stored the audio features uploaded by the server, when the server receives the audio features calculated by the first audio feature algorithm in the intelligent device, the server also generates and sends an error message to prompt the staff to maintain the server.

[0271] It should be noted that, referring to Figure 6 , the server can also receive the audio features calculated by the second audio feature algorithm deployed in the server. When the server receives the third target audio features from the server; after the server parses the source identifier of the third audio feature, it stores the relevant information of the third audio feature correspondingly. Among them, if the audio features corresponding to the audio identifier associated with the third target audio feature have been stored in the server, the original audio features are overwritten by the third target audio features.

[0272] In the audio feature library, the audio features corresponding to the audio identifier are stored, and the source identifier of the stored audio features is the first source identifier, and the server determines the feature category of the stored audio features.

[0273] Referring to Figure 6 , when the audio features received by the server and the stored audio features are both the audio features uploaded by the intelligent device, the server can directly store the audio features uploaded by multiple intelligent devices. In other embodiments, in the case where the feature category of the stored audio features is different from the feature category of the third target audio features, the stored audio features and the third target audio features are stored associatively; in the case where the feature category of the stored audio features is the same as the feature category of the third target audio features, the audio features stored in the current audio feature library are maintained.

[0274] Correspondingly, the server can also compare whether two audio features are of the same feature category. If the audio features received by the server and the stored audio features are of the same feature category, the currently stored audio features are maintained; if the feature categories of the audio features received by the server and the stored audio features are different, the relevant information of the audio features received by the server and the stored audio features is stored correspondingly. In this way, the audio feature library can obtain different audio features from multiple devices, enabling the audio feature library to store rich audio features, so that the intelligent device can obtain audio features, thereby reducing the calculation process of audio features, reducing the delay in obtaining audio features, enabling the application program to obtain the required audio features in real time, and thus optimizing the user experience.

[0275] In the embodiments of the present application, an audio feature library communicatively connected to an intelligent device is constructed. The audio feature library stores audio features of multiple audios, where each audio feature is associated with its corresponding audio identifier. In this way, when an intelligent application needs to obtain the target audio feature of the music function, it only needs to query the audio feature of the target audio required by the intelligent application to implement the target function from the audio feature library, and feedback the first target audio feature of the queried target audio to the intelligent device. After receiving the first target audio feature from the cloud, the intelligent device determines whether the first target audio feature contains the audio feature of the required first feature category, and allocates the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application can implement the target function based on the audio feature that matches the first feature category. In this way, the audio feature required for the target function can be obtained without calculating through the local audio feature algorithm, avoiding occupying the computing resources of the intelligent device. Moreover, the calculation process of the audio feature is reduced, and the delay in obtaining the audio feature is reduced, enabling the target function to obtain the required audio feature in real time, thereby optimizing the user experience.

[0276] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0277] See Figure 7 , which shows a schematic structural diagram of a device for determining audio features provided by the present application. Each module included is used to execute each step in the above embodiments. See Figure 7 , the device for determining audio features includes:

[0278] A first determination module 701, configured to determine a first feature category, where the first feature category is the feature category required for an intelligent application to implement a target function, and obtain the audio identifier of the target audio currently played in the intelligent device, where the intelligent application is an application running in the intelligent device;

[0279] A request module 702, configured to send a feature request message to the server based on the audio identifier, where the feature request message requests the server to obtain the first target audio feature that matches the audio identifier in the audio feature library;

[0280] A control module 703, configured to, in the case of receiving the first target audio feature, allocate the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application can implement the target function based on the audio feature that matches the first feature category.

[0281] In some embodiments, the device further includes:

[0282] A second determination module, configured to, when the first target audio feature does not include all audio features of the first feature category, determine a second target audio feature of the target audio based on a first audio feature algorithm, where the second target audio feature includes audio features of the first feature category that are missing in the first target audio feature, and the first audio feature algorithm is an audio feature algorithm deployed locally;

[0283] The control module 703 is configured to allocate the second target audio feature to the intelligent application, so that the intelligent application implements the target function based on the audio feature matching the first feature category.

[0284] In some embodiments, the second determination module is configured to determine a second feature category based on the first feature category and the feature categories included in the first target audio feature, where the second feature category is the first feature category that is not included in the first target audio feature; determine a first audio feature algorithm for calculating the second feature category; and determine a second target audio feature of the target audio based on the first audio feature algorithm.

[0285] In some embodiments, the apparatus further includes:

[0286] A sending module, configured to send the second target audio feature, a first source identifier associated with the second target audio feature, an audio identifier, and a feature category of the second target audio feature to the server, where the first source identifier is used to indicate that the second target audio feature is calculated by the first audio feature algorithm.

[0287] In some embodiments, the control module 703 is configured to use the audio feature matching the first feature category in the first target audio feature as an execution parameter to generate an execution policy for implementing the target function, where the execution policy includes execution parameters of components in the intelligent device; and control a target component to implement the target function with the execution parameter based on the execution policy.

[0288] In some embodiments, the audio feature library has audio features of one or more audios, and each audio feature is associated with a source identifier, an audio identifier, and a feature category of the audio feature. The source identifier is a first source identifier or a second source identifier. The first source identifier is used to identify that the audio feature is calculated by the first audio feature algorithm, and the second source identifier is used to identify that the audio feature is calculated by the second audio feature algorithm. The algorithm performance of the second audio feature algorithm is higher than that of the first audio feature algorithm.

[0289] In an embodiment of the present application, an audio feature library communicatively connected to an intelligent device is constructed. The audio feature library stores audio features of multiple audios, where each audio feature is associated with its corresponding audio identifier. In this way, when an intelligent application needs to obtain the target audio feature of the music function, it only needs to query the audio feature of the target audio required by the intelligent application to implement the target function from the audio feature library, and feedback the first target audio feature of the queried target audio to the intelligent device. After receiving the first target audio feature from the cloud, the intelligent device determines whether the first target audio feature contains the audio features of the required first feature category, and allocates the audio features in the first target audio feature that match the first feature category to the intelligent application, so that the intelligent application can implement the target function based on the audio features that match the first feature category. In this way, the audio features required for the target function can be obtained without calculating through the local audio feature algorithm, avoiding occupying the computing resources of the intelligent device. Moreover, the calculation process of the audio features is reduced, and the delay in obtaining the audio features is reduced, enabling the target function to obtain the required audio features in real time, thereby optimizing the user experience.

[0290] Figure 8 is a schematic diagram of an intelligent device provided by an exemplary embodiment of the present application. As Figure 8 shown, the intelligent device 8 of this embodiment includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80, such as a program for determining audio features. When the processor 80 executes the computer program 82, the steps in the method embodiments for determining various audio features described above are implemented, such as Figure 2 the steps S201 to S203 shown.

[0291] Exemplarily, the computer program 82 can be divided into one or more units. The one or more units are stored in the memory 81 and executed by the processor 80 to complete the present application. The one or more units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 82 in the intelligent device 8.

[0292] The intelligent device 8 can be any intelligent device with a control function. The intelligent device 8 may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art can understand that Figure 8 merely an example of the intelligent device 8, which does not constitute a limitation on the intelligent device 8. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the intelligent device 8 may further include input / output devices, network access devices, buses, etc.

[0293] The so-called processor 80 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0294] The memory 81 may be an internal storage unit of the intelligent device 8, such as the hard disk or memory of the intelligent device 8. The memory 81 may also be an external storage device of the intelligent device 8, such as a plug-in hard disk equipped on the intelligent device 8, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 81 may also include both the internal storage unit and the external storage device of the intelligent device 8. The memory 81 is used to store the computer program and other programs and data required by the terminal device. The memory 81 may also be used to temporarily store data that has been output or is to be output.

[0295] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0296] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0297] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0298] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0299] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0300] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0301] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0302] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps in the above-described method embodiments are implemented.

[0303] An embodiment of this application also provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can be made to execute the steps in the above-described method embodiments.

[0304] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A method for determining audio features, characterized in that, The method includes: Determine a first feature category, where the first feature category is a feature category required for an intelligent application to implement a target function, and the intelligent application is an application running in the intelligent device; Obtain the audio identifier of the target audio currently played in the intelligent device; Based on the audio identifier, send a feature request message to the server, where the feature request message requests the server to obtain the first target audio feature matching the audio identifier in the audio feature library; In the case of receiving the first target audio feature, allocate the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application implements the target function based on the audio feature in the first target audio feature that matches the first feature category.

2. The method according to claim 1, characterized in that The method further includes: When the first target audio feature does not include all the audio features of the first feature category, determine the second target audio feature of the target audio based on a first audio feature algorithm, where the second target audio feature includes the audio features of the first feature category missing in the first target audio feature, and the first audio feature algorithm is an audio feature algorithm deployed in the intelligent device; Allocate the second target audio feature to the intelligent application, so that the intelligent application implements the target function based on the second target audio feature.

3. The method according to claim 2, characterized in that The determining the second target audio feature of the target audio based on the first audio feature algorithm includes: Based on the first feature category and the feature categories included in the first target audio feature, determine a second feature category, where the second feature category is the first feature category not included in the first target audio feature; Determine the first audio feature algorithm for calculating the second feature category; Based on the first audio feature algorithm, determine the second target audio feature of the target audio.

4. The method according to claim 2, characterized in that, The method further includes: Send the second target audio feature, the first source identifier associated with the second target audio feature, the audio identifier, and the feature category of the second target audio feature to the server, where the first source identifier is used to indicate that the second target audio feature is calculated by the first audio feature algorithm.

5. The method according to any one of claims 1-4, characterized in that, The allocating the audio feature in the first target audio feature that matches the first feature category to the intelligent application includes: Use the audio feature in the first target audio feature that matches the first feature category as an execution parameter to generate an execution policy for implementing the target function, where the execution policy includes the execution parameters of components in the intelligent device; Based on the execution policy, control the target component to implement the target function with the execution parameter.

6. The method according to any one of claims 1 to 4, characterized in that The audio feature library has audio features of one or more audios. Each audio feature is associated with a source identifier, an audio identifier, and a feature category of the audio feature. The source identifier is a first source identifier or a second source identifier. The first source identifier is used to indicate that the audio feature is calculated by a first audio feature algorithm, and the second source identifier is used to indicate that the audio feature is calculated by a second audio feature algorithm. The algorithm performance of the second audio feature algorithm is higher than that of the first audio feature algorithm.

7. A system for determining audio features, characterized in that, The system includes a server and one or more intelligent devices; the audio feature library is deployed in the server, and the audio feature library includes audio features of one or more audios; The server is communicatively connected to one or more of the intelligent devices; The intelligent device is used to implement the method according to any one of claims 1-6; The server is used to query in the audio feature library whether there is an audio feature that matches the audio identifier carried in the feature request message according to the feature request message; When a first target audio feature that matches the audio identifier is found, the server sends the first target audio feature to the intelligent device.

8. The system according to claim 7, wherein The server is further used to receive a third target audio feature sent by another device, as well as the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature; When the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the second source identifier, the audio features stored in the current audio feature library are maintained; When the audio feature corresponding to the audio identifier is not stored in the audio feature library, the third target audio feature, as well as the source identifier, audio identifier, and feature category of the third target audio feature associated with the third target audio feature, are stored in the audio feature library.

9. The system according to claim 8, wherein When the audio feature corresponding to the audio identifier is stored in the audio feature library and the source identifier of the stored audio feature is the first source identifier, the server is further used to determine the feature category of the stored audio feature; When the feature category of the stored audio feature is different from the feature category of the third target audio feature, the stored audio feature and the third target audio feature are stored in an associated manner; When the feature category of the stored audio feature is the same as the feature category of the third target audio feature, the audio features stored in the current audio feature library are maintained.

10. An intelligent device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for determining audio features according to any one of claims 1-6.

11. A computer program product, characterized in that, When the computer program product runs on an intelligent device, the intelligent device is caused to execute the method for determining audio features according to any one of claims 1-6.

12. An apparatus for determining audio features, characterized in that, The device includes: A first determination module, configured to determine a first feature category, where the first feature category is a feature category required for the intelligent application to implement a target function, and the intelligent application is an application running in the intelligent device; An acquisition module, configured to acquire an audio identifier of a target audio currently played in the intelligent device; A request module, configured to send a feature request message to a server based on the audio identifier, where the feature request message requests the server to acquire a first target audio feature matching the audio identifier in an audio feature library; A control module, configured to, when receiving the first target audio feature, allocate the audio feature in the first target audio feature that matches the first feature category to the intelligent application, so that the intelligent application implements the target function based on the audio feature in the first target audio feature that matches the first feature category.