Equipment control method and device, electrical equipment and medium

By building a pre-set fusion voiceprint library in electrical appliances, analyzing users' voiceprint feature codes and attributes, and matching user profiles to control device operation, the problem of personalized service for electrical appliances when user habits are not recorded is solved, thus improving the user experience.

CN121053976APending Publication Date: 2025-12-02WUHU MATY AIR CONDITIONING EQUIP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410703656.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing electrical appliances cannot provide personalized services without recording user habits, resulting in a poor user experience.

Method used

By building a pre-set fusion voiceprint library across multiple devices within the target area, analyzing users' voiceprint feature codes and attributes, and matching user profiles to control device operation, personalized services can be achieved.

Benefits of technology

It improves the user experience by enabling personalized services for users through voiceprint attribute matching, thus meeting user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053976A_ABST
    Figure CN121053976A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment control method and device, electrical equipment and a medium, and the method comprises the steps: obtaining a target voiceprint feature code of a target audio if the target equipment detects the target audio; if the target voiceprint feature code is not included in voiceprint feature codes in a preset fusion voiceprint library, matching voiceprint attributes of the target audio with N groups of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result, the preset fusion voiceprint library comprising N voiceprint feature codes corresponding to at least one device, and the N groups of voiceprint attributes of the preset fusion voiceprint library comprising N voiceprint feature codes corresponding to at least one device; n is a positive integer and a voiceprint attribute corresponding to each voiceprint feature code; and if the matching result shows that the target voiceprint attribute matched with the voiceprint attribute of the target audio exists in the preset fusion voiceprint library, controlling the target equipment to operate based on a user portrait corresponding to the target voiceprint attribute. According to the scheme, the personalized service of the target equipment to the user is realized, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio processing, and particularly relates to a device control method, apparatus, electrical equipment, and medium. Background Technology

[0002] With the continuous development of science and technology, electrical appliances are becoming increasingly intelligent, and many appliances (such as air conditioners and televisions) can support voice control. To improve the user experience, if an appliance records the user's usage habits, it can provide services based on those habits when the user controls the device via voice. However, if the device does not record the user's usage habits, it cannot provide personalized services that meet the user's needs, resulting in a poor user experience. Summary of the Invention

[0003] In view of the above-mentioned technical problems in related technologies, embodiments of the present invention provide a device control method, apparatus, electrical equipment and medium to realize personalized services for each user and improve user experience.

[0004] In a first aspect, embodiments of the present invention provide a device control method, comprising:

[0005] If the target device detects the target audio, the target voiceprint feature code fed back by the voice service provider of the target device after processing the target audio is obtained. The target device is located in the target area, and the target area contains S devices, where S is a positive integer.

[0006] If the target voiceprint feature code is not included in the voiceprint feature code in the preset fusion voiceprint library, the voiceprint attributes of the target audio are matched with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result. The preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer.

[0007] If the matching result indicates that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio, the target device is controlled to operate based on the user profile corresponding to the target voiceprint attribute.

[0008] In some embodiments, the preset fused voiceprint library is stored in a voiceprint fusion device, which is communicatively connected to the S devices. The preset fused voiceprint library is constructed through the following steps:

[0009] Listen to the audio processing results of each of the S devices. The audio processing results include the voiceprint feature code and voiceprint attributes of the audio to be processed. The voiceprint feature code of the audio to be processed in the audio processing results is fed back by the voice service provider of the corresponding device when processing the audio to be processed.

[0010] The similarity of voiceprint attributes in multiple sets of audio processing results is calculated, and the voiceprint feature codes in Q sets of audio processing results with similarity greater than a threshold are fused to obtain a fused voiceprint feature code. The fused voiceprint feature code is used to characterize that the users corresponding to the Q sets of audio processing results are the same user, where Q is an integer greater than 1.

[0011] The multiple sets of audio processing results and the fused voiceprint feature codes are stored in a pre-built voiceprint library to form the preset fused voiceprint library.

[0012] In some implementations, the S devices are all of the same brand, and the voiceprint fusion device is the central control device of the target area or the cloud platform of the brand of the S devices.

[0013] In some implementations, the S devices are of different brands, and the voiceprint fusion device is either a central control device for the target area or a cloud-based device that provides intelligent services to devices within the target area.

[0014] In some embodiments, each set of voiceprint attributes in the preset fused voiceprint library includes a first type of voiceprint attribute for characterizing user characteristics and a second type of voiceprint attribute for characterizing audio characteristics. The method further includes:

[0015] Obtain the first type of voiceprint attribute and the second type of voiceprint attribute of the target audio;

[0016] The step of matching the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fused voiceprint library to obtain the matching result includes:

[0017] The first type of voiceprint attribute of the target audio is matched with the first type of voiceprint attribute in each group of voiceprint attributes;

[0018] If there are M candidate voiceprint attributes among the N groups of voiceprint attributes that match the first type of voiceprint attribute of the target audio, the second type of voiceprint attribute of the target audio is matched with the second type of voiceprint attribute in each group of candidate voiceprint attributes to obtain the matching result, where M is a positive integer.

[0019] In some implementations, the second type of voiceprint attribute includes K seed attributes, where K is a positive integer. The step of matching the second type of voiceprint attribute of the target audio with the second type of voiceprint attribute in each group of candidate voiceprint attributes to obtain the matching result includes:

[0020] For each group of candidate voiceprint attributes, the similarity of each sub-attribute of the second type of voiceprint attribute in the group of candidate voiceprint attributes is calculated with the corresponding sub-attribute in the target audio to obtain the first similarity of each sub-attribute. Based on the first similarity of each sub-attribute, the matching result between the target audio and the group of candidate voiceprint attributes is determined.

[0021] In some implementations, determining the matching result between the target audio and the set of candidate voiceprint attributes based on the first similarity of each sub-attribute includes:

[0022] Determine the weight of each sub-attribute in the K seed attributes;

[0023] For each group of candidate voiceprint attributes, based on the weight of each sub-attribute, the first similarity corresponding to each sub-attribute in the group of candidate voiceprint attributes is weighted, and the weighted result is used as the second similarity between the group of candidate voiceprint attributes and the target audio.

[0024] If there exists a group of P candidate voiceprint attributes with a second similarity greater than a preset similarity, the candidate voiceprint attribute with the largest second similarity among the group of P candidate voiceprint attributes is taken as the target voiceprint attribute that matches the target audio, where P is a positive integer less than or equal to K.

[0025] In some implementations, the sub-attributes of the second type of voiceprint attribute include audio purity, audio resonance, audio speech rate, and audio pitch.

[0026] In some embodiments, the preset fused voiceprint library further includes at least one fused voiceprint feature code and a user profile corresponding to each fused voiceprint feature code, wherein each fused voiceprint feature code corresponds to the same user as at least one voiceprint feature code in the preset fused voiceprint library, and controlling the operation of the target device based on the user profile corresponding to the target voiceprint attribute includes:

[0027] Based on the preset fusion voiceprint library, the voiceprint feature code corresponding to the target voiceprint attribute is determined, and the target fusion voiceprint feature code corresponding to the voiceprint feature code of the target voiceprint attribute is determined from the preset fusion voiceprint library.

[0028] Based on the user profile corresponding to the target fused voiceprint feature code, the operation of the target device is controlled.

[0029] In some embodiments, if the matching result indicates that a target voiceprint attribute exists in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio, the method further includes:

[0030] Associate the target voiceprint feature code of the target audio with the target fused voiceprint feature code.

[0031] In a second aspect, embodiments of the present invention provide a device control apparatus, comprising:

[0032] The acquisition module is used to acquire the target voiceprint feature code fed back by the voice service provider of the target device when the target device detects the target audio. The target device is located in the target area, and the target area contains S devices, where S is a positive integer.

[0033] The first processing module is configured to, if the target voiceprint feature code is not included in the voiceprint feature codes in the preset fusion voiceprint library, match the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result, wherein the preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer;

[0034] The second processing module is used to control the operation of the target device based on the user profile corresponding to the target audio if the matching result indicates that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio.

[0035] Thirdly, embodiments of the present invention provide an electrical device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described device control method.

[0036] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described device control method.

[0037] The embodiments of the present invention provide one or more technical solutions that achieve at least the following technical effects or advantages:

[0038] The device control method provided in this specification, if the target device detects target audio, obtains the target voiceprint feature code fed back by the voice service provider of the target audio. The target device is located in a target area containing S devices. If the target voiceprint feature code is not included in the voiceprint feature codes in the preset fusion voiceprint library, it indicates that the user corresponding to the target audio is a new user of the target device. To provide the user with services that meet their needs, the voiceprint attributes of the target audio are matched with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result. The preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code. If the matching result indicates that the target voiceprint attribute in the preset fusion voiceprint library matches the voiceprint attribute of the target audio, the target device is controlled to operate based on the user profile corresponding to the target voiceprint attribute. In this solution, the voiceprint attributes of the user's target audio can be analyzed. The target voiceprint attributes that match the voiceprint attributes of the target audio can be selected from the preset fusion voiceprint library. Since different users have different voiceprint attributes, the user profile corresponding to the target voiceprint attribute can be used as the user profile corresponding to the target audio through the target voiceprint attribute that matches the voiceprint attribute. Based on this user profile, the target device can be controlled, realizing personalized services for the user corresponding to the target audio and improving the user experience. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0040] Figure 1 A flowchart of a device control method provided in the embodiments of this specification;

[0041] Figure 2 A schematic diagram of a device control system provided as an embodiment of this specification;

[0042] Figure 3 A schematic diagram of a device control apparatus provided in the embodiments of this specification;

[0043] Figure 4 This is a schematic diagram of an electrical device provided as an embodiment of this specification. Detailed Implementation

[0044] This specification provides a device control method, apparatus, electrical equipment, and medium. The method includes: if a target device detects target audio, obtaining a target voiceprint feature code fed back by the voice service provider of the target device after processing the target audio, wherein the target device is located in a target area containing S devices, where S is a positive integer; if the target voiceprint feature code is not included in the voiceprint feature codes in a preset fusion voiceprint library, matching the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result, wherein the preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer; if the matching result indicates that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attributes of the target audio, controlling the operation of the target device based on the user profile corresponding to the target voiceprint attribute.

[0045] The solution in the embodiments of this specification analyzes the voiceprint attributes of the user's target audio and filters out the target voiceprint attributes that match the voiceprint attributes of the target audio from a preset fusion voiceprint library. Since different users have different voiceprint attributes, the user profile corresponding to the target voiceprint attribute can be used as the user profile corresponding to the target audio through the target voiceprint attribute that matches the voiceprint attribute. Based on the user profile, the target device is controlled, realizing personalized services for the user corresponding to the target audio and improving the user experience.

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0047] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0048] This specification provides an embodiment of a device control method, such as... Figure 1 The diagram shown is a flowchart of a device control method provided in an embodiment of this specification. The method includes the following steps:

[0049] Step S101: If the target device detects the target audio, obtain the target voiceprint feature code fed back by the voice service provider of the target device when processing the target audio. The target device is located in the target area, and the target area contains S devices, where S is a positive integer.

[0050] Step S102: If the target voiceprint feature code is not included in the voiceprint feature codes in the preset fusion voiceprint library, match the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result. The preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer.

[0051] Step S103: If the matching result indicates that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio, control the operation of the target device based on the user profile corresponding to the target voiceprint attribute.

[0052] The methods described in this specification can be applied to electrical appliances, such as air conditioners, televisions, smart speakers, etc. They can also be applied to central control devices for multiple terminal devices in a user's area, such as gateways. Furthermore, they can be applied to servers that communicate with electrical appliances and / or central control devices. They can also be applied to systems composed of servers, electrical appliances, and / or central control devices, without any limitations.

[0053] The methods provided in the embodiments of this specification can be applied to the use of household appliances as well as in the use of appliances in public areas. Taking the use of household appliances as an example, a user's home may have multiple electrical appliances installed, such as a television, air conditioner, air purifier, smart speaker, etc. The user can issue voice commands to any of these electrical appliances, and the appliance can respond to the received voice command and provide services to the user according to the user's personalized needs. For example, if the air conditioner detects the user's voice command "turn on the air conditioner," the air conditioner can determine the user's profile based on the user's voice audio. For example, if the user's air conditioner usage habits are set to a temperature of 24°C and a medium fan speed, then the air conditioner will operate according to the user profile to provide services that meet the user's needs and improve the user experience.

[0054] It should be noted that the devices in a user's home can be of the same brand or different brands. The voice service providers offering voice services for devices of the same brand can be the same or different, and the voice service providers offering voice services for devices of different brands can also be the same or different. It should also be noted that processing the same user's audio using voice services from different providers will result in different voiceprint signatures, while processing the same user's audio using the same voice service provider will result in the same voiceprint signature.

[0055] For example, voice service provider 1 provides services to device 1, and voice service provider 2 provides services to device 2. When user A uses device 1, device 1 receives user A's audio and uploads it to voice service provider 1. Voice service provider 1 processes the audio, for example, using its own audio processing algorithm or model to generate a voiceprint feature code 'a', and then sends voiceprint feature code 'a' back to device 1. Similarly, when user A uses device 2, device 2 receives user A's audio and uploads it to voice service provider 2. Voice service provider 2 processes the audio, for example, using its own audio processing algorithm or model to generate a voiceprint feature code 'b'. Therefore, even for the same user, the voiceprint feature codes may differ depending on the voice service provider used by the device.

[0056] In this embodiment, the target area in step S101 can be the user's workplace or the user's home area. For ease of explanation, the user's home area is used as an example here. The target area includes S devices, which can be all home appliances installed in the user's home that support voice services, such as televisions, air conditioners, smart speakers, etc. The S devices can be of the same brand or different brands, and the voice service providers corresponding to each of the S devices can be the same or different. The target device can be any of the S devices. The target audio can be the audio corresponding to the voice command issued by the user. In some embodiments, the target audio can be the audio corresponding to the wake word or the audio corresponding to the voice command issued by the user to perform a specific operation on the device.

[0057] After detecting the target audio, the target device can upload it to the cloud of the corresponding voice service provider. The voice service provider then processes the audio, generates a target voiceprint feature code, and sends it back to the target device. It should be noted that for each device, a unique voiceprint feature code can be obtained for each user's audio received, and each voiceprint feature code uniquely identifies the user.

[0058] In step S102, to determine the user device usage habits corresponding to the target audio, the target voiceprint feature code of the target audio can be matched with the voiceprint feature codes in the preset fusion voiceprint library. In this embodiment of the specification, for a device, after receiving the user's audio, the device can generate the corresponding voiceprint feature code in the above manner. In addition, the user's audio can be processed to obtain the audio's voiceprint attributes, which may include, but are not limited to, the user's gender, age group, pitch, purity, resonance, speech rate, etc. The voiceprint feature codes of each user corresponding to each device, as well as the mapping relationship between the voiceprint feature codes and voiceprint attributes, are stored in the preset fusion voiceprint library. It should be noted that the preset fusion voiceprint library can be pre-built or continuously updated during the use of the device.

[0059] To better understand the preset fusion voiceprint library in the embodiments of this specification, the construction process of the preset fusion voiceprint library is described below. In some embodiments, the preset fusion voiceprint library is stored in a voiceprint fusion device, which is communicatively connected to S devices. The preset fusion voiceprint library can be constructed through the following steps: monitoring the audio processing results of each of the S devices, wherein the audio processing results include the voiceprint feature code and voiceprint attributes of the audio to be processed, wherein the voiceprint feature code of the audio to be processed in the audio processing results is fed back by the voice service provider of the corresponding device when processing the audio to be processed; calculating the similarity of the voiceprint attributes in the multiple sets of monitored audio processing results, and fusing the voiceprint feature codes in Q sets of audio processing results with similarity greater than a threshold to obtain a fused voiceprint feature code, wherein the fused voiceprint feature code is used to characterize that the users corresponding to the Q sets of audio processing results are the same user, where Q is an integer greater than 1; storing the multiple sets of audio processing results and the fused voiceprint feature code in a pre-constructed voiceprint library to form the preset fusion voiceprint library.

[0060] In this embodiment of the specification, considering that the brands of the S devices may be different and the voice service providers of each device may also be different, in order to achieve voiceprint fusion between different brands and different voice service providers, the storage location of the preset fused voiceprint library can be selected according to the device brands of the S devices.

[0061] In some embodiments, the S devices are all of the same brand, and the voiceprint fusion device is the central control device of the target area or the cloud of the brand of the S devices.

[0062] Specifically, although the S devices are of the same brand, the corresponding voice service providers may not be exactly the same. After receiving audio, the devices need to upload the audio to the corresponding voice service provider for processing to obtain voiceprint feature codes. However, different voice service providers cannot communicate with each other, so voiceprint feature codes cannot be shared. In the embodiments of this specification, in order to achieve cross-device and cross-voice service provider voiceprint fusion, a preset fused voiceprint library is constructed and set in a voiceprint fusion device that can communicate with each device. This ensures that each device can access the preset fused voiceprint library when processing audio.

[0063] Since the S devices are from the same brand, and all S devices can communicate with the brand's cloud platform, a preset fusion voiceprint library can be set on the brand's cloud platform. In other words, the voiceprint fusion device is the brand's cloud platform for the S devices. During implementation, when a device receives audio and obtains its voiceprint feature code and attributes, it can send these to the brand's cloud platform. The brand's cloud platform then matches the voiceprint features and attributes with the preset fusion voiceprint library.

[0064] Furthermore, since all S devices can communicate with the central control device (such as a gateway) in the target area, the preset fusion voiceprint library can also be stored in the central control device; that is, the voiceprint fusion device is the central control device. In the specific implementation process, when a device receives audio and obtains the voiceprint feature code and voiceprint attributes of the audio, it can send the voiceprint feature code and voiceprint attributes to the central control device, where the voiceprint features and voiceprint attributes are matched with the preset fusion voiceprint library.

[0065] In other embodiments, the S devices are of different brands, and the voiceprint fusion device is a central control device for the target area or a cloud-based device that provides intelligent services to devices within the target area.

[0066] Specifically, the S devices are from different brands, and the cloud services they communicate with also differ. To ensure that all S devices can access the preset fusion voiceprint library, this library can be stored in a central control device. Since all S devices can communicate with the central control device, it can act as the voiceprint fusion device. In practice, when a device receives audio and obtains its voiceprint feature code and attributes, it sends these to the central control device. The central control device then matches the voiceprint features and attributes with the preset fusion voiceprint library.

[0067] Furthermore, if the target area, such as a user's home, is equipped with whole-house smart home services, then the S devices can communicate with the cloud providing these services. Therefore, a preset fusion voiceprint library can be stored in the central control device, meaning the voiceprint fusion device acts as the cloud providing the smart services. When a device receives audio and obtains its voiceprint feature code and attributes, it can send these to the cloud providing the smart services, where the voiceprint features and attributes are matched against the preset fusion voiceprint library.

[0068] When constructing the pre-built fusion voiceprint library, each of the S devices is monitored. Specifically, for each of the S devices, if it receives a user's wake-up or control voice, it sends the audio corresponding to the wake-up or control voice as the audio to be processed to its voice service provider. The voice service provider processes the audio to be processed and generates a voiceprint feature code for the audio. In some embodiments, the voice service provider can also perform attribute analysis on the audio to be processed to obtain the voiceprint attributes of the audio and feed them back to the device. In other embodiments, the device can analyze the audio to be processed using an audio analysis model or algorithm deployed on the device to obtain the voiceprint attributes of the audio. When the audio processing result of the audio to be processed, i.e., the voiceprint feature code and voiceprint attributes of the audio to be processed, is detected, the audio processing result can be extracted or intercepted and stored in the pre-built voiceprint library. In this way, voiceprint feature codes and voiceprint attributes obtained by different voice service providers from processing the same or different users' audio can be saved.

[0069] To fuse the voiceprint feature code and voiceprint attributes of the same user, in this embodiment, the similarity of voiceprint attributes in any two or more sets of audio processing results can be calculated. If the similarity calculation result is greater than a threshold, it indicates that these sets of voiceprint attributes are similar and correspond to the same user. The threshold can be set according to actual needs, for example, a threshold of 90%, 95%, etc. For example, if the similarity between three sets of voiceprint attributes is greater than the threshold, and the voiceprint feature codes corresponding to these three sets of voiceprint attributes are x, y, and z respectively, then x, y, and z correspond to the feature codes of the same user under different voice service providers. x, y, and z can be fused to obtain a fused voiceprint feature code r, meaning that x, y, z, and r all correspond to the same user. Furthermore, the fused voiceprint feature code, and the correspondence between the fused voiceprint feature code and the voiceprint feature code, are also stored in a pre-built voiceprint library to form a preset fused voiceprint library.

[0070] In step S102, the target voiceprint feature code is matched with the voiceprint feature codes in the preset fusion voiceprint library. This can be done by comparing the target voiceprint feature code with each voiceprint feature code in the preset fusion voiceprint library one by one, or by first filtering out multiple voiceprint feature codes fed back by the voice service provider of the target device from the preset fusion voiceprint library, and then comparing the target voiceprint feature code with the multiple voiceprint feature codes corresponding to the target device one by one. There is no limitation here.

[0071] If the target voiceprint feature code exists in the preset fusion voiceprint library, it indicates that the user is a historical user of the target device. In this case, a user profile corresponding to the target voiceprint feature code can be obtained. The user profile can be generated based on the user's historical usage records, such as historical voice control commands and operation data. To improve user experience, the operation of the target device can be controlled based on the user profile.

[0072] In this embodiment of the specification, if the target voiceprint feature code is not included in the preset fusion voiceprint library, the user's usage habits cannot be determined through the target voiceprint feature code. In order to still provide services that meet the user's needs, the voiceprint attributes of the target audio can be matched with N sets of voiceprint attributes in the preset fusion voiceprint library. The N sets of voiceprint attributes correspond one-to-one with the N voiceprint feature codes included in the preset fusion voiceprint library. It should be understood that although the voiceprint feature codes of the same user differ across different devices, the voiceprint attributes of the same user across different devices have a high degree of similarity. Therefore, in step S102, voiceprint attribute matching can be used to further filter voiceprint attributes similar to the voiceprint attributes of the target audio.

[0073] In some embodiments, each set of voiceprint attributes in the preset fusion voiceprint library includes a first type of voiceprint attribute for characterizing user characteristics and a second type of voiceprint attribute for characterizing audio characteristics. Voiceprint attribute matching can be achieved through the following steps: obtaining the first type of voiceprint attribute and the second type of voiceprint attribute of the target audio; matching the first type of voiceprint attribute of the target audio with the first type of voiceprint attribute in each set of voiceprint attributes; if there are M sets of candidate voiceprint attributes in the N sets of voiceprint attributes that match the first type of voiceprint attribute of the target audio, matching the second type of voiceprint attribute of the target audio with the second type of voiceprint attribute in each set of candidate voiceprint attributes to obtain the matching result, where M is a positive integer.

[0074] The first type of voiceprint attribute can be an attribute used to characterize user characteristics, and the second type of voiceprint attribute can be an attribute used to characterize audio characteristics. In some embodiments, the first type of voiceprint attribute may include sub-attributes such as the user's gender and age group, and the second type of voiceprint attribute may include attributes such as audio purity, audio speech rate, and audio pitch.

[0075] In the embodiments of this specification, audio analysis and processing can be performed on the target audio to obtain a second type of voiceprint attribute. To reduce computational load and improve computational efficiency, the first type of voiceprint attribute of the target audio can be matched with the first type of voiceprint attribute in each group of voiceprint attributes. For example, if the gender in the first type of voiceprint attribute of the target audio is male and the age is elderly, it is determined whether the first type of voiceprint attribute with the gender of male and the age of elderly also exists in N groups of voiceprint attributes. If not, it indicates that the user is a new user, and the device operation can be controlled according to the default operating parameters or based on the usage habits of people of the same type.

[0076] If there are M candidate voiceprint attributes among the N groups of voiceprint attributes that match the first type of voiceprint attributes of the target audio, then the second type of voiceprint attributes are further matched. The matching process for the second type of voiceprint attributes can be implemented in various ways. In some implementations, the second type of voiceprint attributes includes K seed attributes, where K is a positive integer. The second type of voiceprint attributes can be matched through the following steps: For each group of candidate voiceprint attributes, calculate the similarity between each sub-attribute of the second type of voiceprint attributes in that group of candidate voiceprint attributes and the corresponding sub-attribute in the target audio, obtaining a first similarity for each sub-attribute, and determining the matching result between the target audio and that group of candidate voiceprint attributes based on the first similarity for each sub-attribute.

[0077] For ease of explanation, let's take the K-seed attributes, including audio purity, audio resonance, audio speech rate, and audio pitch, as an example. Pitch characterizes the speed of vocal cord vibration; the faster the vibration, the higher the fundamental frequency, and the easier the pitch is to control. Purity characterizes the degree of vocal cord closure, including high purity, medium purity, and low purity. In some embodiments, purity can be significantly affected by illness or overuse of the voice. Resonance characterizes the resonance frequency of the audio. Resonance methods can include oral resonance, chest resonance, and nasal resonance. Oral resonance makes the sound solid and clear, chest resonance makes the sound deep and resonant, and nasal resonance makes the sound high-pitched and bright. Speech rate characterizes the speed at which the user speaks.

[0078] For the similarity calculation of each sub-attribute, the similarity calculation method can be selected according to actual needs. In some embodiments, the Sigmoid function can be used for processing. Below, we take the similarity calculation of the second type of voiceprint attribute of the target audio and the second type of voiceprint attribute of a set of candidate voiceprint attributes as an example. Here, the first similarity corresponding to pitch is represented by α, the first similarity corresponding to purity is represented by β, the first similarity corresponding to resonance is represented by γ, and the first similarity corresponding to speech rate is represented by δ. Then, the similarity of each sub-attribute is as follows:

[0079]

[0080] β=(1-Sigmoid(|β η -β θ |÷β θ ))*2

[0081] γ=(1-Sigmoid(|γ η -γ θ |÷γ θ ))*2

[0082] δ=(1-Sigmoid(|δ η -δ θ |÷δ θ ))*2

[0083] Where η is the target audio, θ is the audio corresponding to the candidate voiceprint attribute, and n is the total number of characters contained in the target audio. i For the i-th word in the target audio, α ηi Let θ be the pitch of the i-th word in the target audio. i For the i-th character in the audio corresponding to the candidate voiceprint attribute, α θi β represents the pitch of the i-th word in the audio corresponding to the candidate voiceprint attribute. η For the purity of the target audio, β θ γ represents the audio purity corresponding to this candidate voiceprint attribute. η For the resonance of the target audio, γ θ δ represents the audio purity corresponding to the candidate voiceprint attribute. η δ represents the speech rate of the target audio. θ The audio speech rate corresponding to the candidate voiceprint attributes.

[0084] It should be understood that the target audio can contain multiple characters. Taking the target audio as a wake word as an example, the target audio can include 4 or 5 characters. Therefore, when filtering candidate voiceprint attributes, the voiceprint attributes corresponding to audio that matches the target audio text need to be used as candidate voiceprint attributes. For example, when the target audio is a wake word, the audio corresponding to each set of candidate voiceprint attributes is also a wake word. In this way, the characters contained in the second type of voiceprint attributes of the target audio and the second type of voiceprint attributes of the candidate voiceprint attributes are consistent, thereby enabling pitch matching of each character and ensuring the accuracy of pitch similarity calculation. In some embodiments, when constructing a preset fused voiceprint library, text recognition can be performed on each audio received by each device. Then, when determining the second type of voiceprint attributes of the audio, the pitch of each character in the audio is determined and saved.

[0085] In this embodiment of the specification, after obtaining the first similarity of each sub-attribute, the matching result between the second type of voiceprint attribute of the target audio and the second type of voiceprint attribute in each group of candidate voiceprint attributes can be determined by the following steps: determining the weight of each sub-attribute in the K seed attributes; for each group of candidate voiceprint attributes, based on the weight of each sub-attribute, performing weighted processing on the first similarity corresponding to each sub-attribute in the group of candidate voiceprint attributes, and using the weighted processing result as the second similarity between the group of candidate voiceprint attributes and the target audio; if there are P groups of candidate voiceprint attributes with a second similarity greater than a preset similarity, the candidate voiceprint attribute with the largest second similarity in the P groups of candidate voiceprint attributes is taken as the target voiceprint attribute that matches the target audio, where P is a positive integer less than or equal to K.

[0086] Specifically, the weight of each sub-attribute can be set according to actual needs. In some embodiments, considering that resonance is more closely related to human physiological characteristics (such as the size and shape of the nasal cavity, larynx, and lungs) and largely determines the recognizability of the voice, the weight of resonance can be set higher. For example, if the weight of resonance is set to 0.4, and the weights of pitch, purity, and speech rate are all set to 0.2, then the second similarity ε can be obtained by weighting using the following formula:

[0087] ε=0.2*α+0.2*β+0.4*γ+0.2*δ

[0088] In this embodiment, the second similarity between the target audio and each group of candidate voiceprint attributes is compared with a preset similarity. The preset similarity can be set according to actual needs, for example, a preset similarity of 0.8, 0.9, etc. If there are P groups of candidate voiceprint attributes with a similarity greater than the preset similarity, where P is a positive integer, then the candidate voiceprint attribute with the highest second similarity can be taken as the target voiceprint attribute most similar to the target audio. It can be assumed that the user corresponding to the target voiceprint attribute and the user of the target audio are the same user.

[0089] It should be understood that if the second similarity between the target audio and each set of candidate voiceprint attributes is less than or equal to the preset similarity, it indicates that the user is a new user, and the device operation can be controlled according to the default operating parameters or based on the usage habits of the same type of people.

[0090] In step S103, if a target voiceprint attribute that matches the target audio is determined, it indicates that the user who output the target audio has used other devices. When using other devices, the user profile of that user can be recorded, and the operation of the target device can be controlled through that user.

[0091] In some embodiments, the preset fusion voiceprint library may further include at least one fusion voiceprint feature code and a user profile corresponding to each fusion voiceprint feature code. The fusion voiceprint feature code and at least one voiceprint feature code in the preset fusion voiceprint library correspond to the same user. Specifically, following the example above, after device 1 receives the audio of user A, it generates voiceprint feature code a. After device 2 receives the audio of user A, it generates voiceprint feature code b. Voiceprint feature code a and voiceprint feature code b are fused to obtain fusion voiceprint feature code c. It can be seen that a, b, and c all correspond to the same user. In addition, each device records the user's usage habits during use. For example, when device 1 is an air conditioner, device 1 can record parameters such as the set temperature and fan speed when the user uses the air conditioner. When device 2 is a smart speaker, device 2 can record the user's preferred music type, playback volume, etc. Based on the user habits collected by each device, a user profile can be generated, and the user profile can be associated with the fusion voiceprint feature code.

[0092] Step S103 can be achieved through the following steps: based on the preset fusion voiceprint library, determine the voiceprint feature code corresponding to the target voiceprint attribute, and determine the target fusion voiceprint feature code corresponding to the voiceprint feature code of the target voiceprint attribute from the preset fusion voiceprint library; based on the user profile corresponding to the target fusion voiceprint feature code, control the operation of the target device.

[0093] Since the preset fusion voiceprint library includes the correspondence between voiceprint feature codes and voiceprint attributes, the voiceprint feature code corresponding to the target voiceprint attribute can be determined based on this correspondence. Then, based on the correspondence between the voiceprint feature code and the fusion voiceprint feature code, the target fusion voiceprint feature code corresponding to the voiceprint feature code of the target voiceprint attribute can be determined. The user profile corresponding to the target fusion voiceprint feature code is used as the user profile of the current output target audio to control the target device.

[0094] For example, if a user sets the temperature of air conditioner D to 24℃ in their usual room, this user profile has been updated in the preset fusion voiceprint library. When the user uses air conditioner E for the first time in a non-usual room and says "turn on the air conditioner," the user profile can be obtained through the above method. The user profile includes the user's habit of setting the temperature to 24℃ when using the air conditioner. Then, air conditioner E will automatically adjust the setting parameter to 24℃, greatly improving the user experience.

[0095] In the embodiments of this specification, if target voiceprint information matching the target audio voiceprint attributes exists, the following steps can be performed: associating the target voiceprint feature code of the target audio with the target fused voiceprint feature code. That is, updating the preset fused voiceprint library and fusing and associating the target voiceprint feature code with other voiceprint feature codes of the user.

[0096] To better understand the device control methods provided in the embodiments of this specification, please refer to... Figure 2 This specification provides an embodiment of a device control system, which includes an audio receiving module, a first type of voiceprint attribute detection module, a second type of voiceprint attribute detection module, and a voiceprint fusion module.

[0097] The audio receiving module, which can be set up on the device, receives the user's wake-up audio or device control audio. The first type of voiceprint attribute detection module, which can be set up on the device or in the cloud of the voice service provider, processes the received user audio to obtain the first type of voiceprint attributes, such as gender and age. These attributes can also be used to generate user profiles and for user identification. The second type of voiceprint attribute detection module, which can also be set up on the device or in the cloud of the voice service provider, processes the user audio to obtain the second type of voiceprint attributes, such as audio purity, resonance, speech rate, and pitch. The voiceprint fusion module, which can be set up on the voiceprint fusion device, builds a pre-set fusion voiceprint library for each household. It fuses voiceprint attributes, voiceprint feature codes, and user profiles to achieve unified voiceprint identity within the same household.

[0098] As described above, the method provided in this specification, in order to achieve unified user identity across all devices in the house, collects, merges, and stores the user's audio voiceprint attributes and the voiceprint feature codes generated by various voice service providers corresponding to the user's audio in a preset fused voiceprint library. When a user uses a device for the first time, the device can first obtain the user's voiceprint feature code through the corresponding voice service provider. If the voiceprint feature code is not included in the preset fused voiceprint library, the user's identity is further determined by comparing it with the voiceprint attributes in the preset fused voiceprint library, and finally a user profile is obtained.

[0099] It should be noted that voice service providers typically charge based on the number of calls when providing voice services. Compared to determining a user's identity by calling different voice service providers, the embodiments in this specification can greatly reduce the number of calls to voice service providers and reduce the cost of user identity verification by constructing a preset fusion voiceprint library.

[0100] Specifically, if a device, after detecting a user's audio, simultaneously sends the user's audio to different voice service providers to obtain voiceprint feature codes from each provider's audio processing, and then determines whether the user has used a particular voice service provider based on these voiceprint feature codes, it will significantly increase the costs for those providers. The preset fused voiceprint library in this embodiment avoids simultaneously calling multiple voice service providers; it only needs to query the preset fused voiceprint library to determine the user's identity, effectively reducing costs.

[0101] In summary, the solution of this embodiment analyzes the voiceprint attributes of the current user's audio to determine the target voiceprint attribute that matches the voiceprint attribute. The user profile corresponding to the target voiceprint attribute is the user profile of the current user. By controlling the target device through the user profile, services tailored to the user's needs can be provided, significantly improving the user experience.

[0102] Furthermore, the solution in the embodiments of this specification, by performing two-level judgments on the audio attributes of the target audio—first-type voiceprint attributes and second-type voiceprint attributes—can effectively reduce the computational load and improve judgment efficiency. Simultaneously, assigning different weights to different sub-attributes within the second-type voiceprint attributes improves matching accuracy.

[0103] The solutions described in this specification, by constructing and updating a preset fused voiceprint library and utilizing the similarity between voiceprint attributes, and the correlation between voiceprint attributes, voiceprint feature codes, fused voiceprint feature codes, and user profiles, can achieve personalized services across devices based on the user. Even if a user has never used a particular device, a user profile can be determined by analyzing the user's audio, thereby enabling device control.

[0104] Based on the same inventive concept, embodiments of the present invention provide a device control apparatus, such as... Figure 3 As shown, the device includes:

[0105] The acquisition module 301 is used to acquire the target voiceprint feature code fed back by the voice service provider of the target device when the target device detects the target audio.

[0106] The first processing module 302 is configured to, if the target voiceprint feature code is not included in the voiceprint feature codes in the preset fusion voiceprint library, match the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result, wherein the preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer;

[0107] The second processing module 303 is used to control the operation of the target device based on the user profile corresponding to the target audio if the matching result shows that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio.

[0108] In some embodiments, the preset fused voiceprint library is stored in a voiceprint fusion device, which is communicatively connected to the S devices. The preset fused voiceprint library is constructed through the following steps:

[0109] Listen to the audio processing results of each of the S devices. The audio processing results include the voiceprint feature code and voiceprint attributes of the audio to be processed. The voiceprint feature code of the audio to be processed in the audio processing results is fed back by the voice service provider of the corresponding device when processing the audio to be processed.

[0110] The similarity of voiceprint attributes in multiple sets of audio processing results is calculated, and the voiceprint feature codes in Q sets of audio processing results with similarity greater than a threshold are fused to obtain a fused voiceprint feature code. The fused voiceprint feature code is used to characterize that the users corresponding to the Q sets of audio processing results are the same user, where Q is an integer greater than 1.

[0111] The multiple sets of audio processing results and the fused voiceprint feature codes are stored in a pre-built voiceprint library to form the preset fused voiceprint library.

[0112] In some implementations, the S devices are all of the same brand, and the voiceprint fusion device is the central control device of the target area or the cloud platform of the brand of the S devices.

[0113] In some implementations, the S devices are of different brands, and the voiceprint fusion device is either a central control device for the target area or a cloud-based device that provides intelligent services to devices within the target area.

[0114] In some embodiments, each set of voiceprint attributes in the preset fused voiceprint library includes a first type of voiceprint attribute for characterizing user characteristics and a second type of voiceprint attribute for characterizing audio characteristics. The device further includes:

[0115] The voiceprint attribute acquisition module is used to acquire the first type of voiceprint attribute and the second type of voiceprint attribute of the target audio.

[0116] The first processing module 302 is used to match the first type of voiceprint attribute of the target audio with the first type of voiceprint attribute in each group of voiceprint attributes; if there are M groups of candidate voiceprint attributes in the N groups of voiceprint attributes that match the first type of voiceprint attribute of the target audio, the second type of voiceprint attribute of the target audio is matched with the second type of voiceprint attribute in each group of candidate voiceprint attributes to obtain the matching result, where M is a positive integer.

[0117] In some implementations, the second type of voiceprint attribute includes a K-seed attribute, where K is a positive integer. The first processing module 302 is used for:

[0118] For each group of candidate voiceprint attributes, the similarity of each sub-attribute of the second type of voiceprint attribute in the group of candidate voiceprint attributes is calculated with the corresponding sub-attribute in the target audio to obtain the first similarity of each sub-attribute. Based on the first similarity of each sub-attribute, the matching result between the target audio and the group of candidate voiceprint attributes is determined.

[0119] In some implementations, the first processing module 302 is configured to:

[0120] Determine the weight of each sub-attribute in the K seed attributes;

[0121] For each group of candidate voiceprint attributes, based on the weight of each sub-attribute, the first similarity corresponding to each sub-attribute in the group of candidate voiceprint attributes is weighted, and the weighted result is used as the second similarity between the group of candidate voiceprint attributes and the target audio.

[0122] If there exists a group of P candidate voiceprint attributes with a second similarity greater than a preset similarity, the candidate voiceprint attribute with the largest second similarity among the group of P candidate voiceprint attributes is taken as the target voiceprint attribute that matches the target audio, where P is a positive integer less than or equal to K.

[0123] In some implementations, the sub-attributes of the second type of voiceprint attribute include audio purity, audio resonance, audio speech rate, and audio pitch.

[0124] In some embodiments, the preset fused voiceprint library further includes at least one fused voiceprint feature code and a user profile corresponding to each fused voiceprint feature code, wherein each fused voiceprint feature code corresponds to the same user as at least one voiceprint feature code in the preset fused voiceprint library, and the second processing module 303 is used for:

[0125] Based on the preset fusion voiceprint library, the voiceprint feature code corresponding to the target voiceprint attribute is determined, and the target fusion voiceprint feature code corresponding to the voiceprint feature code of the target voiceprint attribute is determined from the preset fusion voiceprint library.

[0126] Based on the user profile corresponding to the target fused voiceprint feature code, the operation of the target device is controlled.

[0127] In some embodiments, the apparatus further includes:

[0128] The association module is used to associate the target voiceprint feature code of the target audio with the target fused voiceprint feature code.

[0129] Regarding the above-mentioned device, the specific functions of each module have been described in detail in the embodiments of the device control method provided in this specification, and will not be elaborated here.

[0130] Based on the same inventive concept, embodiments of the present invention provide an electrical device, see reference. Figure 4 As shown, it includes a memory 404, a processor 402, and a computer program stored in the memory 404 and executable on the processor 402. When the processor 402 executes the program, it implements any one of the embodiments of the device control method.

[0131] Among them, Figure 4 In this document, a bus architecture (represented by bus 400) is used. Bus 400 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 402 and memory represented by memory 404. Bus 400 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 405 provides an interface between bus 400 and receiver 401 and transmitter 403. Receiver 401 and transmitter 403 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 402 is responsible for managing bus 400 and general processing, while memory 404 can be used to store data used by processor 402 during operation.

[0132] The functions described herein can be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions can be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and embodiments are within the scope and spirit of this invention and the appended claims. For example, due to the nature of software, the functions described above can be implemented using software executed by a processor, hardware, firmware, hardwired, or any combination thereof. Furthermore, the functional units can be integrated into a single processing unit, or each unit can exist physically separately, or two or more units can be integrated into a single unit.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0134] The units described as separate components may or may not be physically separate. Similarly, the components of the control device may or may not be physical units; they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0135] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0136] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A device control method, characterized in that, include: If the target device detects the target audio, the target voiceprint feature code fed back by the voice service provider of the target device after processing the target audio is obtained. The target device is located in the target area, and the target area contains S devices, where S is a positive integer. If the target voiceprint feature code is not included in the voiceprint feature code in the preset fusion voiceprint library, the voiceprint attributes of the target audio are matched with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result. The preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer. If the matching result indicates that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio, the target device is controlled to operate based on the user profile corresponding to the target voiceprint attribute.

2. The method as described in claim 1, characterized in that, The preset fused voiceprint library is stored in the voiceprint fusion device, which is communicatively connected to the S devices. The preset fused voiceprint library is constructed through the following steps: Listen to the audio processing results of each of the S devices. The audio processing results include the voiceprint feature code and voiceprint attributes of the audio to be processed. The voiceprint feature code of the audio to be processed in the audio processing results is fed back by the voice service provider of the corresponding device when processing the audio to be processed. The similarity of voiceprint attributes in multiple sets of audio processing results is calculated, and the voiceprint feature codes in Q sets of audio processing results with similarity greater than a threshold are fused to obtain a fused voiceprint feature code. The fused voiceprint feature code is used to characterize that the users corresponding to the Q sets of audio processing results are the same user, where Q is an integer greater than 1. The multiple sets of audio processing results and the fused voiceprint feature codes are stored in a pre-built voiceprint library to form the preset fused voiceprint library.

3. The method as described in claim 2, characterized in that, All S devices are of the same brand, and the voiceprint fusion device is either the central control device of the target area or the cloud platform of the brand of the S devices.

4. The method as described in claim 2, characterized in that, The S devices are of different brands, and the voiceprint fusion device is either a central control device for the target area or a cloud-based device that provides intelligent services to devices within the target area.

5. The method as described in claim 1, characterized in that, Each set of voiceprint attributes in the preset fusion voiceprint library includes a first type of voiceprint attribute for characterizing user characteristics and a second type of voiceprint attribute for characterizing audio characteristics. The method further includes: Obtain the first type of voiceprint attribute and the second type of voiceprint attribute of the target audio; The step of matching the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fused voiceprint library to obtain the matching result includes: The first type of voiceprint attribute of the target audio is matched with the first type of voiceprint attribute in each group of voiceprint attributes; If there are M candidate voiceprint attributes among the N groups of voiceprint attributes that match the first type of voiceprint attribute of the target audio, the second type of voiceprint attribute of the target audio is matched with the second type of voiceprint attribute in each group of candidate voiceprint attributes to obtain the matching result, where M is a positive integer.

6. The method as described in claim 5, characterized in that, The second type of voiceprint attribute includes K seed attributes, where K is a positive integer. The step of matching the second type of voiceprint attribute of the target audio with the second type of voiceprint attribute in each group of candidate voiceprint attributes to obtain the matching result includes: For each group of candidate voiceprint attributes, the similarity of each sub-attribute of the second type of voiceprint attribute in the group of candidate voiceprint attributes is calculated with the corresponding sub-attribute in the target audio to obtain the first similarity of each sub-attribute. Based on the first similarity of each sub-attribute, the matching result between the target audio and the group of candidate voiceprint attributes is determined.

7. The method as described in claim 6, characterized in that, The process of determining the matching result between the target audio and the set of candidate voiceprint attributes based on the first similarity of each sub-attribute includes: Determine the weight of each sub-attribute in the K seed attributes; For each group of candidate voiceprint attributes, based on the weight of each sub-attribute, the first similarity corresponding to each sub-attribute in the group of candidate voiceprint attributes is weighted, and the weighted result is used as the second similarity between the group of candidate voiceprint attributes and the target audio. If there exists a group of P candidate voiceprint attributes with a second similarity greater than a preset similarity, the candidate voiceprint attribute with the largest second similarity among the group of P candidate voiceprint attributes is taken as the target voiceprint attribute that matches the target audio, where P is a positive integer less than or equal to K.

8. The method as described in claim 6 or 7, characterized in that, The sub-attributes of the second type of voiceprint attribute include audio purity, audio resonance, audio speech rate, and audio pitch.

9. The method as described in claim 1, characterized in that, The preset fused voiceprint library further includes at least one fused voiceprint feature code and a user profile corresponding to each fused voiceprint feature code. Each fused voiceprint feature code corresponds to the same user as at least one voiceprint feature code in the preset fused voiceprint library. Controlling the operation of the target device based on the user profile corresponding to the target voiceprint attribute includes: Based on the preset fusion voiceprint library, the voiceprint feature code corresponding to the target voiceprint attribute is determined, and the target fusion voiceprint feature code corresponding to the voiceprint feature code of the target voiceprint attribute is determined from the preset fusion voiceprint library. Based on the user profile corresponding to the target fused voiceprint feature code, the operation of the target device is controlled.

10. The method as described in claim 9, characterized in that, If the matching result indicates that there exists a target voiceprint attribute in the preset fused voiceprint library that matches the voiceprint attribute of the target audio, the method further includes: Associate the target voiceprint feature code of the target audio with the target fused voiceprint feature code.

11. A device control apparatus, characterized in that, include: The acquisition module is used to acquire the target voiceprint feature code fed back by the voice service provider of the target device when the target device detects the target audio. The target device is located in the target area, and the target area contains S devices, where S is a positive integer. The first processing module is configured to, if the target voiceprint feature code is not included in the voiceprint feature codes in the preset fusion voiceprint library, match the voiceprint attributes of the target audio with N sets of voiceprint attributes in the preset fusion voiceprint library to obtain a matching result, wherein the preset fusion voiceprint library includes N voiceprint feature codes fed back by the voice service providers corresponding to the S devices, and voiceprint attributes corresponding to each voiceprint feature code, where N is a positive integer; The second processing module is used to control the operation of the target device based on the user profile corresponding to the target audio if the matching result indicates that there is a target voiceprint attribute in the preset fusion voiceprint library that matches the voiceprint attribute of the target audio.

12. An electrical appliance, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-10.