Method, device, apparatus, and computer-readable medium for generating speech data sets
By obtaining the attribute information of smart devices to generate control text and synthesize voice data sets, the problem of low efficiency of manually generating voice data in the existing technology is solved, and efficient automatic generation of voice data sets is achieved.
Patent Information
- Application Number
- CN202210073238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2042-01-21
AI Technical Summary
The existing technology lacks an efficient method for automatically generating voice data for intelligent device testing, resulting in low efficiency of manual speaking or recording.
By obtaining the attribute information of smart devices, generating control text, and using speech synthesis technology to automatically synthesize speech datasets, the automatic generation of speech datasets is achieved.
It realizes efficient and automatic generation of speech data sets and improves the efficiency of generating test speech data.
Smart Images

Figure CN114446295B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, apparatus, device, and computer-readable medium for generating a speech dataset. Background Art
[0002] Testing the voice control features of smart devices like smart TVs, smart air conditioners, and smart curtains requires a large amount of voice data. For example, to test whether a TV can be turned on properly, a test voice message containing the phrase "Turn on the TV" is required. During testing, a tester typically speaks the test voice message manually or pre-records it.
[0003] However, when using the above method to obtain test speech, the following technical problems often occur:
[0004] Manual speaking or recording is inefficient, and there is a lack of methods to automatically generate voice data for smart device testing. Summary of the Invention
[0005] The content of this disclosure is intended to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions. Some embodiments of the present disclosure propose methods, devices, equipment, and computer-readable media for generating speech datasets to address one or more of the technical issues mentioned in the background section above.
[0006] In a first aspect, some embodiments of the present disclosure provide a method for generating a speech dataset, the method comprising:
[0007] In a second aspect, some embodiments of the present disclosure provide a speech data set device, the device comprising:
[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.
[0010] The above-described embodiments of the present disclosure have the following beneficial effects: they enable the automatic generation of voice datasets. Specifically, the inefficiency of manual speech or recording is due to the lack of methods for automatically generating voice data for smart device testing. Based on this, the present disclosure automatically generates control text based on the attribute information of smart devices and automatically synthesizes a voice dataset based on the control text. This provides a method for automatically generating voice datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0012] Figure 1 is a schematic diagram of an application scenario of a method for generating a speech dataset according to some embodiments of the present disclosure;
[0013] Figure 2 is a flowchart of some embodiments of the method for generating a speech dataset according to the present disclosure;
[0014] Figure 3 An exemplary scenario diagram for generating control text is shown;
[0015] Figure 4 is a flowchart of other embodiments of the method for generating a speech dataset according to the present disclosure;
[0016] Figure 5 is a schematic structural diagram of some embodiments of a speech dataset device according to the present disclosure;
[0017] Figure 6 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] Figure 1 It is a schematic diagram of an application scenario of the method for generating a speech dataset according to some embodiments of the present disclosure.
[0025] like Figure 1 As shown, the execution entity for generating the voice data set can be a computing device 101. Based on this, the computing device 101 can first obtain the attribute information of multiple smart devices (such as smart device 1021, smart device 1022, and smart device 1023 in the figure) through the interface provided by the server, thereby obtaining a smart device attribute information set 103. In practice, each attribute information includes at least one attribute, such as device name, device type, device serial number, device control attribute (such as switch, temperature adjustment), etc. On this basis, the computing device 101 can generate multiple control texts 104 corresponding to each attribute information based on the attribute value corresponding to the at least one attribute and the device control attribute. For example: "Adjust the wind speed to high speed" "Turn on the TV switch", etc. Then, the control text can be synthesized into voice data using speech synthesis technology to obtain a voice data set 105.
[0026] Alternatively, you can use the generated voice dataset by playing it through a voice device. This allows the smart speaker to receive the played voice data. The smart speaker then converts the received voice data into execution instructions and sends them to the server. The server then sends the instructions to the gateway, which then forwards them to the corresponding smart device in the gateway. This allows testing to be performed based on the smart device's response.
[0027] It should be noted that the computing device 101 can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitations are given here.
[0028] Continue to refer Figure 2 , shows a process 200 of some embodiments of the method for generating a speech dataset according to the present disclosure. The method for generating a speech dataset comprises the following steps:
[0029] Step 201: Acquire attribute information of multiple smart devices through an interface provided by the server to obtain a smart device attribute information set, wherein each attribute information includes at least one attribute, and the at least one attribute includes a device control attribute.
[0030] In some embodiments, the execution entity of the voice data set generation method (e.g., electronic device 101) can obtain attribute information of multiple smart devices through an interface provided by the server. Among them, smart devices can support voice control devices, including but not limited to: smart TVs, smart air conditioners, smart curtains, etc. In practice, multiple smart devices can be bound through the mobile terminal application as needed. On this basis, the server will provide an interface for obtaining attribute information of each smart home device.
[0031] The attribute information for each smart device may include, but is not limited to, the device name, device serial number, device ID, and device control attributes. Device control attributes may be attributes related to device control, including but not limited to at least one of the following: on / off, air conditioning temperature, main light brightness, spotlight color, etc. The specific content of the attribute information for each smart device may be the same or different, and can be selected based on actual needs.
[0032] Optionally, in practice, to extract valid attributes, at least one attribute included in each acquired attribute information may be filtered, thereby extracting valid attributes. For example, one or more device control-related attributes such as the device name, device control attribute, and the attribute value of the control attribute may be extracted.
[0033] In practice, the attribute information of a smart device may include one device control attribute or multiple device control attributes. For example, the attribute information of a smart TV may include the device control attributes "switch" and "volume".
[0034] Step 202: For the attribute information in the smart device attribute information set, generate multiple control texts corresponding to the attribute information based on at least one attribute included in the attribute information and the attribute value corresponding to the device control attribute included in the at least one attribute, wherein each attribute value corresponds to at least one control text.
[0035] In some embodiments, for a certain attribute information in the smart device attribute information set, the above-mentioned execution entity can generate multiple control texts corresponding to the attribute information based on at least one attribute included in the attribute information and the attribute value corresponding to the device control attribute.
[0036] Optionally, the device name, device control attributes, and attribute values can be input into a pre-trained text generation model to generate control text. The network structure of the text generation model can be a recurrent neural network (RNN), a long short-term memory network (LSTM), or other similar architectures. Training samples for the text generation model can include keywords and sentences containing keywords. Machine learning methods can be used to train the text generation model using these training samples.
[0037] Optionally, for each attribute value corresponding to a device control attribute included in the attribute information, multiple control texts corresponding to the attribute information are generated based on the attribute value, the device control attribute corresponding to the attribute value, and the device name included in at least one attribute. For example, the device name, device control attribute, and attribute value can be entered into a preset text template to generate the control text. Multiple text templates can be pre-set as needed to generate multiple control texts for a single attribute value.
[0038] As an example, Figure 3 As shown, for smart devices, the smart fresh air panel 301 is taken as an example. The device control attribute 302 included therein is "wind speed", which corresponds to three attribute values, namely attribute value 3021 ("high speed"), attribute value 3022 ("medium speed"), and attribute value 3023 ("low speed"). On this basis, for each attribute value, such as attribute value 3021 ("high speed"), at least one corresponding control text can be generated. As an example, the device name, device control attribute, and attribute value can be filled into the text template according to a preset text template. For example, the text template is "Set the [device control attribute] of [device name] to [attribute value]". Thus, three control texts corresponding to the device control attribute 302 being "wind speed" can be obtained, namely "Set the wind speed of the smart fresh air panel to high speed", "Set the wind speed of the smart fresh air panel to medium speed", and "Set the wind speed of the smart fresh air panel to low speed". On this basis, if there are multiple device control attributes, multiple control texts corresponding to each device control attribute can be generated in a similar manner.
[0039] In some optional implementations of some embodiments, generating multiple control texts corresponding to the attribute information based on the attribute value, the device control attribute corresponding to the attribute value, and the device name included in at least one attribute includes: generating standard control text based on the attribute value, the device control attribute corresponding to the attribute value, and the device name included in at least one attribute; replacing the attribute value in the standard control text with associated words to obtain associated control text; and determining the standard control text and the associated control text as the multiple control texts corresponding to the attribute information. In these optional implementations, generating associated control texts can make the control texts more diverse.
[0040] Step 203: Perform speech synthesis on each control text to obtain a speech data set.
[0041] In some embodiments, the execution entity may perform speech synthesis on each control text to obtain a speech data set. Various speech synthesis (TTS, Text To Speech) algorithms may be used to synthesize the speech data.
[0042] Some embodiments of the present disclosure provide methods that automatically generate control text using attribute information of a smart device and automatically synthesize a speech dataset based on the control text, thereby providing a method for automatically generating a speech dataset.
[0043] Further references Figure 4 , which shows a process 400 of another embodiment of a method for generating a speech dataset. The process 400 of the method for generating a speech dataset includes the following steps:
[0044] Step 401: Acquire attribute information of multiple smart devices through an interface provided by the server to obtain a smart device attribute information set, wherein each attribute information includes at least one attribute, and the at least one attribute includes a device control attribute.
[0045] In some embodiments, the specific implementation of step 401 and the technical effects thereof can be referred to in Figure 2 Step 201 in the corresponding embodiment will not be described again here.
[0046] Step 402: Divide the smart device attribute information set according to the device control attributes to obtain multiple smart device attribute information subsets.
[0047] In some embodiments, the execution subject of the voice data set generation method can divide the smart device attribute information set according to the number of attribute values corresponding to the device control attributes to obtain multiple smart device attribute information subsets. For example, for the attribute information in the smart device attribute information set, if the number of attribute values corresponding to each attribute in at least one attribute included in the attribute information is two, the attribute information is added to the first smart device attribute information subset; if the number of attribute values corresponding to each attribute in at least one attribute included in the attribute information is greater than two, the attribute information is added to the second smart device attribute information subset. In practice, if it is a switch-type attribute, the number of corresponding attribute values is generally two, namely "on" and "off". If it is a non-switch-type attribute such as temperature, brightness, etc., the number of corresponding attribute values is generally greater than two.
[0048] Optionally, the execution entity may also implement the division of the smart device attribute information set through keyword recognition. For a particular attribute information, it may be matched against multiple preset keywords. Each keyword corresponds to a subset of the smart device attribute information. Based on this, if the attribute information successfully matches the target keyword, the attribute information may be added to the subset of the smart device attribute information corresponding to the target keyword.
[0049] Step 403: For the attribute information in each smart device attribute information subset among the multiple smart device attribute information subsets, the attribute values corresponding to at least one attribute included in the attribute information and the device control attribute included in at least one attribute are input into the model corresponding to the smart device attribute information subset to generate multiple control texts corresponding to the attribute information.
[0050] In some embodiments, different subsets of smart device attribute information may correspond to different models. Models may be used to generate control text, including but not limited to regular expressions, artificial neural networks, and the like. For examples of models, see the description of step 202. In practice, different subsets of smart device attribute information may correspond to different scenarios, thereby enabling differentiated control text generation for different scenarios.
[0051] Step 404: Perform speech synthesis on each control text to obtain a speech data set.
[0052] In some embodiments, the specific implementation of step 404 and the technical effects thereof can be referred to in Figure 2 The corresponding step 203 in the embodiments will not be described in detail here.
[0053] from Figure 4 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 4The process 400 of the speech data set generation method in some corresponding embodiments adds the steps of dividing the smart device attribute information set and generating control text according to the corresponding model, thereby achieving differentiated control text generation.
[0054] Further references Figure 5 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a speech data set device. These device embodiments are similar to Figure 2 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0055] like Figure 5 As shown, the speech data set apparatus 500 of some embodiments includes: an acquisition unit 501, a generation unit 502, and a synthesis unit 503. The acquisition unit 501 is configured to acquire attribute information of multiple smart devices through an interface provided by the server to obtain a smart device attribute information set, wherein each attribute information includes at least one attribute, and at least one attribute includes a device control attribute. The generation unit 502 is configured to generate, for the attribute information in the smart device attribute information set, multiple control texts corresponding to the attribute information based on at least one attribute included in the attribute information and the attribute value corresponding to the device control attribute included in the at least one attribute, wherein each attribute value corresponds to at least one control text. The synthesis unit 503 is configured to perform speech synthesis on each control text to obtain a speech data set.
[0056] In an optional implementation of some embodiments, the apparatus further includes: a division unit configured to divide the smart device attribute information set according to device control attributes to obtain a plurality of smart device attribute information subsets.
[0057] In an optional implementation of some embodiments, the generation unit 502 is further configured to input the attribute information in each smart device attribute information subset in multiple smart device attribute information subsets into a model corresponding to the smart device attribute information subset, including at least one attribute included in the attribute information and the attribute value corresponding to the device control attribute included in at least one attribute, to generate multiple control texts corresponding to the attribute information.
[0058] In an optional implementation of some embodiments, the generation unit 502 is further configured to generate multiple control texts corresponding to the attribute information for each attribute value corresponding to the device control attribute included in the attribute information, based on the attribute value, the device control attribute corresponding to the attribute value, and the device name included in at least one attribute.
[0059] In some optional implementations of the embodiments, the generation unit 502 is further configured to: generate a standard control text based on the attribute value, the device control attribute corresponding to the attribute value, and the device name included in at least one attribute; replace the attribute value in the standard control text with associated words to obtain associated control text; and determine the standard control text and the associated control text as multiple control texts corresponding to the attribute information.
[0060] In an optional implementation of some embodiments, the division unit is configured to, for the attribute information in the smart device attribute information set, add the attribute information to the first smart device attribute information subset if the number of attribute values corresponding to each attribute in at least one attribute included in the attribute information is two; and add the attribute information to the second smart device attribute information subset if the number of attribute values corresponding to each attribute in at least one attribute included in the attribute information is greater than two.
[0061] It is understood that the units described in the device 500 are similar to those in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 500 and the units included therein, and will not be repeated here.
[0062] Reference below Figure 6 , which shows an electronic device (eg, Figure 1 A structural diagram of a server or terminal device in (a) 600. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0063] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0064] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0065] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0066] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0067] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0068] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: obtains attribute information of multiple smart devices through an interface provided by the server to obtain a smart device attribute information set, wherein each attribute information includes at least one attribute, and at least one attribute includes a device control attribute; for the attribute information in the smart device attribute information set, generates multiple control texts corresponding to the attribute information based on the attribute values corresponding to at least one attribute included in the attribute information and the device control attribute included in the at least one attribute, wherein each attribute value corresponds to at least one control text; and performs speech synthesis on each control text to obtain a speech data set.
[0069] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0070] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0071] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor. For example, they may be described as: a processor including an acquisition unit, a generation unit, and a synthesis unit. The names of these units do not, in some cases, constitute a limitation on the units themselves. For example, the acquisition unit may also be described as a "unit for acquiring attribute information of multiple smart devices through an interface provided by the server."
[0072] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0073] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for generating a speech dataset, comprising: Acquire attribute information of multiple smart devices through an interface provided by the server to obtain a smart device attribute information set, wherein each attribute information includes at least one attribute, and the at least one attribute includes a device control attribute; For the attribute information in the smart device attribute information set, input the device name and device control attributes included in the attribute information, as well as the attribute values corresponding to the device control attributes, into a pre-trained text generation model to generate a plurality of control texts corresponding to the attribute information, wherein each attribute value corresponds to at least one control text; Perform speech synthesis on each control text to obtain a speech dataset.
2. The method according to claim 1, wherein Before inputting the device name and device control attributes included in the attribute information in the smart device attribute information set, and the attribute values corresponding to the device control attributes, into a pre-trained text generation model to generate a plurality of control texts corresponding to the attribute information, the method further includes: The smart device attribute information set is divided according to the device control attribute to obtain a plurality of smart device attribute information subsets.
3. The method according to claim 2, wherein: For the attribute information in the smart device attribute information set, the device name and device control attributes included in the attribute information, as well as the attribute values corresponding to the device control attributes, are input into a pre-trained text generation model to generate multiple control texts corresponding to the attribute information, including: For the attribute information in each smart device attribute information subset among the multiple smart device attribute information subsets, the attribute values corresponding to at least one attribute included in the attribute information and the device control attribute included in the at least one attribute are input into the model corresponding to the smart device attribute information subset to generate multiple control texts corresponding to the attribute information.
4. The method according to claim 3, wherein: The multiple control texts are generated by the following steps: generating a standard control text based on the attribute value, the device control attribute corresponding to the attribute value, and the device name included in the at least one attribute; Performing associated word replacement on the attribute value in the standard control text to obtain associated control text; The standard control text and the associated control text are determined as a plurality of control texts corresponding to the attribute information.
5. The method according to claim 2, wherein: The smart device attribute information set is divided according to the device control attribute to obtain multiple smart device attribute information subsets, including: For the attribute information in the smart device attribute information set, if the number of attribute values corresponding to each attribute in at least one attribute included in the attribute information is two, adding the attribute information to the first smart device attribute information subset; If the number of attribute values corresponding to each attribute in at least one attribute included in the attribute information is greater than two, the attribute information is added to the second smart device attribute information subset.
6. A speech data set generation device, comprising: an acquiring unit configured to acquire attribute information of a plurality of smart devices through an interface provided by the server to obtain a set of smart device attribute information, wherein each attribute information includes at least one attribute, and the at least one attribute includes a device control attribute; a generating unit configured to input, for attribute information in the smart device attribute information set, a device name and device control attributes included in the attribute information, and attribute values corresponding to the device control attributes, into a pre-trained text generation model to generate a plurality of control texts corresponding to the attribute information, wherein each attribute value corresponds to at least one control text; The synthesis unit is configured to perform speech synthesis on each control text to obtain a speech data set.
7. The device according to claim 6, wherein The device further comprises: The dividing unit is configured to divide the smart device attribute information set according to the device control attribute to obtain a plurality of smart device attribute information subsets.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Intelligent household control method and intelligent household control system
CN107688329A
Intelligent household electrical appliance intelligent level test system and method based on voice interaction
CN112383451A