System construction, information recording, model training method and device, equipment and medium
By constructing a simulated voice recording system, the problems of insufficient convenience and stability in voice information recording were solved, enabling efficient voice information recording even when acquisition is inconvenient, and improving the flexibility and matching of the recording process.
Patent Information
- Application Number
- CN202111626379.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2041-12-28
AI Technical Summary
Existing technologies for recording voice information in voice systems suffer from insufficient convenience and stability, making it difficult to record voice information efficiently, especially when access is inconvenient.
By constructing a simulated voice recording system, determining the simulated voice modules and their location distribution, and building the simulated voice recording system within the simulated space, the system simulates voice output and signal acquisition to achieve simulated voice information recording.
It improves the convenience and stability of voice recording, reduces labor costs, avoids information discrepancies caused by differences in recording environments, and enhances the flexibility and compatibility of the recording process.
Smart Images

Figure CN114220422B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of artificial intelligence technology such as voice technology and vehicle networking technology. Specifically, it relates to a method for constructing a voice recording system, a method for recording voice information, a method for training a voice recognition model, a device for constructing a voice recording system, a device for recording voice information, a device for training a voice recognition model, an electronic device, and a non-transient computer-readable storage medium. Background Technology
[0002] With the development of artificial intelligence technology, voice systems are being widely applied in various fields. For example, in-vehicle terminals can be used for voice recognition. To enable a voice system to have voice recognition capabilities, it is usually necessary to collect a large amount of voice information for the system to learn from. Summary of the Invention
[0003] This disclosure provides a method for constructing a speech recording system, a method for recording speech information, a method for training a speech recognition model, a device for constructing a speech recording system, a device for recording speech information, a device for training a speech recognition model, an electronic device, and a non-transitory computer-readable storage medium.
[0004] According to one aspect of this disclosure, a method for constructing a voice recording system is provided, comprising:
[0005] Based on the system type of the voice system to be simulated, determine the simulation voice module of the simulation voice recording system and the simulation position distribution between different simulation voice modules; wherein, the simulation voice module includes a sound acquisition module and a sound output module;
[0006] Based on the simulated speech module and the simulated location distribution, a simulated speech recording system is constructed in the simulated space for recording simulated speech information.
[0007] According to another aspect of this disclosure, a method for recording voice information is also provided, comprising:
[0008] According to the voice recording requirements, the sound output module of the simulated voice recording system is controlled to output a sound signal; wherein, the simulated voice recording system is constructed based on any of the voice recording system construction methods provided in the embodiments of this disclosure.
[0009] The sound acquisition module in the simulated speech recording system is controlled to acquire sound signals and obtain simulated speech information.
[0010] According to another aspect of this disclosure, a method for training a speech recognition model is also provided, comprising:
[0011] Acquire simulated voice information; wherein, the simulated voice information is acquired based on any of the voice information recording methods provided in the embodiments of this disclosure;
[0012] Based on the simulated speech information, the speech recognition model in the simulated speech system is trained.
[0013] According to another aspect of this disclosure, an electronic device is also provided, comprising:
[0014] At least one processor; and
[0015] A memory that is communicatively connected to at least one processor; wherein,
[0016] The memory stores instructions that can be executed by at least one processor, which enables the at least one processor to perform any one of the speech recording system construction method, speech information recording method, and speech recognition model training method provided in the embodiments of this disclosure.
[0017] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause a computer to execute any one of the speech recording system construction method, speech information recording method, and speech recognition model training method provided in the embodiments of this disclosure.
[0018] The technology disclosed herein improves the convenience and stability of voice information recording.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0020] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0021] Figure 1 This is a flowchart of a method for constructing a voice recording system provided in an embodiment of this disclosure;
[0022] Figure 2 This is a flowchart of a voice information recording method provided in an embodiment of this disclosure;
[0023] Figure 3 This is a flowchart of a speech recognition model training method provided in an embodiment of this disclosure;
[0024] Figure 4 This is a structural diagram of a voice recording system construction device provided in an embodiment of this disclosure;
[0025] Figure 5 This is a structural diagram of a voice information recording device provided in an embodiment of this disclosure;
[0026] Figure 6 This is a structural diagram of a speech recognition model training device provided in an embodiment of this disclosure;
[0027] Figure 7 This is a block diagram of an electronic device used to implement the speech recording system construction method, speech information recording method, or speech recognition model training method of the embodiments of this disclosure. Detailed Implementation
[0028] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0029] The speech recording system construction method provided in this disclosure is applicable to the construction scenario of a simulated speech recording system for recording simulated speech information. The speech recording system construction methods provided in this disclosure can be executed by a speech recording system construction device, which can be implemented in software and / or hardware and specifically configured in an electronic device.
[0030] To facilitate understanding, the construction method of the voice recording system will be explained in detail first.
[0031] See Figure 1 The method for constructing a voice recording system, as shown, includes:
[0032] S101. Based on the system type of the voice system to be simulated, determine the simulation voice module of the simulation voice recording system and the simulation position distribution between different simulation voice modules; wherein, the simulation voice module includes a sound acquisition module and a sound output module.
[0033] The voice system to be simulated can be understood as a closed system with voice recording and voice recognition capabilities. For example, the voice system to be simulated can be a vehicle equipped with an in-vehicle terminal, which includes a sound acquisition module (such as a microphone) and a sound output module (such as a speaker).
[0034] The system type is used to uniquely identify the voice system to be simulated. Continuing the previous example, if the voice system to be simulated is a vehicle, the system type can be a vehicle model identifier or a vehicle series identifier, etc.
[0035] Among them, the simulated position distribution is used to characterize the relative positional relationship between different simulated speech modules.
[0036] The simulated speech module includes a sound output module, which simulates the sound output module in the simulated speech system to output sound signals; and a sound acquisition module, which simulates the sound acquisition module in the simulated speech system to acquire sound signals. The sound output module and the sound acquisition module in the simulated speech module can be the same as or different from those in the simulated speech system. It is understood that, to ensure compatibility between the simulated speech recording system and the simulated speech system, the sound output module and the sound acquisition module in the simulated speech module are typically at least partially identical to each other.
[0037] It should be noted that this disclosure does not impose any restrictions on the types of sound acquisition modules or sound output modules in the simulated voice module.
[0038] In one specific implementation, the sound acquisition module in the simulated voice module can be a microphone.
[0039] In another specific implementation, the sound output module in the simulated voice module may include a speaker to output noise information, thereby simulating the internal noise environment of the voice system to be simulated and providing data support for subsequent recording of simulated voice information. For example, the noise information may be music, navigation voice, or radio audio.
[0040] In another specific implementation, the sound output module in the simulated voice module may include a simulated human head, which is used to replace a real person to output sound information, thereby improving the stability of the output sound information, providing data support for the subsequent recording of simulated voice information, and reducing labor costs.
[0041] For example, if the voice system to be simulated is a vehicle, the simulated human head can correspond to a seat in the vehicle to simulate the driver or passenger's voice signal output. Similarly, if the voice system to be simulated is a system built from devices within the smart speaker's placement area, the simulated human head can correspond to home appliances within the smart speaker's placement area to simulate the noise signal output of those appliances.
[0042] In one optional embodiment, a simulation location distribution map or table can be pre-stored. This map or table stores the number of simulated speech modules corresponding to different system types, as well as the simulation location distribution among the different simulated speech modules. Accordingly, during the construction of the speech recording system, based on the system type of the speech system to be simulated, the number of each simulated speech module and the simulation location distribution among the different simulated speech modules are determined by querying the simulation location distribution map or table.
[0043] In another optional embodiment, the actual positional distribution between different actual voice modules in the voice device to be simulated can be determined according to the system type of the voice system to be simulated; and the simulated positional distribution between different simulated voice modules can be determined according to the actual positional distribution; wherein, the simulated voice module corresponds one-to-one with the actual voice module.
[0044] In this context, the actual speech module can be understood as the sound acquisition module and sound output module in the speech system to be simulated; the actual position distribution is used to characterize the relative positional relationships between different actual speech modules in the speech system to be simulated. The actual speech module can include a sound output module and a sound acquisition module. The sound output module can be a sound-generating entity, such as a driver, passenger, or sound-generating device in a vehicle.
[0045] Understandably, since the actual positional distribution directly reflects the relative positional relationships between different actual speech modules in the speech system to be simulated, and this relative positional relationship is already determined during the construction of the speech system to be simulated, determining the actual positional distribution directly based on the system type of the speech system to be simulated is not affected by other factors, resulting in higher accuracy and better stability. Correspondingly, based on this actual positional distribution, determining the simulation positional distribution between different simulated speech modules allows for differentiated determination of simulation positional distribution according to different actual positional distributions, improving the accuracy of the simulation positional distribution determination results. By setting a one-to-one matching relationship between simulated speech modules and actual speech modules, the matching degree between the subsequently constructed simulated speech recording system and the speech system to be simulated is improved.
[0046] Optionally, based on the actual position distribution, the simulated position distribution between different simulated speech modules can be determined. This can be achieved by directly converting the actual position distribution between different actual speech modules into the simulated position distribution between the corresponding different simulated speech modules. This restores the position distribution of the speech modules between the constructed simulated speech recording system and the speech system to be simulated, which helps to improve the matching degree between the simulated speech recording system and the speech system to be simulated.
[0047] Alternatively, the simulation position distribution between different simulated speech modules can be determined based on the actual position distribution. This can be achieved by: determining the scaling ratio based on the size of the simulation space; and determining the simulation position distribution between different simulated speech modules based on the scaling ratio and the actual position distribution.
[0048] The simulation space can be understood as the construction space for the simulated voice recording system, providing a recording environment for the system. In a specific example, the simulation space can be an indoor space.
[0049] For example, determining the scaling ratio based on the size of the simulation space can be achieved by: pre-setting the correspondence between the size of different simulation spaces and the scaling ratio, and then using this correspondence to find and determine the scaling ratio.
[0050] Since the simulated speech recording system is used to spatially simulate the speech system to be simulated, the spatial ratio can be determined based on the size of the simulated space and the size of the speech system to be simulated; and the scaling ratio can be determined based on the spatial ratio. Accordingly, the actual distance information in the actual position distribution between different actual speech modules is weighted according to the scaling ratio to obtain the simulated distance information of the simulated position distribution between the corresponding different simulated speech modules.
[0051] For example, determining the scaling ratio based on the spatial proportion can be done by: determining the scaling ratio based on a preset scaling ratio determination function, where the independent variable of the preset scaling ratio determination function is the spatial proportion, and the dependent variable is the scaling ratio; the preset scaling ratio determination function is an increasing function of the spatial proportion.
[0052] Alternatively, for example, determining the scaling ratio based on the spatial ratio can be achieved by: pre-setting the correspondence between different spatial ratios and scaling ratios, and using this correspondence to find and determine the scaling ratio corresponding to the spatial ratio.
[0053] It is understandable that by introducing the size of the simulation space and determining the scaling ratio, and then establishing a mapping relationship between the actual position distribution and the simulated position distribution based on the scaling ratio, it is possible to construct simulation voice recording systems corresponding to different voice systems to be simulated, even when the simulation space is single or limited, thereby improving the flexibility and universality of the simulation voice recording system construction process.
[0054] S102. Based on the distribution of the simulated voice module and the simulated position, construct a simulated voice recording system in the simulated space for recording simulated voice information.
[0055] For example, the position information of each simulated voice module in the simulation space can be determined according to the simulation position distribution of different simulated voice modules; and each simulated voice module can be set to the corresponding position information to obtain a simulated voice recording system for recording simulated voice information.
[0056] This embodiment of the disclosure constructs a simulated voice recording system based on the system type of the voice system to be simulated, determines the simulated voice modules, and establishes the simulated position distribution among different simulated voice modules. Using the simulated voice modules and their simulated position distribution as reference data, the construction of the simulated voice recording system ensures compatibility between the simulated voice recording system and the voice system to be simulated. Furthermore, by using the simulated voice recording system to replace the voice system to be simulated for recording simulated voice information, it enables voice information recording even when obtaining the voice system to be simulated is inconvenient, thus improving the convenience of voice information recording. Since the simulated voice recording system operates in a simulated space, it is less affected by real environmental factors, thereby also improving the stability of voice information recording.
[0057] Based on the above technical solutions, this disclosure also provides a voice information recording method, applicable to scenarios where simulated voice information is recorded using the aforementioned constructed simulated voice recording system. The voice information recording method provided in this disclosure can be executed using a voice information recording device, which can be implemented in software and / or hardware and specifically configured in an electronic device. This electronic device and the electronic device executing the aforementioned voice recording system construction method can be the same or different, and this disclosure does not impose any limitations in this regard.
[0058] See Figure 2 A method for recording voice information, as shown, includes:
[0059] S201. Based on the voice recording requirements, control the sound output module of the simulated voice recording system to output the sound signal.
[0060] The simulated voice recording system is constructed based on any of the voice recording system construction methods provided in the embodiments of this disclosure.
[0061] Optionally, the voice recording requirements may include recording content requirements, such as at least one of the following: text topic, text keywords, text content, and recording duration in the voice information to be recorded. Alternatively, the voice recording requirements may include the direction of the recording sound source, used to characterize the directional position of the sound source emitting the voice information to be recorded in the simulation voice recording system within the simulation space. Alternatively, the voice recording information may include the direction of noise interference, used to characterize the directional position of the sound source corresponding to the noise information when noise interference is introduced into the simulation voice recording system. Alternatively, the voice recording information may include the intensity of the recorded sound source, used to characterize the signal strength of the sound signal corresponding to the voice information to be recorded emitted in the simulation voice recording system. Alternatively, the voice recording information may include the intensity of noise interference, used to characterize the signal strength of the noise interference signal emitted in the simulation voice recording system.
[0062] For example, at least one target audio output module can be selected from the audio output modules of the simulated audio recording system according to the audio recording requirements; the target audio output module is then controlled to output the audio signal corresponding to the audio recording requirements. The audio signal includes the audio signal to be recorded and / or noise interference signals. For instance, when recording the audio message "turn on music," the audio signal corresponding to the background conversation is the noise interference signal; the audio signal corresponding to "turn on music" is the audio signal to be recorded corresponding to the audio message to be recorded.
[0063] In an optional embodiment, if the voice recording requirement includes recording the direction of the sound source, then controlling the sound output module in the simulated voice recording system to output voice information according to the voice recording requirement may include: controlling the sound output module in the simulated voice recording system corresponding to the direction of the sound source to output the sound signal to be recorded.
[0064] For example, a directional module correspondence between different sound output modules and the direction of sound emission can be established in advance in the simulated voice recording system; accordingly, based on the directional module correspondence, the sound output module corresponding to the direction of the recorded sound source is searched to obtain the first sound output module; the first sound output module is controlled to output the sound signal to be recorded.
[0065] In another optional embodiment, if the voice recording requirement includes the direction of noise interference, then controlling the sound output module in the simulated voice recording system to output voice information according to the voice recording requirement may include: controlling the sound output module in the simulated voice recording system corresponding to the direction of noise interference to output a noise interference signal.
[0066] For example, based on the aforementioned correspondence between direction modules, the sound output module corresponding to the noise interference direction can be searched to obtain the second sound output module; the second sound output module can then be controlled to output a noise interference signal.
[0067] It is understandable that by introducing the recording sound direction and / or noise interference direction, the corresponding sound output module is controlled to output the sound signal to be recorded and / or noise interference signal, thereby improving the richness of the sound signal output by the sound output module, which helps to improve the richness of the simulated speech information recorded subsequently.
[0068] In one alternative embodiment, the signal strength of the audio signal to be recorded can be preset to a fixed value.
[0069] To enhance the richness of the recorded audio signal from the perspective of signal strength, and thus improve the richness of the recorded simulated speech information, in another optional embodiment, the speech recording requirement may further include recording the intensity of the sound source; correspondingly, controlling the sound output module corresponding to the direction of the recorded sound source in the simulated speech recording system to output the audio signal to be recorded may include: controlling the sound output module corresponding to the direction of the recorded sound source in the simulated speech recording system to output the audio information to be recorded based on the intensity of the recorded sound source.
[0070] In another alternative embodiment, the signal strength of the noise interference signal can be preset to a fixed value.
[0071] To enhance the richness of noise interference signals from the signal strength dimension, thereby increasing the richness of the recorded simulated speech information, in another optional embodiment, the speech recording requirement may further include noise interference intensity; correspondingly, controlling the sound output module corresponding to the noise interference direction in the simulated speech recording system to output the noise interference signal may include: controlling the sound output module corresponding to the noise interference direction in the simulated speech recording system to output the noise interference signal according to the noise interference intensity.
[0072] S202. Control the sound acquisition module in the simulated voice recording system to acquire sound signals and obtain simulated voice information.
[0073] The sound acquisition module in the control simulation speech recording system acquires the sound signal output by the sound output module, that is, the sound signal to be recorded and / or noise interference signal, and synthesizes the recorded sound signals to obtain simulated speech information.
[0074] This embodiment of the disclosure simulates a speech system to be simulated using a simulated speech recording system. By recording simulated speech information within the simulated speech recording system, it eliminates the need to acquire the speech system to be simulated, thus enabling speech information recording when acquiring the speech system to be simulated is inconvenient, thereby improving the convenience of speech information recording. Furthermore, by incorporating speech recording requirements to control the sound signal output and acquisition of the simulated speech recording system, the flexibility of the simulated speech information recording process is improved. Simultaneously, the automated recording of simulated speech information reduces labor costs and avoids variations in recorded speech information caused by differences in personnel or recording environments, thereby improving the stability of the simulated speech information.
[0075] Based on the above technical solutions, this disclosure also provides a speech recognition model training method, applicable to application scenarios where speech recognition models are trained based on speech information recorded by a simulated speech recording system. The speech recognition model training methods provided in the embodiments of this disclosure can be executed by a speech recognition model training device, which can be implemented in software and / or hardware and specifically configured in an electronic device. This electronic device is typically configured in the speech system to be simulated.
[0076] See Figure 3 The method for training a speech recognition model shown includes:
[0077] S301. Obtain simulated voice information.
[0078] The simulated voice information is collected based on any of the voice information recording methods provided in the embodiments of this disclosure.
[0079] S302. Train the speech recognition model in the simulated speech system based on the simulated speech information.
[0080] Simulated speech information is used as training samples to train the speech recognition model in the simulated speech system until the model training cutoff condition is met, thereby enabling the speech recognition model in the simulated speech system to have speech recognition capabilities.
[0081] The model training cutoff conditions can be: the amount of simulated speech information reaches a set threshold, the trained speech recognition model becomes stable, or the accuracy of the trained speech recognition model meets a set accuracy threshold. The specific values for the set quantity threshold and the set accuracy threshold can be determined by technical personnel based on needs or experience, or through extensive experimentation.
[0082] It should be noted that this disclosure does not impose any restrictions on the specific network structure of the speech recognition model.
[0083] This embodiment uses simulated speech information recorded by a simulated speech recording system as training samples to train a speech recognition model in the simulated speech system. This eliminates the need to acquire the simulated speech system during the training sample preparation stage, reducing the difficulty of obtaining training samples and thus lowering the time cost of the training sample preparation stage. This improves the training efficiency of the speech recognition model. Furthermore, since the acquired simulated speech information is automatically recorded by the simulated speech recording system, the stability of the simulated speech information is ensured, thereby improving the accuracy of the speech recognition model training results.
[0084] Based on the above technical solutions, since the recording environment of the simulated speech recording system and the real environment of the speech system to be simulated have certain inherent differences, the speech recognition model trained with simulated speech information may not be well adapted to the real environment of the speech system to be simulated. To further improve the accuracy and stability of the speech recognition model, in an optional embodiment, online speech information from the speech system to be simulated can also be collected; based on the online speech information, the trained speech recognition model can be retrained.
[0085] Online voice information refers to voice information recorded in the real environment of the voice system to be simulated.
[0086] It is understandable that using online speech information as online training samples for the aforementioned trained speech recognition model, and then performing secondary training on the previously trained speech recognition model to fine-tune the network parameters of the speech recognition model, helps to improve the adaptability of the trained speech recognition model to the speech system to be simulated, thereby improving the accuracy and stability of the speech recognition results of the speech recognition model.
[0087] Based on the above technical solutions, this disclosure also provides an optional embodiment. In this optional embodiment, the process of using the aforementioned trained speech recognition model is described.
[0088] Optionally, the speech information to be tested can be obtained; the speech information to be tested can be used as input data and input into the speech recognition model trained by the aforementioned speech recognition model training method to obtain the speech prediction result; the speech recognition model can be evaluated based on the speech prediction result and the standard result of the speech information to be tested.
[0089] Alternatively, the speech information to be recognized can be obtained; the speech information to be recognized can be used as input data and input into the speech recognition model trained using the aforementioned speech recognition model training method to obtain the speech recognition result.
[0090] As an implementation of the above-described methods for constructing voice recording systems, this disclosure also provides an optional embodiment of an execution device for implementing the voice recording system construction methods. See further details. Figure 4 The voice recording system construction device 400 shown includes: a simulation position distribution determination module 401 and a simulation system construction module 402. Among them,
[0091] The simulation position distribution determination module 401 is used to determine the simulation position distribution between the simulation voice modules of the simulation voice recording system and different simulation voice modules according to the system type of the voice system to be simulated; wherein, the simulation voice module includes a sound acquisition module and a sound output module;
[0092] The simulation system construction module 402 is used to construct the simulation voice recording system in the simulation space according to the simulation voice module and the simulation position distribution, for simulating voice information recording.
[0093] This embodiment of the disclosure constructs a simulated voice recording system based on the system type of the voice system to be simulated, determines the simulated voice modules, and establishes the simulated position distribution among different simulated voice modules. Using the simulated voice modules and their simulated position distribution as reference data, the construction of the simulated voice recording system ensures compatibility between the simulated voice recording system and the voice system to be simulated. Furthermore, by using the simulated voice recording system to replace the voice system to be simulated for recording simulated voice information, it enables voice information recording even when obtaining the voice system to be simulated is inconvenient, thus improving the convenience of voice information recording. Since the simulated voice recording system operates in a simulated space, it is less affected by real environmental factors, thereby also improving the stability of voice information recording.
[0094] In an optional embodiment, the simulation location distribution determination module 401 includes:
[0095] The actual location distribution determination unit is used to determine the actual location distribution between different actual voice modules in the voice device to be simulated based on the system type of the voice system to be simulated.
[0096] The simulation position distribution determination unit is used to determine the simulation position distribution between different simulation voice modules based on the actual position distribution;
[0097] The simulated voice module corresponds one-to-one with the actual voice module.
[0098] In an optional embodiment, the simulation location distribution determination unit includes:
[0099] The scaling ratio determination subunit is used to determine the scaling ratio based on the size of the simulation space.
[0100] The simulation position distribution determination subunit is used to determine the simulation position distribution between different simulation voice modules based on the scaling ratio and the actual position distribution.
[0101] In an alternative embodiment, the sound output module includes a speaker and / or a simulated human head.
[0102] The above-described speech recording system construction apparatus can execute the speech recording system construction method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing each speech recording system construction method.
[0103] As an implementation of the above-described voice information recording methods, this disclosure also provides an optional embodiment of an execution device for implementing the voice information recording methods. See further details. Figure 5 The voice information recording device 500 shown includes: a sound information output module 501 and a simulated voice information acquisition module 502. Among them,
[0104] The sound information output module 501 is used to control the sound output module of the simulated voice recording system to output sound signals according to the voice recording requirements; wherein, the simulated voice recording system is constructed based on any of the voice recording system construction devices provided in the embodiments of this disclosure;
[0105] The simulated voice information acquisition module 502 is used to control the sound acquisition module in the simulated voice recording system to acquire the sound signal and obtain simulated voice information.
[0106] This embodiment of the disclosure simulates a speech system to be simulated using a simulated speech recording system. By recording simulated speech information within the simulated speech recording system, it eliminates the need to acquire the speech system to be simulated, thus enabling speech information recording when acquiring the speech system to be simulated is inconvenient, thereby improving the convenience of speech information recording. Furthermore, by incorporating speech recording requirements to control the sound signal output and acquisition of the simulated speech recording system, the flexibility of the simulated speech information recording process is improved. Simultaneously, the automated recording of simulated speech information reduces labor costs and avoids variations in recorded speech information caused by differences in personnel or recording environments, thereby improving the stability of the simulated speech information.
[0107] In one optional embodiment, the voice recording requirement includes recording the direction of the sound source and the direction of noise interference;
[0108] The sound information output module 501 includes:
[0109] The audio signal output unit is used to control the audio output module corresponding to the direction of the recorded sound source in the simulated speech recording system to output the audio signal to be recorded; and...
[0110] The noise interference signal output unit is used to control the sound output module corresponding to the noise interference direction in the simulated voice recording system to output a noise interference signal.
[0111] In one alternative embodiment, the voice recording requirement includes recording the intensity of the sound source;
[0112] The audio signal output unit to be recorded includes:
[0113] The sound source intensity control subunit is used to control the sound output module corresponding to the direction of the recorded sound source in the simulated speech recording system according to the recorded sound source intensity, so as to output the sound information to be recorded.
[0114] In one alternative embodiment, the voice recording requirement includes noise interference intensity;
[0115] The audio signal output unit to be recorded includes:
[0116] The interference intensity control subunit is used to control the sound output module corresponding to the noise interference direction in the simulated voice recording system to output a noise interference signal according to the noise interference intensity.
[0117] The above-described voice information recording device can execute the voice information recording method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing each voice information recording method.
[0118] As an implementation of the above-mentioned speech recognition model training methods, this disclosure also provides an optional embodiment of an execution device for implementing the speech recognition model training methods. See further details. Figure 6 The speech recognition model training device 600 shown includes: a simulated speech information acquisition module 601 and a training module 602. Among them,
[0119] The simulated voice information acquisition module 601 is used to acquire simulated voice information; wherein, the simulated voice information is acquired based on any of the voice information recording devices provided in the embodiments of this disclosure;
[0120] The training module 602 is used to train the speech recognition model in the simulated speech system based on the simulated speech information.
[0121] This embodiment uses simulated speech information recorded by a simulated speech recording system as training samples to train a speech recognition model in the simulated speech system. This eliminates the need to acquire the simulated speech system during the training sample preparation stage, reducing the difficulty of obtaining training samples and thus lowering the time cost of the training sample preparation stage. This improves the training efficiency of the speech recognition model. Furthermore, since the acquired simulated speech information is automatically recorded by the simulated speech recording system, the stability of the simulated speech information is ensured, thereby improving the accuracy of the speech recognition model training results.
[0122] In an optional embodiment, the device further includes:
[0123] An online voice information acquisition module is used to acquire online voice information in the voice system to be simulated.
[0124] The secondary training module is used to perform secondary training on the trained speech recognition model based on the online speech information.
[0125] The above-described speech recognition model training device can execute the speech recognition model training method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing each speech recognition model training method.
[0126] The technical solutions disclosed herein, including the collection, storage, use, processing, transmission, provision, and disclosure of system types, simulation space size, voice recording requirements, etc., all comply with relevant laws and regulations and do not violate public order and good morals.
[0127] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0128] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0129] like Figure 7As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded from storage unit 708 into random access memory (RAM) 703. RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0130] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0131] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs at least one of the methods and processes described above, such as a method for constructing a voice recording system, a method for recording voice information, and a method for training a voice recognition model. For example, in some embodiments, at least one of the methods for constructing a voice recording system, a method for recording voice information, and a method for training a voice recognition model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the method for constructing a voice recording system, a method for recording voice information, or a method for training a voice recognition model described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured by any other suitable means (e.g., by means of firmware) to perform at least one of a speech recording system construction method, a speech information recording method, and a speech recognition model training method.
[0132] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0133] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0134] In one alternative embodiment, the aforementioned electronic device may be an in-vehicle terminal.
[0135] In another alternative embodiment, this disclosure also provides a vehicle in which the aforementioned vehicle-mounted terminal is provided.
[0136] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0138] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0139] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem that addresses the management difficulties and weak business scalability inherent in traditional physical hosting and VPS services. Servers can also be servers for distributed systems or servers integrated with blockchain technology.
[0140] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0141] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0142] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A voice information recording method, comprising: controlling a sound output module in a simulated voice recording system to output a sound signal according to a voice recording requirement, wherein the voice recording requirement comprises a recording content requirement, a recording sound source direction, a noise interference direction, a recording sound source intensity, and a noise interference intensity; determining a sound output module for outputting a to-be-recorded sound signal based on the recording sound source direction and a preset direction module correspondence relationship between different sound output modules and sound emission directions in the simulated voice recording system; determining a sound output module for outputting a noise interference signal based on the noise interference direction and the preset module correspondence relationship; controlling a sound collection module in the simulated voice recording system to collect the sound signal to obtain simulated voice information; wherein the simulated voice recording system is constructed based on a voice recording system construction method as follows: determining a simulated voice module of a simulated voice recording system and a simulated position distribution between different simulated voice modules according to a system type of a to-be-simulated voice system, wherein the simulated voice module comprises a sound collection module and a sound output module, and the to-be-simulated voice system is a vehicle and the system type is a vehicle model identifier or a vehicle series identifier; constructing the simulated voice recording system in a simulation space according to the simulated voice module and the simulated position distribution for simulated voice information recording; wherein the number of the simulated voice module and the simulated position distribution between different simulated voice modules are determined by querying a simulated position distribution map or a simulated position distribution table based on the system type of the to-be-simulated voice system; wherein the number of the simulated voice module corresponding to different system types and the simulated position distribution between different simulated voice modules are stored in the simulated position distribution map or the simulated position distribution table.
2. The method of claim 1, wherein, The determination of the simulated position distribution between different simulated voice modules according to the system type of the to-be-simulated voice system comprises: determining an actual position distribution between different actual voice modules in the to-be-simulated voice system according to the system type of the to-be-simulated voice system; determining the simulated position distribution between different simulated voice modules according to the actual position distribution; wherein the simulated voice module and the actual voice module correspond to each other.
3. The method of claim 2, wherein, The determination of the simulated position distribution between different simulated voice modules according to the actual position distribution comprises: determining a scaling ratio according to a space size of the simulation space; determining the simulated position distribution between different simulated voice modules according to the scaling ratio and the actual position distribution.
4. The method according to any one of claims 1 to 3, wherein, The sound output module comprises a loudspeaker and / or a simulated human head.
5. The method of claim 1, wherein, The voice recording requirement comprises a recording sound source direction and a noise interference direction. The controlling of the sound output module in the simulated voice recording system to output the sound signal according to the voice recording requirement comprises: controlling a sound output module corresponding to the recording sound source direction in the simulated voice recording system to output a to-be-recorded sound signal; and controlling a sound output module corresponding to the noise interference direction in the simulated voice recording system to output a noise interference signal. The voice recording requirement comprises a recording sound source intensity.
6. The method of claim 5, wherein, The control sound output module corresponding to the recording sound source direction in the simulation voice recording system outputs the to-be-recorded sound signal, and includes: According to the recording sound source intensity, the sound output module corresponding to the recording sound source direction in the simulation voice recording system outputs the to-be-recorded sound signal.
7. The method of claim 5, wherein, The voice recording requirement includes noise interference intensity; The control sound output module corresponding to the noise interference direction in the simulation voice recording system outputs the noise interference signal, and includes: According to the noise interference intensity, the sound output module corresponding to the noise interference direction in the simulation voice recording system outputs the noise interference signal.
8. A voice recognition model training method, comprising: Obtaining simulation voice information; wherein the simulation voice information is collected based on the voice information recording method of any one of claims 1-7; According to the simulation voice information, the voice recognition model in the to-be-simulated voice system is trained; Collecting online voice information in the to-be-simulated voice system; wherein the online voice information refers to the voice information recorded in the real environment of the to-be-simulated voice system; According to the online voice information, the trained voice recognition model is further trained.
9. A voice information recording device, comprising: A sound information output module for controlling the sound output module in the simulation voice recording system to output a sound signal according to voice recording requirements; wherein the voice recording requirements include recording content requirements, recording sound source direction, noise interference direction, recording sound source intensity, and noise interference intensity; the sound output module for outputting the to-be-recorded sound signal is determined based on the recording sound source direction and the preset direction module correspondence relationship between different sound output modules and sound emission directions in the simulation voice recording system; the sound output module for outputting the noise interference signal is determined based on the noise interference direction and the preset module correspondence relationship; A simulation voice information obtaining module for controlling the sound collection module in the simulation voice recording system to collect the sound signal and obtain simulation voice information; The simulation voice recording system is constructed based on the voice recording system construction device as follows: A simulation position distribution determination module for determining the simulation voice module of the simulation voice recording system and the simulation position distribution between different simulation voice modules according to the system type of the to-be-simulated voice system; wherein the simulation voice module includes a sound collection module and a sound output module; the to-be-simulated voice system is a vehicle, and the system type is a vehicle model identifier or a vehicle series identifier; A simulation system construction module for constructing the simulation voice recording system in a simulation space according to the simulation voice module and the simulation position distribution, for simulation voice information recording; The number of simulation voice modules and the simulation position distribution between different simulation voice modules are determined by querying a simulation position distribution diagram or a simulation position distribution table based on the system type of the to-be-simulated voice system. The simulation position distribution table or the simulation position distribution diagram stores the number of simulation voice modules corresponding to different system types and the simulation position distribution between different simulation voice modules.
10. The apparatus of claim 9, wherein, The simulation position distribution determining module comprises: An actual position distribution determining unit configured to determine an actual position distribution between different actual voice modules in the voice system to be simulated according to the system type of the voice system to be simulated; A simulation position distribution determining unit configured to determine a simulation position distribution between different simulation voice modules according to the actual position distribution. The simulation voice module corresponds to the actual voice module one by one.
11. The apparatus of claim 10, wherein, The simulation position distribution determining unit comprises: A scaling ratio determining sub-unit configured to determine a scaling ratio according to the space size of the simulation space; A simulation position distribution determining sub-unit configured to determine a simulation position distribution between different simulation voice modules according to the scaling ratio and the actual position distribution.
12. The apparatus of any of claims 9-11, wherein, The sound output module comprises a loudspeaker and / or a simulated human head.
13. The apparatus of claim 9, wherein, The voice recording requirement comprises a recording sound source direction and a noise interference direction; The sound information output module comprises: A to-be-recorded sound signal output unit configured to control the sound output module corresponding to the recording sound source direction in the simulation voice recording system to output a to-be-recorded sound signal; and A noise interference signal output unit configured to control the sound output module corresponding to the noise interference direction in the simulation voice recording system to output a noise interference signal. The voice recording requirement comprises a recording sound source intensity.
14. The apparatus of claim 13, wherein, The to-be-recorded sound signal output unit comprises: A sound source intensity control sub-unit configured to control the sound output module corresponding to the recording sound source direction in the simulation voice recording system to output the to-be-recorded sound information according to the recording sound source intensity. The voice recording requirement comprises a noise interference intensity.
15. The apparatus of claim 13, wherein, The to-be-recorded sound signal output unit comprises: An interference intensity control sub-unit configured to control the sound output module corresponding to the noise interference direction in the simulation voice recording system to output a noise interference signal according to the noise interference intensity.
16. A voice recognition model training apparatus, comprising: A simulation voice information acquisition module configured to acquire simulation voice information; wherein the simulation voice information is collected based on the voice information recording apparatus of any one of claims 9-15; A training module configured to train a voice recognition model in a voice system to be simulated according to the simulation voice information; An online voice information collection module configured to collect online voice information in the voice system to be simulated; wherein the online voice information refers to voice information recorded in a real environment of the voice system to be simulated; A secondary training module configured to perform secondary training on the trained voice recognition model according to the online voice information.
17. An electronic device, comprising: At least one processor; And A memory in communication connection with the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of the voice information recording method in claims 1-7 and the voice recognition model training method in claim 8.
18. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform any one of the voice information recording method in claims 1-7 and the voice recognition model training method in claim 8.
19. A computer program product comprising computer programs / instructions which, when executed by a processor, implement any one of the voice information recording method in claims 1-7 and the voice recognition model training method in claim 8.
Citation Information
Patent Citations
Mining area unmanned transportation simulation test platform and mining area unmanned transportation simulation method
CN112365216A
Method and apparatus for tuning speech recognition systems to accommodate ambient noise
US9978399B2