Speech stream processing methods, deep learning model training methods, devices and intelligent agents
By extracting features from speech frame sequences and fusing them with an attention mechanism, combined with preset speech attributes, the problem of low-latency, high-quality speech conversion in the context of infinitely streaming speech data is solved, achieving speech coherence and semantic integrity.
Patent Information
- Application Number
- CN202510334495.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Existing speech conversion methods struggle to meet low-latency requirements in scenarios with unlimited streaming speech data, and are prone to losing semantics and exhibiting incoherent and unnatural prosody.
By extracting features from the speech frame sequence in the speech stream to be processed, using an attention mechanism to fuse the first and second speech features, and combining them with preset speech attributes for conversion, high-quality speech conversion with low latency is achieved.
It achieves low-latency, high-quality speech conversion in scenarios with unlimited streaming voice data, while maintaining the coherence and semantic integrity of the speech.
Smart Images

Figure CN120183435B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and more particularly to the fields of deep learning, speech processing, and speech conversion technology. More specifically, it designs a speech stream processing method, a deep learning model training method, apparatus, device, medium, program product, and intelligent agent. Background Technology
[0002] With the rapid development of artificial intelligence technology, speech conversion can be achieved based on AI. For example, speech conversion can convert the speech features of a source speaker into the speech features of a target speaker, while preserving the semantic content of the source speaker's speech. Summary of the Invention
[0003] This disclosure provides a speech stream processing method, a deep learning model training method, an apparatus, a device, a medium, a program product, and an intelligent agent.
[0004] According to one aspect of this disclosure, a speech stream processing method is provided, comprising: extracting features from a first speech frame sequence in a speech stream to obtain first speech features, wherein the first speech frame sequence overlaps with at least one second speech frame in a second speech frame sequence, and the second speech frame sequence is arranged before the first speech frame sequence in the speech stream; fusing the first speech features and the second speech features determined based on the second speech frame sequence based on an attention mechanism to obtain speech fusion features; and converting the speech fusion features based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence.
[0005] According to another aspect of this disclosure, a method for training a deep learning model is provided, comprising: acquiring a sample speech stream, wherein a first sample speech frame sequence in the sample speech stream overlaps with at least one second speech frame in a second sample speech frame sequence, and the second sample speech frame sequence is arranged before the first sample speech frame sequence in the sample speech stream; extracting features from the first sample speech frame sequence using a feature extraction layer to obtain first sample speech features; masking the first sample speech features to obtain masked sample speech features; fusing the masked sample speech features and the second sample speech features determined based on the second speech frame sequence using a feature fusion layer to obtain fused sample speech features; and training a speech feature extraction network using the fused sample speech features based on a self-supervised mechanism to obtain a trained deep learning model.
[0006] According to another aspect of this disclosure, a speech stream processing apparatus is provided, comprising: a first extraction module, configured to extract features from a first speech frame sequence in a speech stream to be processed, to obtain first speech features, wherein the first speech frame sequence overlaps with at least one second speech frame in a second speech frame sequence, and the second speech frame sequence is arranged before the first speech frame sequence in the speech stream; a fusion module, configured to fuse the first speech features and the second speech features determined based on the second speech frame sequence based on an attention mechanism, to obtain speech fusion features; and a conversion module, configured to convert the speech fusion features based on preset speech attributes, to obtain converted speech data corresponding to the first speech frame sequence.
[0007] According to another aspect of this disclosure, a training apparatus for a deep learning model is provided, comprising: an acquisition module for acquiring a sample speech stream, wherein a first sample speech frame sequence in the sample speech stream overlaps with at least one second speech frame in a second sample speech frame sequence, and the second sample speech frame sequence is arranged before the first sample speech frame sequence in the sample speech stream; a second extraction module for extracting features from the first sample speech frame sequence using a feature extraction layer to obtain first sample speech features; a masking module for masking the first sample speech features to obtain masked sample speech features; a second fusion module for fusing the masked sample speech features and the second sample speech features determined based on the second speech frame sequence using a feature fusion layer to obtain fused sample speech features; and a training module for training a speech feature extraction network using the fused sample speech features based on a self-supervised mechanism to obtain a trained deep learning model.
[0008] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by invoking the large model to execute a method provided according to an embodiment of this disclosure; and an output module for outputting the output information obtained by the processing module.
[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods described above.
[0011] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described above.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 The illustration schematically shows an exemplary system architecture to which speech stream processing methods, deep learning model training methods and apparatus can be applied according to embodiments of the present disclosure;
[0015] Figure 2 A flowchart illustrating a speech stream processing method according to an embodiment of the present disclosure is shown schematically.
[0016] Figure 3 The illustration shows a schematic diagram of obtaining speech fusion features based on attention mechanism fusion window speech features and first speech features according to an embodiment of the present disclosure;
[0017] Figure 4 The illustration shows a schematic diagram of attention masking of input features based on a preset window sliding according to an embodiment of the present disclosure;
[0018] Figure 5 This schematically illustrates a flowchart of processing a speech stream to obtain converted speech data according to an embodiment of the present disclosure;
[0019] Figure 6 A flowchart illustrating a method for training a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0020] Figure 7 A block diagram of a voice stream processing apparatus according to an embodiment of the present disclosure is shown schematically;
[0021] Figure 8 A block diagram illustrating a training apparatus for a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0022] Figure 9 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and
[0023] Figure 10 A schematic block diagram of an example electronic device is shown that can be used to implement the speech stream processing method or deep learning model training method of the embodiments of this disclosure. Detailed Implementation
[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.
[0026] With the rapid development of artificial intelligence technology, speech conversion can be achieved based on AI. For example, speech conversion can convert the speech features of a source speaker into the speech features of a target speaker, while preserving the semantic content of the source speaker's speech.
[0027] In realizing the inventive concept of this disclosure, the inventors discovered that related speech conversion methods typically require complete source speech data as input in order to convert the complete source speech data and output complete target speech data. However, in the scenario of unlimited streaming speech data, since unlimited streaming speech data is continuously and in real time, related speech conversion methods are unable to meet the requirements of low latency and suffer from problems such as easy loss of semantics and incoherent and unnatural prosody.
[0028] Figure 1 The illustration schematically depicts an exemplary system architecture to which content processing methods and apparatus can be applied according to embodiments of the present disclosure.
[0029] It is important to note that Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the content processing methods and apparatus can be applied may include a terminal device, but the terminal device may implement the content processing methods and apparatus provided by the embodiments of this disclosure without interacting with the server.
[0030] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0031] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0032] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0033] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0034] A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. It solves the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. A server can also be a server for a distributed system or a server that incorporates blockchain technology.
[0035] It should be noted that the content processing method provided in the embodiments of this disclosure can generally be executed by terminal devices 101, 102, or 103. Accordingly, the content processing apparatus provided in the embodiments of this disclosure can also be disposed in terminal devices 101, 102, or 103.
[0036] Alternatively, the content processing method provided in this embodiment can generally be executed by server 105. Correspondingly, the content processing apparatus provided in this embodiment can generally be located in server 105. The content processing method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the content processing apparatus provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0037] For example, when a user is reading an ebook online, terminal devices 101, 102, and 103 can acquire the target content in the ebook that the user is looking at, and then send the acquired target content to server 105. Server 105 analyzes the target content to determine its feature information; predicts content that the user is interested in based on the feature information; and extracts the content that the user is interested in. Alternatively, a server or server cluster capable of communicating with terminal devices 101, 102, and 103 and / or server 105 can analyze the target content and ultimately extract the content that the user is interested in.
[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0039] Figure 2 A flowchart illustrating a speech stream processing method according to an embodiment of the present disclosure is shown schematically.
[0040] like Figure 2 As shown, the method 200 includes operations S210 to S230.
[0041] In operation S210, feature extraction is performed on the first speech frame sequence in the speech stream to be processed to obtain the first speech feature, wherein the first speech frame sequence overlaps with at least one second speech frame in the second speech frame sequence, and the second speech frame sequence is arranged before the first speech frame sequence in the speech stream.
[0042] In operation S220, the first speech feature and the second speech feature determined based on the second speech frame sequence are fused based on the attention mechanism to obtain the speech fusion feature.
[0043] In operation S230, the speech fusion features are converted based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence.
[0044] According to embodiments of this disclosure, the voice stream to be processed can be understood as a voice stream to be converted into speech. The voice stream to be processed is transmitted and processed in real time in a continuous streaming form, and may include multiple voice frames. For example, the voice stream to be processed may include real-time communication voice streams, online live broadcast voice streams, etc., and the input duration of the voice stream to be processed may be several minutes, tens of minutes, several hours, or even longer.
[0045] According to one embodiment of this disclosure, in response to receiving a voice stream signal to be processed, the continuous voice stream can be segmented into multiple voice frames. Further, the multiple voice frames can be sequentially processed into multiple voice data packets, each voice data packet being understood as a frame sequence comprising multiple voice frames. The voice stream to be processed can be processed in batches in real time based on multiple voice data packets, thereby allowing analysis and processing to be performed immediately upon receiving a certain amount of voice stream signal, without waiting for the complete voice stream data, thus achieving low-latency real-time conversion.
[0046] According to embodiments of this disclosure, the first speech frame sequence can be understood as the currently processed speech data packet in the speech stream to be processed. The second speech frame sequence can be understood as the previous speech data packet adjacent to the current speech data packet in the speech stream to be processed. The second speech frame sequence is arranged before the first speech frame sequence in the speech stream, and at least one speech frame in the first speech frame sequence overlaps with at least one speech frame in the second speech frame sequence. In other words, the first speech frame sequence includes at least one speech frame from the preceding adjacent speech frame sequence (the second speech frame sequence) in the speech stream.
[0047] In one embodiment, the second speech frame sequence may include, for example, speech frames 1 to 20 (e.g., denoted as P1-P20), and the first speech frame sequence may include, for example, speech frames 11 to 30 (e.g., denoted as P11-P30). It should be noted that those skilled in the art can reasonably set the number of overlapping speech frames between the first and second speech frame sequences according to actual needs or application scenarios, and no specific limitation is made here.
[0048] According to embodiments of this disclosure, features can be extracted from a second speech frame sequence in the speech stream to be processed to obtain second speech features. Features can also be extracted from a first speech frame sequence in the speech stream to be processed to obtain first speech features. The first and second speech features can be fused based on an attention mechanism to obtain fused speech features.
[0049] It is understandable that, since the first speech frame sequence contains at least one speech frame in the second speech frame sequence, feature extraction of the first speech frame sequence can make the first speech features contain the contextual speech information in the second speech frame sequence. The first speech features and the second speech features can be fused based on the attention mechanism to obtain speech fusion features, so that the speech fusion features can maintain speech coherence and semantic integrity based on the contextual speech information of the second speech frame sequence.
[0050] According to embodiments of this disclosure, in speech conversion applications, the speech stream to be processed can be understood as the source speech stream data of the source speaker. Preset speech attributes can characterize the timbre, pitch, and speech style of the target speaker; for example, preset speech attributes may include the target speaker's gender, age, timbre, emotional style, etc. The source speech stream data can be converted based on the preset speech attributes to obtain converted speech stream data, thereby preserving the semantics and prosody of the source speech stream data while converting the speaker from the source speaker to the target speaker.
[0051] According to embodiments of this disclosure, speech fusion features can characterize the speech information of a first speech frame sequence and its preceding speech information. The speech fusion features can be converted based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence. Thus, the obtained converted speech data can express the speech content of the first speech frame sequence based on preset speech attributes under the condition that the speech rhythm is natural and the semantics are complete and accurate, thereby achieving high-quality speech conversion of the speech stream.
[0052] According to embodiments of this disclosure, by processing the speech stream to be processed in batches based on multiple speech frame sequences in real time, analysis and processing can be performed immediately upon receiving a speech stream signal of a preset data length, without waiting for the complete speech stream data, thereby achieving low-latency real-time interaction. The first speech frame sequence includes at least one speech frame from the preceding adjacent speech frame sequence (the second speech frame sequence) in the speech stream. Therefore, by extracting features from the first speech frame sequence, the first speech features can contain the preceding speech information from the second speech frame sequence. The first and second speech features can be fused based on an attention mechanism to obtain speech fusion features, which can maintain speech coherence and semantic integrity based on the preceding speech information of the second speech frame sequence. The speech fusion features can be converted based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence. This allows the obtained converted speech data to express the speech content of the first speech frame sequence based on preset speech attributes, while maintaining natural speech rhythm and accurate semantic completeness. Thus, high-quality, low-latency speech conversion operations can be achieved by real-time processing and conversion of the current data packets in the speech stream.
[0053] It should be noted that the voice stream processing method provided in this disclosure can be applied to real-time voice interaction application scenarios, such as intelligent customer service, voice communication, and online live streaming. The voice stream processing method provided in this disclosure does not limit the application scenarios.
[0054] According to embodiments of this disclosure, the speech fusion feature obtained by fusing a first speech feature and a second speech feature determined based on a second speech frame sequence based on an attention mechanism includes: masking the second speech feature based on a window mechanism to obtain a window speech feature; and fusing the window speech feature and the first speech feature based on an attention mechanism to obtain the speech fusion feature.
[0055] According to embodiments of this disclosure, a second speech feature can be masked based on a window mechanism to obtain windowed speech features. For example, a portion of the second speech feature can be masked based on a preset window to obtain windowed speech features. Alternatively, a portion of the first sub-features in the first speech feature can be masked based on a preset window to determine the sub-features from the first speech feature that require contextual semantic attention calculation among multiple first speech sub-features, thereby improving the information completeness and accuracy of the speech fusion features. Those skilled in the art can reasonably set the preset window according to actual needs or application scenarios, and no specific limitations are made here.
[0056] According to embodiments of this disclosure, speech fusion features can be obtained by fusing window speech features and first speech features based on an attention mechanism.
[0057] Understandably, the second speech feature can represent the preceding speech information adjacent to the first speech feature. Window-based masking operations can effectively capture local, nearby preceding speech information in the second speech feature while improving computational efficiency. Based on this, by fusing the windowed speech feature and the first speech feature using an attention mechanism, a speech fusion feature is obtained. This ensures that the speech fusion feature can maintain speech coherence and semantic integrity based on the local, nearby preceding speech information in the second speech feature.
[0058] According to embodiments of this disclosure, the second speech feature includes a plurality of second sub-features arranged in sequence, and the first speech feature includes a first target sub-feature; wherein, masking the second speech feature based on a window mechanism to obtain a window speech feature includes: determining at least one window sub-feature adjacent to the first target sub-feature from the plurality of second sub-features based on a preset window; and masking the other second sub-features in the second speech feature other than the at least one window sub-feature to obtain a window speech feature corresponding to the first target sub-feature.
[0059] According to one embodiment of this disclosure, the first target sub-feature in the first speech feature can be determined based on the length of a preset window. For example, the first speech feature may include five first sub-features arranged in sequence. The length of the preset window may be, for example, 2, in which case the first two of the five first sub-features can be determined as the first target sub-features.
[0060] According to one embodiment of this disclosure, at least one window sub-feature in the second speech feature can be determined based on the length of a preset window. For example, the second speech feature may include five second sub-features arranged in sequence, and the length of the preset window may be, for example, 2. In this case, the last two of the five second sub-features adjacent to the first target sub-feature can be determined as window sub-features.
[0061] As an example, the second speech features, except for at least one window sub-feature, can be masked to obtain the window speech features corresponding to the first target sub-feature.
[0062] Figure 3 The illustration shows a schematic diagram of obtaining speech fusion features based on attention mechanism fusion window speech features and first speech features according to an embodiment of the present disclosure.
[0063] like Figure 3 As shown, the second speech frame sequence 302 may include, for example, the first to the 20th speech frames in the speech stream to be processed, such as P1-P20. The first speech frame sequence 301 may include, for example, the 21st to the 30th speech frames in the speech stream to be processed, such as P21-P30. The second speech frame sequence 302 is arranged before the first speech frame sequence 301 in the speech stream, and the first speech frame sequence 301 overlaps with the 11th to the 20th speech frames in the second speech frame sequence 302, for example, the overlapping speech frames may be P11-P20.
[0064] like Figure 3 As shown, feature extraction can be performed on the second speech frame sequence 302 to obtain second speech feature 304. The second speech feature 304 may include multiple second sub-features arranged in sequence, such as A1, A2, A3, A4, and A5. Feature extraction can be performed on the first speech frame sequence 301 to obtain first speech feature 303. The first speech feature 303 may include multiple first sub-features B1, B2, B3, B4, and B5 arranged in sequence.
[0065] like Figure 3 As shown, the length of the preset window 305 can be 2. Based on the preset window 305, B1 and B2 among multiple first sub-features can be determined as the first target sub-features. Based on the preset window 305, A4 and A5, which are adjacent to B1 and B2 among multiple second sub-features, can be determined as window sub-features. On this basis, the other second sub-features (i.e., A1-A3) in the second speech feature 304, except for A4 and A5, can be masked to obtain the window speech feature 306 corresponding to the first target sub-features.
[0066] like Figure 3As shown, the speech fusion sub-feature corresponding to the first target sub-feature B1 can be obtained by fusing window speech feature 306 and first speech feature 303 based on the attention mechanism. Based on the fusion results of the speech fusion sub-features corresponding to multiple first sub-features B1 to B5, the speech fusion feature 307 can be determined.
[0067] According to embodiments of this disclosure, masking other second sub-features in the second speech features, excluding at least one window sub-feature, to obtain window speech features corresponding to the target sub-feature includes: masking other second sub-features in the second speech features, excluding at least one window sub-feature, to obtain at least one first window sub-feature, wherein the second speech features include the first window sub-feature; determining the second window sub-feature from other first sub-features in the first speech features that are arranged before the first target sub-feature; and determining the window speech features based on the first window sub-feature and the second window sub-feature.
[0068] According to embodiments of this disclosure, in the attention mechanism, the query vector, key vector, and value vector can be obtained from the input features through a linear transformation. For example, windowed speech features and first speech features can be used as input features, and the query vector, key vector, and value vector can be obtained based on second speech features and parameter matrices WQ, Wk, and Wv, respectively.
[0069] According to one embodiment of this disclosure, attention masking can be applied to the input second speech features based on a preset window sliding mechanism. This limits the attention calculation range of the first target sub-feature to a local area within a preset window adjacent to the first target sub-feature, allowing each first sub-feature to focus only on its local contextual speech information. Since prosody, intonation, and timbre in speech stream signals typically change within a short timeframe, window-sliding attention masking can effectively capture these local contextual features, thereby better preserving speech coherence and semantic integrity.
[0070] According to one embodiment of this disclosure, other second sub-features in the second speech features, besides at least one window sub-feature, can be masked to obtain at least one first window sub-feature, wherein the second speech features include the first window sub-features. The second window sub-features can be determined from other first sub-features in the first speech features that precede the first target sub-feature. Window speech features can be determined based on the first window sub-features and the second window sub-features.
[0071] Figure 4 The illustration shows a schematic diagram of attention masking of input features based on a preset window sliding according to an embodiment of the present disclosure.
[0072] like Figure 4As shown, the window length of the preset window can be set to, for example, 2, and the sliding step size can be set to, for example, 1. The first speech feature 401 includes multiple first sub-features B1 to B5 arranged in sequence, and the second speech feature 402 includes multiple second sub-features A1 to A5 arranged in sequence.
[0073] like Figure 4 As shown, for example, the second speech feature 402 is masked based on a window mechanism. The first sub-feature B1 is used as the first target sub-feature. The second sub-features A1, A2, and A3 in the second speech feature 402 are masked based on a preset window, resulting in the window sub-features A4 and A5 corresponding to the first sub-feature B1. The second sub-features A4 and A5 are used as key and value vectors, respectively, and the first sub-feature B1 is used as the query vector to perform attention calculation, resulting in the speech fusion sub-feature corresponding to the first sub-feature B1. The first sub-feature B2 is used as the first target sub-feature. The second sub-features A1, A2, A3, and A4 in the second speech feature 402 are masked based on a preset window, resulting in the window sub-features A5 and the first feature B1 corresponding to the first sub-feature B2. The second sub-features A5 and B1 in the window sub-features are used as key and value vectors, respectively, and the first feature B2 is used as the query vector to perform attention calculation, resulting in the speech fusion sub-feature corresponding to the first feature B2. It should be understood that the window sub-feature corresponding to the first sub-feature B3 can be the first sub-features B1 and B2. Using the first sub-features B1 and B2 as key and value vectors, and the first sub-feature B3 as the query vector, attention calculation is performed to obtain the speech fusion sub-feature corresponding to the first sub-feature B3. Similarly, the window sub-feature corresponding to the first sub-feature B4 can be the first sub-features B2 and B3. Using the first sub-features B2 and B3 as key and value vectors, and the first sub-feature B4 as the query vector, attention calculation is performed to obtain the speech fusion sub-feature corresponding to the first sub-feature B4. Likewise, the window sub-feature corresponding to the first sub-feature B5 can be the first sub-features B3 and B4. Using the first sub-features B3 and B4 as key and value vectors, and the first sub-feature B5 as the query vector, attention calculation is performed to obtain the speech fusion sub-feature corresponding to the first sub-feature B5. Thus, multiple speech fusion features corresponding to each of the first sub-features B1 to B5 can be obtained. By fusing multiple speech fusion sub-features, the speech fusion feature can be determined. For example, a normalization algorithm can be used to fuse multiple speech fusion sub-features to obtain speech fusion features.
[0074] According to embodiments of this disclosure, feature extraction of a first speech frame sequence in a speech stream to obtain a first speech feature includes: performing at least one convolution operation on the first speech frame sequence based on a first convolution kernel to obtain an initial speech feature; and performing at least one convolution operation on the initial speech feature based on a second convolution kernel to obtain the first speech feature, wherein the step size of the first convolution kernel is greater than the step size of the second convolution kernel.
[0075] According to one embodiment of this disclosure, each speech frame in the first speech frame sequence can be processed to obtain the mel-spectrum corresponding to each speech frame. The mel-spectrum can convert the spectral representation of the audio signal to a mel frequency scale to simulate the human ear's perception of frequency.
[0076] In one embodiment, the Mel spectrum corresponding to each speech frame in the first speech frame sequence can be input into a convolutional network for feature extraction to obtain the first speech features.
[0077] As an example, the convolutional network described above may include: a first convolutional layer based on a first convolutional kernel and a second convolutional layer based on a second convolutional kernel. The stride of the first convolutional kernel is greater than the stride of the second convolutional kernel.
[0078] In one embodiment, the kernel size of the first convolutional kernel can be set to, for example, 5, and the stride can be set to, for example, 2. The kernel size of the second convolutional kernel can be set to, for example, 5, and the stride can be set to, for example, 1. Exemplarily, at least one convolution operation can be performed on the first speech frame sequence based on the first convolutional kernel to obtain initial speech features. At least one convolution operation can be performed on the initial speech features based on the second convolutional kernel to obtain the first speech features.
[0079] According to one embodiment of this disclosure, the above convolution operation can be selected as causal convolution. Causal convolution can be used to process time series data. It only uses information from the current time and previous times for convolution, thereby avoiding interference from future information. This convolution method can introduce positional information without the need for positional encoding.
[0080] According to embodiments of this disclosure, convolutional networks can incorporate positional information based on at least one causal convolutional layer. By controlling the number of convolutional layers and the size of the convolutional kernels within each layer, the receptive field of the convolutional network can be effectively expanded, enabling it to capture a wider range of contextual information. By controlling the stride of the convolutional kernels, computational efficiency can be optimized, reducing latency in real-time processing of speech stream data.
[0081] According to embodiments of this disclosure, the process of converting speech fusion features based on preset speech attributes to obtain converted speech data corresponding to a first speech frame sequence includes: performing an upsampling convolution operation on the speech fusion features to obtain target fusion features; performing feature fusion on the preset speech attributes and the speech fusion features to obtain converted speech features; and determining the converted speech data corresponding to the first speech frame sequence based on the converted speech features.
[0082] According to one embodiment of this disclosure, speech fusion features can be input into an encoder for upsampling convolution to obtain target fusion features, such that the data size of the target fusion features is consistent with the data size of the first speech frame sequence. For example, the data size of the first speech frame sequence can be 20 frames. Feature extraction can be performed on the first speech frame sequence to obtain first speech features with a data size of 5 frames. The first speech features and second speech features determined based on the second speech frame sequence can be fused based on an attention mechanism to obtain speech fusion features with a data size of 5 frames. Upsampling convolution can be performed on the speech fusion features to obtain a target fusion operation with a data size of 20 frames.
[0083] According to embodiments of this disclosure, in speech conversion applications, preset speech attributes can characterize the timbre, pitch, and speech style of the target speaker. Feature fusion can be performed on the preset speech attributes and speech fusion features to obtain converted speech features. The feature fusion method can be, for example, based on an attention mechanism, and is not specifically limited thereto.
[0084] According to one embodiment of this disclosure, the converted speech features can be input into a decoder to obtain converted speech data corresponding to a first speech frame sequence. Thus, the speaker can be converted from the source speaker to the target speaker while preserving the semantics and prosody of the first speech frame sequence.
[0085] Figure 5 A flowchart illustrating the processing of a speech stream to obtain converted speech data according to an embodiment of the present disclosure is shown.
[0086] like Figure 5 As shown, features can be extracted from the second speech frame sequence 502 in the speech stream to be processed to obtain second speech features 504. Features can be extracted from the first speech frame sequence 501 in the speech stream to be processed to obtain first speech features 503. The second speech frame sequence 502 is arranged before the first speech frame sequence 501 in the speech stream, and at least one second speech frame in the first speech frame sequence 501 overlaps with the second speech frame sequence 502.
[0087] like Figure 5As shown, the second speech feature 504 can be masked based on a window mechanism to obtain the windowed speech feature 505. The windowed speech feature 505 and the first speech feature 503 can be fused based on an attention mechanism to obtain the speech fusion feature 506.
[0088] like Figure 5 As shown, an upsampling convolution operation can be performed on the speech fusion feature 506 to obtain the target fusion feature 507. Feature fusion can be performed on the preset speech attribute 508 and the speech fusion feature 507 to obtain the converted speech feature 509. Based on the converted speech feature 509, the converted speech data 510 corresponding to the column of the first speech frame sequence 501 can be determined.
[0089] Figure 6 A flowchart illustrating a training method for a deep learning model according to an embodiment of the present disclosure is shown. The speech feature extraction network of the deep learning model includes a feature extraction layer and a feature fusion layer.
[0090] like Figure 6 As shown, the method 600 includes operations S610~S3650.
[0091] In operation S610, a sample speech stream is acquired. The sample first speech frame sequence in the sample speech stream overlaps with at least one second speech frame in the sample second speech frame sequence. The sample second speech frame sequence is arranged before the sample first speech frame sequence in the sample speech stream.
[0092] In operation S620, the feature extraction layer is used to extract features from the first speech frame sequence of the sample to obtain the first speech features of the sample.
[0093] In operation S630, the first speech feature of the sample is masked to obtain the masked speech feature of the sample.
[0094] In operation S640, attention features are fused between the sample mask speech features and the sample second speech features determined based on the second speech frame sequence using the feature fusion layer to obtain the sample speech fusion features.
[0095] When operating the S650, a speech feature extraction network is trained using sample speech fusion features based on a self-supervised mechanism to obtain a trained deep learning model.
[0096] According to embodiments of this disclosure, a first sample speech frame sequence can be determined from an obtained sample speech stream. The first sample speech frame sequence overlaps with at least one second speech frame in a second sample speech frame sequence, and the second sample speech frame sequence precedes the first sample speech frame sequence in the sample speech stream.
[0097] According to embodiments of this disclosure, a feature extraction layer can be used to extract features from a first speech frame sequence to obtain sample first speech features. A feature extraction layer can also be used to extract features from a second speech frame sequence to obtain sample second speech features. The sample first speech features can be randomly masked to obtain sample masked speech features. Finally, a feature fusion layer can be used to fuse the sample masked speech features and the sample second speech features using attention features to obtain sample fused speech features.
[0098] According to embodiments of this disclosure, sample speech fusion features can be used as input to a deep learning model to train the deep learning model to predict masked features. The model parameters are then optimized using a loss function to obtain the trained deep learning network. By combining random masking and prediction tasks, the deep learning model can learn the speech features of the first sample speech frame sequence and its adjacent contextual speech features based on a self-supervised mechanism. This allows it to accurately represent the prosodic and semantic features of the first sample speech frame sequence.
[0099] Figure 7 A block diagram of a voice stream processing apparatus according to an embodiment of the present disclosure is shown schematically.
[0100] like Figure 7 As shown, the voice stream processing device 700 may include a first extraction module 710, a fusion module 720, and a conversion module 730.
[0101] The first extraction module 710 is used to extract features from the first speech frame sequence in the speech stream to be processed, and obtain the first speech features, wherein the first speech frame sequence overlaps with at least one second speech frame in the second speech frame sequence, and the second speech frame sequence is arranged before the first speech frame sequence in the speech stream.
[0102] The fusion module 720 is used to fuse the first speech features and the second speech features determined based on the second speech frame sequence based on the attention mechanism to obtain speech fusion features.
[0103] The conversion module 730 is used to convert speech fusion features based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence.
[0104] According to embodiments of this disclosure, the fusion module may include a mask submodule and a first fusion submodule.
[0105] The masking submodule is used to mask the second speech features based on the window mechanism to obtain windowed speech features.
[0106] The first fusion submodule is used to fuse window speech features and first speech features based on an attention mechanism to obtain speech fusion features.
[0107] According to embodiments of this disclosure, the second speech feature includes a plurality of second sub-features arranged in sequence, the first speech feature includes a first target sub-feature, and the masking submodule may include a determining unit and a masking unit.
[0108] The determining unit is used to determine at least one window sub-feature that is adjacent to the first target sub-feature from a plurality of second sub-features based on a preset window.
[0109] The masking unit is used to mask the second sub-features in the second speech feature except for at least one window sub-feature, to obtain the window speech feature corresponding to the first target sub-feature.
[0110] According to embodiments of this disclosure, a masking unit may include a masking subunit, a first determining subunit, and a second determining subunit.
[0111] The masking sub-unit is used to mask the other second sub-features in the second speech feature except for at least one window sub-feature, to obtain at least one first window sub-feature, wherein the second speech feature includes the first window sub-feature.
[0112] The first determining sub-unit is used to determine the second window sub-feature from other first sub-features arranged before the first target sub-feature in the first speech features.
[0113] The second determining subunit is used to determine the window speech features based on the first window sub-features and the second window sub-features.
[0114] According to embodiments of this disclosure, the first extraction module may include a first convolution submodule and a second convolution submodule.
[0115] The first convolutional submodule is used to perform at least one convolution operation on the first speech frame sequence based on the first convolutional kernel to obtain initial speech features.
[0116] The second convolutional submodule is used to perform at least one convolution operation on the initial speech features based on the second convolutional kernel to obtain the first speech features, wherein the step size of the first convolutional kernel is greater than the step size of the second convolutional kernel.
[0117] According to embodiments of this disclosure, the conversion module may include an upsampling submodule, a second fusion submodule, and a conversion submodule.
[0118] The upsampling submodule is used to perform upsampling convolution operations on the speech fusion features to obtain the target fusion features.
[0119] The second fusion submodule is used to perform feature fusion on preset speech attributes and speech fusion features to obtain converted speech features.
[0120] The conversion submodule is used to determine the converted speech data corresponding to the first speech frame sequence based on the converted speech features.
[0121] Figure 8 A block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure is shown schematically. The speech feature extraction network of the deep learning model includes a feature extraction layer and a feature fusion layer.
[0122] like Figure 8 As shown, the training device 800 for the deep learning model may include an acquisition module 810, a second extraction module 820, a mask module 830, a second fusion module 840, and a training module 850.
[0123] The acquisition module 810 is used to acquire a sample speech stream, wherein the sample first speech frame sequence in the sample speech stream overlaps with at least one second speech frame in the sample second speech frame sequence, and the sample second speech frame sequence is arranged before the sample first speech frame sequence in the sample speech stream.
[0124] The second extraction module 820 is used to extract features from the first speech frame sequence of the sample using the feature extraction layer to obtain the first speech features of the sample.
[0125] The masking module 830 is used to mask the first speech feature of the sample to obtain the masked speech feature of the sample.
[0126] The second fusion module 840 is used to perform attention feature fusion on the sample mask speech features and the sample second speech features determined based on the second speech frame sequence using the feature fusion layer to obtain sample speech fusion features.
[0127] Training module 850 is used to train a speech feature extraction network based on a self-supervised mechanism and using sample speech fusion features to obtain a trained deep learning model.
[0128] It should be noted that the speech stream processing device part in the embodiments of this disclosure corresponds to the speech stream processing method part in the embodiments of this disclosure. For a detailed description of the speech stream processing device part, please refer to the speech stream processing method part, which will not be repeated here.
[0129] The training device part of the deep learning model in the embodiments of this disclosure corresponds to the training method part of the deep learning model in the embodiments of this disclosure. The specific description of the training device part of the deep learning model is referred to the training method part of the deep learning model, and will not be repeated here.
[0130] Figure 9 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0131] In embodiments of this disclosure, such as Figure 9 As shown, the AI agent 900 may include an input module 910, a processing module 920, and an output module 930.
[0132] Input module 910 is used to receive the voice stream to be processed;
[0133] The processing module 920 is used to obtain converted speech data by calling a trained deep learning model to execute the speech stream processing method provided according to the embodiments of this disclosure based on the speech stream to be processed received by the input module.
[0134] Output module 930 is used to output the converted speech data obtained by the processing module.
[0135] According to embodiments of this disclosure, the input module 910 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI agent 900 can understand and process. The input module 910 is the primary link for the AI agent 900 to interact with the outside world, enabling the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0136] In the example, input module 910 can input the speech stream to be processed as described above.
[0137] In the example, processing module 920 is the core support for the AI agent 900's ability to handle complex tasks. Processing module 920 can execute the speech stream processing methods described above.
[0138] In the example, the performance of the processing module 920 is closely related to the large model on which the AI agent 900 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 920 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0139] In the example, after the AI agent 900 acquires the speech stream to be processed, the processing module 920 can use a trained deep learning model to process the first speech frame sequence in the speech stream to obtain converted speech data, and then pass the converted speech data to the output module 930.
[0140] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. However, once the AI agent 900 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to generate weather forecasts.
[0141] In the example, output module 930 can output the converted speech data described above.
[0142] The AI agent 900 according to embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.
[0143] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0144] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0145] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.
[0146] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0147] Figure 10 A schematic block diagram of an example electronic device is shown that can be used to implement the speech stream processing method or deep learning model training method of the embodiments of this disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the disclosure described and / or claimed herein.
[0148] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1002 or a computer program loaded into random access memory (RAM) 1003 from storage unit 1008. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.
[0149] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0150] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as speech stream processing methods and deep learning model training methods. For example, in some embodiments, the speech stream processing methods and deep learning model training methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the speech stream processing methods and deep learning model training methods described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured in any other suitable manner (e.g., by means of firmware) to perform speech stream processing methods or deep learning model training methods.
[0151] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0152] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0153] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0155] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0156] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0157] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0158] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A speech stream processing method, comprising: Feature extraction is performed on a first speech frame sequence in the speech stream to be processed to obtain a first speech feature, wherein the first speech frame sequence overlaps with at least one second speech frame in a second speech frame sequence, and the second speech frame sequence is arranged before the first speech frame sequence in the speech stream. The first speech feature includes multiple first sub-features, and the second speech feature determined based on the second speech frame sequence includes multiple second sub-features arranged in sequence. A window sub-feature adjacent to a first target sub-feature among the multiple first sub-features is determined from the multiple sub-features based on a preset window, and the multiple sub-features include at least one of the first sub-features and the second sub-features. Mask the sub-features other than the window sub-feature among the multiple sub-features to obtain the window speech feature corresponding to the first target sub-feature; Based on the attention mechanism, the window speech features and the first target sub-feature are fused to obtain the speech fusion sub-feature corresponding to the first target sub-feature; Based on the aforementioned speech fusion sub-features, determine the speech fusion features; The speech fusion features are converted based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence.
2. The method according to claim 1, wherein, The feature extraction of the first speech frame sequence in the speech stream to obtain the first speech features includes: At least one convolution operation is performed on the first speech frame sequence based on the first convolution kernel to obtain initial speech features; and The initial speech features are convolved at least once based on the second convolution kernel to obtain the first speech features, wherein the stride of the first convolution kernel is greater than the stride of the second convolution kernel.
3. The method according to claim 1 or 2, wherein, The step of performing speech conversion on the speech fusion features based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence includes: The speech fusion features are upsampled and convolutionally processed to obtain the target fusion features; The preset speech attributes and the speech fusion features are fused to obtain the converted speech features; and Based on the converted speech features, the converted speech data corresponding to the first speech frame sequence is determined.
4. A method for training a deep learning model, wherein, The speech feature extraction network of the deep learning model includes a feature extraction layer and a feature fusion layer; the method includes: Acquire a sample speech stream, wherein at least one second speech frame in the sample first speech frame sequence and the sample second speech frame sequence in the sample speech stream overlap, and the sample second speech frame sequence is arranged before the sample first speech frame sequence in the sample speech stream. The feature extraction layer is used to extract features from the first speech frame sequence of the sample to obtain the first speech features of the sample. The first speech feature of the sample is masked to obtain the masked speech feature of the sample; The sample speech fusion features are obtained by using the feature fusion layer to perform attention feature fusion on the sample mask speech features and the sample second speech features determined based on the second speech frame sequence. Based on a self-supervised mechanism, the speech feature extraction network is trained using the sample speech fusion features to obtain a trained deep learning model. The sample speech fusion features are determined based on the following operations: Based on a preset window, a sample window sub-feature is determined from multiple sample sub-features that is adjacent to the sample first target sub-feature among multiple sample first sub-features. The sample mask speech feature includes multiple sample first sub-features, and the sample second speech feature includes multiple sample second sub-features arranged in sequence. The multiple sample sub-features include at least one of the sample first sub-features and the sample second sub-features. Mask the sample sub-features other than the sample window sub-features among the multiple sample sub-features to obtain the sample window speech features corresponding to the first target sub-feature of the sample; Based on the attention mechanism, the speech features of the sample window and the first target sub-feature of the sample are fused to obtain the sample speech fusion sub-feature corresponding to the first target sub-feature of the sample. Based on the sample speech fusion sub-features, the sample speech fusion features are determined.
5. A voice stream processing apparatus, comprising: The first extraction module is used to extract features from the first speech frame sequence in the speech stream to be processed, and obtain the first speech feature. The first speech frame sequence overlaps with at least one second speech frame in the second speech frame sequence. The second speech frame sequence is arranged before the first speech frame sequence in the speech stream. The first speech feature includes multiple first sub-features. The second speech feature determined based on the second speech frame sequence includes multiple second sub-features arranged in sequence. The fusion module is used to fuse the first speech features and the second speech features determined based on the second speech frame sequence using an attention mechanism to obtain speech fusion features; and The conversion module is used to convert the speech fusion features based on preset speech attributes to obtain converted speech data corresponding to the first speech frame sequence. The fusion module is configured as follows: Based on a preset window, a window sub-feature is determined from multiple sub-features that is adjacent to a first target sub-feature among multiple first sub-features, wherein the multiple sub-features include at least one of the first sub-feature and the second sub-feature; Mask the sub-features other than the window sub-feature among the multiple sub-features to obtain the window speech feature corresponding to the first target sub-feature; Based on the attention mechanism, the window speech features and the first target sub-feature are fused to obtain the speech fusion sub-feature corresponding to the first target sub-feature; Based on the aforementioned speech fusion sub-features, speech fusion features are determined.
6. A training device for a deep learning model, wherein, The speech feature extraction network of the deep learning model includes a feature extraction layer and a feature fusion layer; the device includes: An acquisition module is used to acquire a sample speech stream, wherein at least one second speech frame in a sample first speech frame sequence and a sample second speech frame sequence in the sample speech stream overlaps, and the sample second speech frame sequence is arranged before the sample first speech frame sequence in the sample speech stream. The second extraction module is used to extract features from the first speech frame sequence of the sample using the feature extraction layer to obtain the first speech features of the sample. The masking module is used to mask the first speech feature of the sample to obtain the masked speech feature of the sample; The second fusion module is used to perform attention feature fusion on the sample mask speech features and the sample second speech features determined based on the second speech frame sequence using the feature fusion layer to obtain sample speech fusion features. The training module is used to train the speech feature extraction network based on the sample speech fusion features using a self-supervised mechanism, so as to obtain a trained deep learning model. The sample speech fusion features are determined based on the following operations: Based on a preset window, a sample window sub-feature is determined from multiple sample sub-features that is adjacent to the sample first target sub-feature among multiple sample first sub-features. The sample mask speech feature includes multiple sample first sub-features, and the sample second speech feature includes multiple sample second sub-features arranged in sequence. The multiple sample sub-features include at least one of the sample first sub-features and the sample second sub-features. Mask the sample sub-features other than the sample window sub-features among the multiple sample sub-features to obtain the sample window speech features corresponding to the first target sub-feature of the sample; Based on the attention mechanism, the speech features of the sample window and the first target sub-feature of the sample are fused to obtain the sample speech fusion sub-feature corresponding to the first target sub-feature of the sample. Based on the sample speech fusion sub-features, the sample speech fusion features are determined.
7. An intelligent agent, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method of any one of claims 1 to 4 by calling the large model to obtain output information; An output module is used to output the output information obtained by the processing module.
8. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 4.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 4.
10. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Superimposed sound detection method, device and equipment
CN111640456A
Real-time speech recognition method, model training method, device and equipment
CN114596841A