Sample construction method and device, equipment and storage medium

CN121866618APending Publication Date: 2026-04-14FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional speech translation data construction requires a large number of professionals to participate in annotation, resulting in low efficiency and a large amount of manpower consumption.

Method used

By acquiring first speech content and first text content associated with the first language, determining the first time distribution associated with the first text content, and constructing sample data for training the speech processing model based on the correspondence between the first time distribution and the second text content.

Benefits of technology

It improves the efficiency of sample data construction, reduces the waste of human resources, and enables the rapid construction of large batches of high-quality training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121866618A_ABST
    Figure CN121866618A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a sample construction method and device, equipment and a storage medium. The method provided herein includes: acquiring first voice content associated with a first language and first text content corresponding to the first voice content; determining a first time distribution related to the first text content based on the first voice content; based on the first time distribution and the corresponding relation between the second text content and the first text content, second time distribution associated with the second text content is determined, and the second text content is associated with a second language; and based on the first voice content, the second text content and the second time distribution, constructing sample data for training a voice processing model. In this way, the construction efficiency of the sample data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]The technical field of a method, device and equipment for constructing a sample and a storage medium relates to the computer field, and particularly to a method, device, equipment and computer readable storage medium for constructing a sample. Background Art Speech translation data construction refers to a process of translating and labeling text recognized by speech recognition to realize high-quality and multilingual translation model training. Traditional speech translation data construction needs to involve a large number of professionals in labeling, and this method not only consumes a large amount of manpower, but also is inefficient. Summary In a first aspect of the present disclosure, a method for constructing a sample is provided, comprising: obtaining first speech content associated with a first language and first text content corresponding to the first speech content; determining a first time distribution related to the first text content based on the first speech content; determining a second time distribution associated with second text content based on the first time distribution and a corresponding relationship between the second text content and the first text content, the second text content being associated with a second language; and constructing sample data for training a speech processing model based on the first speech content, the second text content and the second time distribution. In a second aspect of the present disclosure, a device for constructing a sample is provided. The device comprises: a text content obtaining module configured to obtain first speech content associated with a first language and first text content corresponding to the first speech content; a first determining module configured to determine a first time distribution related to the first text content based on the first speech content; a second determining module configured to determine a second time distribution associated with second text content based on the first time distribution and a corresponding relationship between the second text content and the first text content, the second text content being associated with a second language; and a sample data constructing module configured to construct sample data for training a speech processing model based on the first speech content, the second text content and the second time distribution. In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect. In a fourth aspect of the present disclosure, a computer readable storage medium is provided. The computer readable storage medium has stored thereon a computer program executable by a processor to implement the method of the first aspect. It should be understood that the content described in this part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description.The above-described and other features and advantages of various embodiments of the present disclosure will be more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference characters refer to like elements throughout. In the drawings: FIG. 1 shows a schematic diagram of an example environment in which some embodiments of the present disclosure can be implemented; FIG. 2 shows a flowchart of a process of constructing a sample according to some embodiments of the present disclosure; FIG. 3 shows an example diagram of streaming sample data construction according to some embodiments of the present disclosure; FIG. 4 shows an example diagram of offline sample data construction according to some embodiments of the present disclosure; FIG. 5 shows a schematic structural block diagram of an example apparatus for constructing a sample according to some embodiments of the present disclosure; and FIG. 6 shows a block diagram of a device capable of implementing various embodiments of the present disclosure. Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and should not be used to limit the scope of protection of the present disclosure. It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or in a different section / subsection in any manner. In the description of embodiments of the present disclosure, the term “comprising” and its conjugations should be understood to be open-ended, i.e., “including but not limited to.” The term “based on” should be understood as “based at least in part on.” The term “one embodiment” or “an embodiment” should be understood as “at least one embodiment.” The term “some embodiments” should be understood as “at least some embodiments.” Other explicit and implicit definitions can be included below. The terms “first,” “second,” etc. can refer to different or same objects. Other explicit and implicit definitions can be included below. Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the present disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms. Accordingly, in the implementation of various embodiments of the present disclosure, the type of data or information that can be involved, the scope of use, the scenario of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.The specific notification and / or authorization manner can vary according to actual conditions and application scenarios, and the scope of the present disclosure is not limited in this regard. In the specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed scope. Users who refuse to process personal information other than the necessary information required for basic functions will not affect the use of basic functions. As briefly mentioned above, the voice translation data construction refers to the process of translating and labeling the text recognized by voice recognition to achieve high-quality, multilingual translation model training. Traditional voice translation data construction requires a large number of professionals to participate in labeling, and this method not only consumes a lot of manpower, but also is not efficient. The embodiments of the present disclosure propose a scheme for constructing samples. According to various embodiments of the present disclosure, first voice content associated with a first language and first text content corresponding to the first voice content are obtained; based on the first voice content, a first time distribution related to the first text content is determined; based on the first time distribution and a corresponding relationship between second text content and the first text content, a second time distribution associated with the second text content is determined, the second text content being associated with a second language; and based on the first voice content, the second text content, and the second time distribution, sample data for training a voice processing model is constructed. In this way, the embodiments of the present disclosure can effectively improve the construction efficiency of sample data. Various example implementations of the scheme are described in detail below in combination with the accompanying drawings. Example environment FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include an electronic device 110 and a target model 136o In the environment 100 of FIG. 1, the electronic device 110 can receive first voice content 130o from a user 140. The electronic device 110 can call an alignment model to process the first voice content 130, and obtain a first time distribution related to the first voice content. Further, the electronic device 110 can input the first text content to the target model 136, and obtain second text content associated with a second language. Then, the electronic device 110 can determine a second time distribution corresponding to the second text content based on the first text content, the second text content, and the first time distribution. Finally, the electronic device 110 can construct sample data 120 based on the first voice content, the second text content, and the second time distribution.In some embodiments, the target model 136 can run on a local device or a remote device. In some embodiments, the electronic device 110 can include various types of computing systems / servers capable of providing computing capabilities, and the electronic device 110 can include an end device. Such an end device can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), a voice / video recorder, a digital still / video camera, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals for these devices or any combination thereof. The electronic device 110 can include various types of computing systems / servers capable of providing computing capabilities, such as mainframes, edge computing nodes, computing devices in a cloud environment, virtual machines, and the like, for example. Although shown as a single device, the electronic device 110 can include multiple physical devices. It should be understood that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure. Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings. Example process diagram of constructing a sample FIG. 2 shows a flowchart of a process 200 of constructing a sample according to some embodiments of the present disclosure. The process 200 can be implemented at the electronic device 110. The process 200 will be described below with reference to FIG. 1. As shown in FIG. 2, at block 210, the electronic device 110 obtains first speech content associated with a first language and first text content corresponding to the first speech content. As an example, as shown in FIG. 3, which shows a streaming sample data construction process of an embodiment of the present disclosure. After obtaining the first speech content 310 associated with the first language, the electronic device 110 can invoke a speech recognition model, such as an ASR (Automatic Speech Recognition) model, to convert the first speech content into the first text content 320. With continued reference to FIG. 2, at block 220, the electronic device 110 determines a first time distribution related to the first text content based on the first speech content. As an example, as shown in FIG. 3, the electronic device 110 can input the first text content 320 and the first speech content 310 into an alignment model 330 to obtain the first time distribution through the alignment model 330.In some embodiments, the electronic device 110 can determine the timestamp information corresponding to multiple characters in the first text content and determine the first time distribution related to the first text content based on the timestamp information. As an example, as shown in FIG. 3, after the electronic device 110 inputs the first text content 320 and the first speech content 310 into the alignment model 330, the alignment model 330 can determine the first time distribution 335 corresponding to the first text content based on the time nodes (i.e., timestamps) at which each character appears in the first speech content 310. For example, if the first text content 320 is "Hello, world", the first time distribution obtained by aligning the first text content 320 with the first speech content 310 by the alignment model 330 is {"Hello": 0.1, "world": 0.2, "world": 0.35, "world": 0.42}. Continuing to refer to FIG. 2, in block 230, the electronic device 110 determines the second time distribution of the second text content based on the first time distribution and the correspondence between the second text content and the first text content, and the second text content is associated with a second language. As an example, the electronic device 110 can match the time nodes corresponding to the second text content from the first time distribution based on the correspondence between the second text content and the first text content, so as to obtain the second time distribution. In some embodiments, the correspondence between the first text content and the second text content indicates the correspondence between the first set of semantic units in the first text content and the second set of semantic units in the second text content. For example, the first text content is "Hello, world", and the second text content is "hello, world". An example of the correspondence between the first text content and the second text content is: "Hello" in the first text content corresponds to "hello" in the second text content, and "world" in the first text content corresponds to "world" in the second text content. In some embodiments, the electronic device 110 can divide the first text content into multiple semantic units. For example, if the first text content is "Hello, world", the multiple semantic units obtained by semantic segmentation of the first text content by the electronic device 110 are "Hello / world / ". Further, the electronic device 110 can determine the timestamps corresponding to multiple delimiters based on the timestamp information, and the multiple delimiters correspond to multiple semantic units. As an example, as shown in FIG. 3, the electronic device 110 can use the time node corresponding to the last character in the semantic unit as the time node corresponding to the semantic unit.As an example, as shown in FIG. 3, the electronic device 110 can determine the time nodes corresponding to each semantic unit in the second text content based on the segmented second text content 346 and the time nodes 350 corresponding to the set of semantic units of the first text content, thereby obtaining a second time distribution 360. With continued reference to FIG. 2, at block 240, the electronic device 110 constructs sample data for training the speech processing model based on the first speech content, the second text content, and the second time distribution. For example, the first speech content can be speech data of “Hello, world”, the second text content can be text data of the first text content translated into another language, e.g., “hello, world”, and the second time distribution can be {“hello”: 0.2, “world”: 0.42}. In some scenarios, the electronic device 110 can segment the first speech content based on the time nodes corresponding to the set of semantic units of the first text content, to obtain a set of speech segments. The sample data for training the speech processing model is constructed based on the set of semantic segments, the second text content, and the second time distribution. In this way, some embodiments of the present disclosure can automatically construct the sample data for training the speech processing model through the language model, effectively improving the construction efficiency of the sample data.In some embodiments, the speech processing model is configured to process a speech data stream to generate a translation result of the speech data stream, the translation structure including translated text and / or speech corresponding to a second language. As an example, the input data of the speech processing model is audio data of "hello, world", and the output result obtained by the electronic device 110 after inputting the input data to the speech processing model is audio data and text data of "hello, world". In some embodiments, the electronic device 110 can perform speech recognition on the first speech content to obtain speech recognition text. As an example, as shown in FIG. 4, FIG. 4 illustrates a process of offline construction of sample data according to an embodiment of the present application. The electronic device 110 first receives speech audio input by the user 140, converts the speech audio into text form through a speech recognition model such as an ASR model, and obtains speech recognition text 410. Further, the electronic device 110 can determine the first text content associated with the first language by optimizing the speech recognition text. The process of optimizing the speech recognition text may, for example, be performing inverse text regularization and / or performing a spoken language smoothing process. As an example, as shown in FIG. 4, since the speech recognition text 410 obtained by the ASR model has no punctuation and sentence division, the quality of the translated text directly based on the speech recognition text 410 is low. Based on this, the electronic device 110 can perform inverse text regularization and spoken language smoothing processing on the speech recognition text 410 through a target model 420 to obtain the first text content 430 associated with the first language. The target model 420 may, for example, be a language model. Finally, the electronic device 110 can generate the second text content associated with the second language based on the first text content. As an example, as shown in FIG. 4, the electronic device 110 can input the first text content to the target model 420 to generate the second text content associated with the second language. In some embodiments, the electronic device 110 can translate the first text content into intermediate text content associated with the second language. Further, the electronic device 110 can optimize the intermediate text content using a language model to obtain the second text content. As an example, the electronic device 110 can use part of the speech recognition text and the first text content as a component of a prompt word to activate the inference capability of the target model 420 based on the prompt word, so that the target model 420 knows the task to be performed (i.e., the translation task). Further, the electronic device 110 inputs the first text content to the target model 420 to generate the intermediate text content. The user 140 can instruct the target model 420 to correct the intermediate text content via the electronic device 110, thereby obtaining the second text content 440.Based on this, the thinking chain for constructing offline sample data is built. After obtaining the intermediate text content, the intermediate text content is corrected again, which can effectively improve the translation quality of the target model. In some embodiments, the electronic device 110 can correct at least one translation error in the intermediate text content and simplify the expression of the intermediate text content based on the first text content and / or the speech recognition text. As an example, as shown in FIG. 4, after obtaining the intermediate text content, the electronic device 110 can instruct the target model 420 to correct the intermediate text content based on the first text content 430 and the language recognition text 410 to correct the translation error, and the electronic device 110 can also instruct the target model 420 to simplify the expression of the intermediate text content. In addition, the electronic device 110 can also instruct the target model 420 to correct the missing translation and multiple translation problems in the intermediate text. In this way, the quality of the offline speech data constructed can be effectively improved. Based on the above-described process, the embodiments of the present disclosure can significantly improve the efficiency of data construction and the consistency of data standards by constructing speech translation data through a model, reduce the waste of human resources, and quickly construct a large amount of high-quality training data in a short time. The embodiments of the present disclosure also provide a corresponding device for implementing the above-mentioned method or process. FIG. 5 shows a schematic structural block diagram of an example device 500 for constructing samples according to certain embodiments of the present disclosure. The device 500 can be implemented as or included in an electronic device. The various modules / components in the device 500 can be implemented by hardware, software, firmware, or any combination thereof. As shown in FIG. 5, the device 500 includes a text content acquisition module 510 configured to acquire first speech content associated with a first language and first text content corresponding to the first speech content; a first determination module 520 configured to determine a first time distribution related to the first text content based on the first speech content; a second determination module 530 configured to determine a second time distribution associated with second text content based on the first time distribution and a correspondence between the second text content and the first text content, the second text content being associated with a second language; and a sample data construction module 540 configured to construct sample data for training a speech processing model based on the first speech content, the second text content, and the second time distribution. In some embodiments, the device 500 further includes a second text content generation module configured to perform speech recognition on the first speech content to obtain speech recognition text; determine the first text content associated with the first language by optimizing the speech recognition text; and generate the second text content associated with the second language based on the first text content.In some embodiments, the second text content generation module is further configured to: perform an inverse text regularization process and / or perform a spoken language smoothing process. In some embodiments, the second text content generation module is further configured to: translate the first text content into intermediate text content associated with a second language; optimize the intermediate text content using a language model to obtain the second text content. In some embodiments, the language model is instructed to perform at least one of the following optimization processes: correct at least one translation error in the intermediate text content based on the first text content and / or speech-recognized text; simplify the expression of the intermediate text content. In some embodiments, the correspondence between the first text content and the second text content indicates: the correspondence between a first set of semantic units in the first text content and a second set of semantic units in the second text content. In some embodiments, the first determining module 520 is further configured to: determine timestamp information corresponding to multiple characters in the first text content; and determine a first time distribution associated with the first text content based on the timestamp information. In some embodiments, the first determining module 520 is further configured to: segment the first text content into a plurality of semantic units; determine a plurality of timestamps corresponding to a plurality of delimiters based on timestamp information, the plurality of delimiters corresponding to a plurality of semantic units; and determine a second time distribution associated with the second text content based on the plurality of timestamps. In some embodiments, the speech processing model is configured to process a speech data stream to generate a translation result of the speech data stream, the translation result including translated text and / or translated speech corresponding to a second language. FIG6 shows a block diagram of a computing device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the computing device 600 shown in FIG6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The computing device 600 shown in FIG6 can be used to implement the electronic device in FIG1. ​​As shown in FIG6, the computing device 600 is in the form of a general-purpose computing device. The components of the computing device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. oThe processing unit 610 can be a real or virtual processor and capable of performing various processing according to programs stored in the memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing power of the computing device 600. The computing device 600 typically includes a plurality of computer storage media. Such media can be removable computer storage media 630 and / or non-removable computer storage media implemented in a suitable device or devices. The memory 620 can be volatile (such as registers, cache, random access memory (RAM)), non-volatile (such as read-only memory (ROM), EEPROM, flash memory), or some combination of the two. The storage device 630 can be a removable storage and / or non-removable storage including, but not limited to, magnetic disks, optical disks, or tape. Such storage media can be used to store data and / or computer programs and can be accessed by a computer. The computing device 600 can further include other removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile memory media such as a floppy disk, a ZIP® disk, a magnetic tape, or an optical disk. In such instances, each drive can be connected to the bus by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure. The communication unit 640 enables communication with other electronic devices over a communication media. Additionally, the functionality of the components of the computing device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, the computing device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in a networking environment. The input device 650 can be one or more input devices such as a mouse, a keyboard, a trackball, a remote control, etc. The output device 660 can be one or more output devices such as a display, a projector, a television, etc.The computing device 600 can also communicate with one or more external devices (not shown) such as a storage device or other computing devices using the communication unit 640, in accordance with one or more embodiments of the present disclosure. In this regard, the communication unit 640 can include a wired communication device and / or a wireless communication device to enable the computing device 600 to communicate with one or more devices. The communication can be facilitated via an input / output (I / O) interface (not shown) in accordance with exemplary implementations of the present disclosure. In accordance with exemplary implementations of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, wherein the computer executable instructions are executed by a processor to implement the methods described above. In accordance with exemplary implementations of the present disclosure, a computer program product is also provided tangibly stored on a non-transitory computer readable medium and including computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the methods described above. Various aspects of the disclosure can be described in the context of flow diagrams and / or block diagrams that illustrate the functions and operations of methods, apparatuses, devices, and computer program products according to implementations of the present disclosure. It will be understood that each block of the flow diagrams and / or block diagrams, and combinations of blocks in the flow diagrams and / or block diagrams, can be implemented by computer readable program instructions. These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flow diagrams and / or block diagrams. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flow diagrams and / or block diagrams. The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flow diagrams and / or block diagrams. The flow diagrams and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various implementations of the present disclosure.In this regard, each block within a flowchart or block diagram can represent a module, a portion of a program, or a portion of instructions that comprise an executable procedure that operates to implement a specified logical function. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs a specified function or action, or a combination of special-purpose hardware and computer instructions. Having thus described the functionality of various implementations of the disclosure, it is manifest to those skilled in the art that various modifications can be made without departing from the scope and spirit of the described implementations. Numerous other implementations will be apparent to those skilled in the art in view of the foregoing description. The embodiments disclosed herein are to be considered merely exemplary and are not restrictive in nature. The scope of the various implementations of the disclosure is not to be limited to the embodiments disclosed herein but is only limited by the claims that follow.

Claims

CLAIM 1. A method of constructing a sample, comprising: obtaining first speech content associated with a first language and first text content corresponding to the first speech content; determining, based on the first speech content, a first time distribution related to the first text content; determining, based on the first time distribution and a correspondence between second text content and the first text content, a second time distribution associated with the second text content, the second text content being associated with a second language; and constructing, based on the first speech content, the second text content, and the second time distribution, sample data for training a speech processing model.

2. The method of claim 1, further comprising: performing speech recognition on the first speech content to obtain speech recognition text; determining, based on the first text content, the second text content associated with the second language.

3. The method of claim 2, wherein determining, by optimizing the speech recognition text, the first textual content associated with the first language comprises: performing an inverse text regularization process and / or performing a spoken language smoothing process.

4. The method of claim 2, wherein generating the second text content associated with the second language based on the first text content comprises: translating the first text content into intermediate text content associated with the second language; optimizing the intermediate text content using a language model to obtain the second text content.

5. The method of claim 4, wherein the language model is instructed to perform at least one optimization process including: correcting at least one translation error in the intermediate text content based on the first text content and / or the speech recognition text; and simplifying expressions in the intermediate text content.

6. The method of claim 1, wherein the correspondence between the first text content and the second text content indicates a correspondence between a first set of semantic units in the first text content and a second set of semantic units in the second text content.

7. The method of claim 1, wherein determining, based on the first speech content, a first temporal profile related to the first text content comprises: determining timestamp information corresponding to a plurality of characters in the first text content; and determining, based on the timestamp information, the first time distribution related to the first text content. segmenting the first text content into a plurality of semantic units; 8. The method of claim 7, wherein determining a second temporal profile of the second textual content based on the first temporal profile and a correspondence between the second textual content and the first textual content comprises: determining, based on the timestamp information, a plurality of timestamps corresponding to a plurality of delimiters, the plurality of delimiters corresponding to the plurality of semantic units; and determining, based on the plurality of timestamps, the second time distribution associated with the second text content.

9. The method of claim 1, wherein the speech processing model is configured to process a speech data stream to generate a translation result of the speech data stream, the translation result including translated text and / or translated speech corresponding to the second language. a text content obtaining module configured to obtain first speech content associated with a first language and first text content corresponding to the first speech content; a first determining module configured to determine, based on the first speech content, a first time distribution related to the first text content; 10. An apparatus for constructing a sample, comprising: a second determining module configured to determine, based on the first time distribution and a correspondence between second text content and the first text content, a second time distribution associated with the second text content, the second text content being associated with a second language; and a constructing module configured to construct, based on the first speech content, the second text content, and the second time distribution, sample data for training a speech processing model. ​ ​ a correspondence between the first text content and the first speech content, determining a second time distribution associated with the second text content, the second text content being associated with a second language; and a sample data construction module configured to construct, based on the first speech content, the second text content, and the second time distribution, sample data for training a speech processing model.

11. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to carry out the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9. 16