Voice text matching method and device based on reinforcement learning, equipment and medium
By constructing a semantic feature space and using reinforcement learning, the optimal matching path for the target is determined, and the speech-text matching model is optimized. This solves the problem of insufficient accuracy in speech-text matching in existing technologies and achieves higher matching accuracy and diversity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YUNSHANG TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, cross-modal speech-text matching methods rely on fixed feature alignment mechanisms and lack adaptive feature matching capabilities, which makes it impossible to effectively identify and process long-distance semantic dependencies between speech and text, thus reducing the accuracy of speech-text matching.
By constructing a semantic feature space, utilizing key anchors and cumulative reward functions in reinforcement learning, the optimal matching path for the target is determined, the speech-text matching model is optimized, the information acquisition capability within the semantic feature space is enhanced, and the model is further optimized using the optimal matching path for the target to improve the accuracy of the matching results.
It improves the accuracy of speech-text matching, enhances the diversity and comprehensiveness of matching between candidate training samples to be matched and candidate training samples to be matched, and improves the accuracy of the model in obtaining matching results.
Smart Images

Figure CN122024705A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for voice-text matching based on reinforcement learning. Background Technology
[0002] With the rapid development of multimedia technology, cross-modal speech-text matching technology is increasingly widely used in fields such as speech recognition, intelligent human-computer interaction, and multimedia retrieval. This technology aims to build semantic relationships between speech data and text data, achieving intelligent matching of data from different modalities. However, traditional cross-modal speech-text matching methods mainly rely on static feature extraction and simple similarity measurements, making it difficult to fully capture the complex semantic relationships between speech and text.
[0003] Based on this, deep learning technology has made significant breakthroughs in feature extraction and pattern recognition in recent years. Advanced models such as convolutional neural networks and recurrent neural networks have been applied to feature extraction of speech and text, and methods such as metric learning have been introduced to construct efficient cross-modal matching models. Meanwhile, reinforcement learning, as a method that learns optimal decision-making strategies through continuous interaction with the environment, has shown great potential in handling complex decision-making problems; however, its application in cross-modal matching is still in its early stages of exploration. For example, metric learning-based methods can utilize Siamese or Triplet networks to embed speech and text into a shared semantic space through contrastive learning, and evaluate the degree of matching using Euclidean distance or cosine similarity. Furthermore, attention-based alignment models can be employed, using cross-modal Transformers and their multi-head attention mechanisms to achieve dynamic alignment of speech frames and text words, thereby completing speech-text matching.
[0004] However, existing technologies, relying on fixed feature alignment mechanisms, lack the ability for adaptive feature matching. Furthermore, the speech-to-text matching process focuses too much on calculating the similarity of local features, neglecting global semantic consistency, thus failing to effectively identify and handle long-distance semantic dependencies between speech and text. These limitations collectively lead to a decrease in the accuracy of speech-to-text matching. Summary of the Invention
[0005] The purpose of this invention is to provide a speech-to-text matching method based on reinforcement learning, which addresses the problem that existing technologies lack adaptive feature matching capabilities due to their reliance on fixed feature alignment mechanisms. Furthermore, the speech-to-text matching process focuses excessively on calculating the similarity of local features, neglecting global semantic consistency, thus failing to effectively identify and process long-distance semantic dependencies between speech and text, thereby reducing the accuracy of speech-to-text matching.
[0006] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a speech-text matching method based on reinforcement learning, comprising: Obtain multiple training samples to be matched and matching training samples corresponding to each training sample to be matched. The training sample to be matched is either the speech feature to be matched or the text feature to be matched. The matching training sample is either the speech feature to be matched or the text feature to be matched. The speech feature to be matched and the text feature to be matched are in one-to-one correspondence. Based on the multiple training samples to be matched and the multiple training samples to be matched, a semantic feature space is constructed, wherein the semantic feature space includes: multiple semantic feature candidate spaces, and each semantic feature candidate space includes: a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. Based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function, a target optimal matching path is determined for each candidate training sample to be matched, wherein the target optimal matching path is used to determine the candidate training sample to be matched. Based on the cumulative reward value corresponding to the optimal matching path of the target, the speech-text matching model is updated, and the relevant matching data of the current multiple candidate training samples to be matched are stored in the experience pool. Then, the execution is returned to construct the semantic feature space based on the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched, until the training termination condition is reached, and the trained target speech-text matching model is obtained. The target speech-text matching model is used to obtain the matching result of the target object to be matched.
[0007] In one embodiment, constructing a semantic feature space based on the plurality of training samples to be matched and the plurality of matching training samples includes: Based on the multiple training samples to be matched and the multiple matching training samples, an initial semantic feature space is constructed; Based on the distribution density of the multiple training samples to be matched and the multiple matching training samples, the initial semantic feature space is divided to obtain an intermediate semantic feature space. The intermediate semantic feature space includes multiple intermediate semantic feature subspaces, and each intermediate semantic feature subspace includes multiple intermediate training samples to be matched and multiple intermediate matching training samples. Based on the semantic center vector of each intermediate semantic feature subspace, the initial anchor position of each intermediate semantic feature subspace is initialized, wherein the semantic center vector is determined based on multiple intermediate training samples to be matched and multiple intermediate matching training samples. Based on the initial anchor point position, the preset radius, and the matching success rate of the training samples to be matched and the matching training samples in the intermediate semantic feature subspace, the initial anchor point position is adjusted to obtain the key anchor point, and the key anchor point is verified. Based on the verification results, multiple semantic feature candidate spaces are determined in multiple intermediate semantic feature subspaces. Each semantic feature candidate space includes: a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. The semantic feature space is constructed based on multiple semantic feature candidate spaces.
[0008] In one embodiment, determining the target optimal matching path corresponding to each of the multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function includes: Based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, and multiple candidate training samples to be matched, multiple candidate matching paths are determined for each candidate training sample to be matched, wherein the candidate matching paths are used to determine the candidate matching training samples corresponding to the candidate training samples to be matched. The cumulative reward value corresponding to multiple candidate matching paths is obtained according to the preset cumulative reward function, and the candidate matching path corresponding to the largest cumulative reward value is determined as the target optimal matching path based on the multiple cumulative reward values.
[0009] In one embodiment, determining multiple candidate matching paths corresponding to each candidate training sample to be matched based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, and multiple candidate training samples to be matched includes: Starting from each key anchor point, the matching is performed on multiple candidate training samples to be matched and multiple candidate training samples to be matched corresponding to each key anchor point to obtain multiple initial candidate matching paths. Each initial candidate matching path includes: a first initial candidate matching sub-path corresponding to the candidate training sample to be matched and a second initial candidate matching sub-path corresponding to the candidate training sample to be matched. For the first initial candidate matching sub-path and the second initial candidate matching sub-path, the path is expanded in multiple candidate training samples to be matched and multiple candidate training samples corresponding to all key anchor points to obtain the first candidate matching sub-path corresponding to the first initial candidate matching sub-path and the second candidate matching sub-path corresponding to the second initial candidate matching sub-path. The first candidate matching sub-path and the second candidate matching sub-path are combined to obtain a candidate matching path.
[0010] In one embodiment, the candidate matching path includes: multiple matching connection nodes, wherein the matching connection node is a candidate training sample to be matched, or a candidate training sample that matches the candidate training sample to be matched, and the step of obtaining the cumulative reward value corresponding to the multiple candidate matching paths according to a preset cumulative reward function includes: For each candidate matching path, the local activation value of each matching connection node is obtained based on the cosine similarity. Obtain the global incentive value for each candidate matching path; The path reward value is obtained by weighted summation of the local incentive value and the global incentive value; The path reward value is substituted into a preset cumulative reward function for calculation to obtain the cumulative reward value corresponding to the candidate matching path.
[0011] In one embodiment, the speech-to-text matching model includes a policy network and a value network, and updating the speech-to-text matching model based on the cumulative reward value corresponding to the target optimal matching path includes: The strategy gradient is determined based on the cumulative reward value corresponding to the optimal matching path of the target. Adjust the weight parameters of the policy network according to the policy gradient; The value error is determined based on the cumulative reward value and the preset value error loss function; The weight parameters of the value network are adjusted based on the value error.
[0012] In one embodiment, the method further includes: Obtain the target object to be matched, wherein the target object to be matched is either the speech feature target object to be matched or the text feature target object to be matched; The target object to be matched is input into the target speech-text matching model, and the matching result of the target object is obtained through the policy network in the target speech-text matching model.
[0013] Secondly, embodiments of the present invention provide a speech-text matching device based on reinforcement learning, comprising: The training sample acquisition module is used to acquire multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched. The training sample to be matched is any one of the speech feature to be matched and the text feature to be matched. The matching training sample is any one of the speech feature to be matched and the text feature to be matched. The speech feature to be matched and the text feature to be matched correspond one-to-one. The semantic feature space construction module is used to construct a semantic feature space based on the multiple training samples to be matched and the multiple matching training samples. The semantic feature space includes multiple semantic feature candidate spaces, and each semantic feature candidate space includes a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. The target optimal matching path determination module is used to determine the target optimal matching path corresponding to each of the multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function, wherein the target optimal matching path is used to determine the candidate training samples to be matched corresponding to the candidate training samples to be matched. The target speech-text matching model acquisition module is used to update the speech-text matching model according to the cumulative reward value corresponding to the optimal matching path of the target, and store the relevant matching data of multiple candidate training samples to be matched into the experience pool. Then, it returns to execute the construction of semantic feature space according to the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched, until the training termination condition is reached, and obtains the trained target speech-text matching model. The target speech-text matching model is used to obtain the matching result of the target object to be matched.
[0014] Thirdly, embodiments of the present invention provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the reinforcement learning-based speech-text matching method described in the first aspect.
[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the reinforcement learning-based speech-text matching method described in the first aspect.
[0016] The technical solution provided by the embodiments of the present invention has the following advantages compared with the prior art: This invention provides a speech-text matching method based on reinforcement learning. It acquires multiple training samples to be matched and corresponding matching training samples for each training sample. Each training sample to be matched is either a speech feature to be matched or a text feature to be matched, and each matching training sample is either a speech feature to be matched or a text feature to be matched. There is a one-to-one correspondence between the speech features to be matched and the text features to be matched, and vice versa. Based on these multiple training samples, a semantic feature space is constructed. This semantic feature space includes multiple semantic feature candidate spaces, each including a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. Based on the multiple key anchor points, the multiple candidate training samples to be matched corresponding to each key anchor point, the multiple candidate training samples to be matched, and a preset cumulative reward function, a target optimal matching path is determined for each candidate training sample to be matched. This target optimal matching path is used to determine the candidate training samples to be matched corresponding to the candidate training samples to be matched. Based on the cumulative reward value corresponding to the optimal matching path, the speech-text matching model is updated, and the relevant matching data of multiple candidate training samples to be matched are stored in the experience pool. The process then returns to the previous step, constructing a semantic feature space based on the multiple candidate training samples and their corresponding matching training samples, until the training termination condition is met. This yields the trained target speech-text matching model, which is used to obtain the matching results for the target object. This approach enhances the ability to obtain semantic alignment information with the candidate training samples within the semantic feature space, based on the key anchor points corresponding to multiple candidate semantic feature spaces. This increases the diversity and comprehensiveness of the matching between the candidate training samples and the matching candidate training samples. Furthermore, by utilizing the cumulative reward value corresponding to the optimal matching path, the speech-text matching model is optimized, thereby improving the accuracy of the model in obtaining the matching results for the target object. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 A flowchart illustrating a speech-text matching method based on reinforcement learning, provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a speech-text matching device based on reinforcement learning, provided as an embodiment of the present invention. Detailed Implementation
[0018] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.
[0019] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0020] In this invention, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between the associated objects, indicating that three relationships can exist.
[0021] like Figure 1 As shown, Figure 1 A flowchart illustrating a reinforcement learning-based speech-to-text matching method provided in this embodiment of the invention includes the following steps: S10: Obtain multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched.
[0022] Among them, the training sample to be matched is either the speech feature to be matched or the text feature to be matched, and the matching training sample is either the matching speech feature or the matching text feature. The speech feature to be matched and the matching text feature are in one-to-one correspondence, and the text feature to be matched and the matching speech feature are in one-to-one correspondence.
[0023] The speech features to be matched and the matching speech features can be extracted using speech feature extraction networks such as wav2vec2.0, VGGish, or convolutional recurrent hybrid networks. Specifically, the speech feature extraction network extracts context-aware temporal frame-level feature vector sequences or globally aggregated feature vectors from the original speech waveform or Mel spectrogram. The text features to be matched and the matching text features can be extracted using text feature extraction networks such as BERT, RoBERTa, or BiLSTM-Attention networks. Specifically, the text feature extraction network encodes the input text into context-sensitive word-level context embedding sequences and sentence-level semantic vectors. It should be noted that the dimensions of the speech features and text features are unified to the same dimension through a linear projection layer.
[0024] S11: Construct a semantic feature space based on multiple training samples to be matched and multiple matching training samples.
[0025] The semantic feature space includes multiple semantic feature candidate spaces. Each semantic feature candidate space includes a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. There is a one-to-one correspondence between the candidate training samples to be matched and the candidate training samples to be matched. The key anchor point is used to obtain the candidate training samples to be matched corresponding to each candidate training sample to be matched.
[0026] Specifically, multiple training samples to be matched and matching training samples corresponding to each training sample to be matched are obtained. After obtaining multiple training samples to be matched and matching training samples corresponding to each training sample to be matched, a semantic feature space is constructed based on the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched. The semantic feature space includes multiple semantic feature candidate spaces. Each semantic feature candidate space includes a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. There is a one-to-one correspondence between the candidate training samples to be matched and the candidate training samples to be matched.
[0027] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S11 may be: S111: Construct an initial semantic feature space based on multiple training samples to be matched and multiple matching training samples.
[0028] Specifically, an initial semantic feature space is constructed based on the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched.
[0029] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S111 may be: Multiple training samples to be matched, and the matching training samples corresponding to each training sample to be matched, are concatenated, weighted, or fused across modalities to generate multiple joint feature vectors, thereby obtaining the initial semantic feature space.
[0030] S112: Based on the distribution density of multiple training samples to be matched and multiple matching training samples, the initial semantic feature space is divided to obtain the intermediate semantic feature space.
[0031] The intermediate semantic feature space includes multiple intermediate semantic feature subspaces, and each intermediate semantic feature subspace includes multiple intermediate training samples to be matched and multiple intermediate matching training samples.
[0032] Specifically, for multiple training samples to be matched and multiple matching training samples, the distribution density is calculated using a density-aware clustering algorithm. Then, based on the distribution density, the initial semantic feature space is divided to obtain an intermediate semantic feature space. The intermediate semantic feature space includes multiple intermediate semantic feature subspaces, and each intermediate semantic feature subspace includes multiple intermediate training samples to be matched and multiple intermediate matching training samples.
[0033] It should be noted that for each intermediate semantic feature subspace, each intermediate training sample to be matched and the intermediate matching training sample corresponding to each intermediate training sample to be matched have an internal similarity greater than 0.75 with other intermediate training samples to be matched and their corresponding intermediate matching training samples, and a similarity less than 0.3 between each intermediate semantic feature subspace. This ensures that each intermediate training sample to be matched and each intermediate matching training sample in each intermediate semantic feature subspace have similar cross-modal semantic relevance.
[0034] S113: Initialize the initial anchor point position of each intermediate semantic feature subspace based on the semantic center vector of each intermediate semantic feature subspace.
[0035] The semantic center vector is determined based on multiple intermediate training samples to be matched and multiple intermediate matching training samples.
[0036] Optionally, based on the above embodiments, in some embodiments of the present invention, the semantic center vector is determined based on multiple intermediate training samples to be matched and multiple intermediate matching training samples. One possible implementation is as follows: The initial semantic center vector is obtained by averaging the feature vectors of multiple intermediate training samples to be matched and multiple intermediate matching training samples. The cosine similarity between each intermediate training sample to be matched and the corresponding intermediate matching training sample and the initial semantic center vector is obtained. The semantic center vector is obtained by averaging the feature vectors of the first N intermediate training samples to be matched and the corresponding intermediate matching training samples according to the magnitude of the cosine similarity.
[0037] Specifically, based on the multiple intermediate training samples to be matched and the multiple intermediate matching training samples included in each intermediate semantic feature subspace, the semantic center vector of each intermediate semantic feature subspace is calculated, and the initial anchor point position of each intermediate semantic feature subspace is initialized based on the semantic center vector.
[0038] S114: Based on the initial anchor point position, preset radius, and the matching success rate of the training samples to be matched and the matching training samples in the intermediate semantic feature subspace, adjust the initial anchor point position to obtain the key anchor point, and verify the key anchor point.
[0039] The preset radius can be the size of each intermediate semantic feature subspace, or it can be an adaptively adjustable radius.
[0040] Specifically, for the initial anchor point position of each intermediate semantic feature subspace, the initial anchor point position of each intermediate semantic feature subspace is adjusted within a preset radius, centered on the initial anchor point position, based on the matching success rate of the training samples to be matched and the matching training samples in the intermediate semantic feature subspace, to obtain key anchor points. The key anchor points of multiple intermediate semantic feature subspaces after adjustment are then verified.
[0041] Optionally, based on the above embodiments, in some embodiments of the present invention, one way to verify key anchor points may be: For each intermediate semantic feature subspace, the adjusted key anchor points are used to calculate the cross-modal cosine similarity of each intermediate training sample to be matched and its corresponding intermediate matching training sample included in each intermediate semantic feature subspace. The average semantic similarity is obtained by averaging the multiple cross-modal cosine similarities. Each average semantic similarity is compared with a preset threshold to determine whether the average semantic similarity is greater than the preset threshold.
[0042] It should be noted that cross-modal cosine similarity is used to characterize the alignment quality of the intermediate training samples to be matched in the intermediate semantic feature subspace and their corresponding intermediate matching training samples.
[0043] S115: Based on the verification results, determine multiple semantic feature candidate spaces in multiple intermediate semantic feature subspaces.
[0044] Each semantic feature candidate space includes: a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched.
[0045] Optionally, based on the above embodiments, the verification results include: the average semantic similarity is greater than a preset threshold, or the average semantic similarity is not greater than a preset threshold. Therefore, in some embodiments of the present invention, one implementation of S115 may be: When the average semantic similarity is greater than a preset threshold, the intermediate semantic feature subspace corresponding to that average semantic similarity is retained. When the average semantic similarity is not greater than the preset threshold, the intermediate semantic feature subspace corresponding to that average semantic similarity is filtered out. Based on this, multiple semantic feature candidate spaces are determined from multiple intermediate semantic feature subspaces.
[0046] S116: Construct a semantic feature space based on multiple semantic feature candidate spaces.
[0047] Specifically, after obtaining multiple semantic feature candidate spaces, a semantic feature space is constructed based on these multiple semantic feature candidate spaces.
[0048] S12: Based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function, determine the target optimal matching path corresponding to each candidate training sample to be matched.
[0049] Among them, the target optimal matching path is used to determine the matching candidate training sample corresponding to the candidate training sample to be matched.
[0050] Specifically, after obtaining multiple key anchor points, the optimal target matching path corresponding to each candidate training sample to be matched is determined based on the multiple key anchor points, the multiple candidate training samples to be matched corresponding to the multiple key anchor points, the multiple candidate training samples to be matched, and the preset cumulative reward function.
[0051] Optionally, based on the above embodiments, in some embodiments of the present invention, S12 may be implemented as follows: S121: Based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, and multiple candidate training samples to be matched, determine multiple candidate matching paths corresponding to each candidate training sample to be matched.
[0052] Among them, the candidate matching path is used to determine the candidate matching training sample corresponding to the candidate training sample to be matched.
[0053] Optionally, based on the above embodiments, in some embodiments of the present invention, one implementation of S121 may be: S1211: Starting from each key anchor point, match multiple candidate training samples to be matched and multiple candidate training samples to be matched corresponding to each key anchor point to obtain multiple initial candidate matching paths.
[0054] Each initial candidate matching path includes: a first initial candidate matching sub-path corresponding to the candidate training sample to be matched, and a second initial candidate matching sub-path corresponding to the candidate training sample to be matched. Both the first and second initial candidate matching sub-paths include: a starting connection matching node, an intermediate connection matching node, and a target connection matching node. The starting connection matching node is used to represent the potential center of semantic alignment between the candidate training sample to be matched and the candidate training sample to be matched. The intermediate connection matching node is used to represent the semantic transition points between each candidate training sample to be matched and the candidate training sample to be matched in the semantic feature candidate space. The target connection matching node is used to represent the key features of the candidate training sample to be matched.
[0055] For example, using the first key anchor point as the starting connection matching node A, among multiple candidate training samples to be matched, the intermediate connection matching node B most semantically relevant to the starting connection matching node A is obtained. The starting connection matching node A and the intermediate connection matching node B are connected, and then the intermediate connection matching node B is connected to the target connection matching node C, resulting in the first initial candidate matching sub-path A→B→C. Among multiple candidate training samples, using the first key anchor point as the starting connection matching node A, among multiple candidate training samples to be matched, the intermediate connection matching node D most semantically relevant to the starting connection matching node A is obtained. The starting connection matching node A and the intermediate connection matching node D are connected, and then the intermediate connection matching node D is connected to the target connection matching node E, resulting in the second initial candidate matching sub-path A→D→E. However, this invention is not limited to this, and those skilled in the art can set it according to the actual situation.
[0056] S1212: For the first initial candidate matching sub-path and the second initial candidate matching sub-path, perform path expansion in multiple candidate training samples to be matched and multiple candidate training samples corresponding to all key anchor points to obtain the first candidate matching sub-path corresponding to the first initial candidate matching sub-path and the second candidate matching sub-path corresponding to the second initial candidate matching sub-path.
[0057] Specifically, for the first initial candidate matching sub-path, path expansion is performed in multiple candidate training samples to be matched corresponding to all key anchor points to obtain the first candidate matching sub-path corresponding to the first initial candidate matching sub-path. For the second initial candidate matching sub-path, path expansion is performed in multiple candidate training samples to match corresponding to all key anchor points to obtain the second candidate matching sub-path corresponding to the second initial candidate matching sub-path.
[0058] For example, continuing the above example, for the first initial candidate matching sub-path A→B→C, based on the semantic relevance matrix corresponding to the semantic feature candidate space, among the multiple candidate training samples to be matched corresponding to all key anchor points, obtain the adjacent connection matching node F whose semantic similarity to the target connection matching node C is greater than a preset semantic similarity threshold, connect the target connection matching node C with the adjacent connection matching node F, further obtain the adjacent connection matching node G whose semantic similarity to the adjacent connection matching node F is greater than a preset semantic similarity threshold, and connect it, until all the multiple candidate training samples to be matched corresponding to all key anchor points have been traversed, and the first candidate matching sub-path such as A→B→C→F→G is obtained. For the second initial candidate matching sub-path A→D→E, based on the semantic relevance matrix corresponding to the semantic feature candidate space, among multiple matching candidate training samples corresponding to all key anchor points, adjacent connecting matching nodes H with a semantic similarity greater than a preset semantic similarity threshold to the target connecting matching node E are obtained. The target connecting matching node E is then connected to the adjacent connecting matching node H. Further, adjacent connecting matching nodes J with a semantic similarity greater than a preset semantic similarity threshold to the adjacent connecting matching node H are obtained and connected, until all multiple matching candidate training samples corresponding to all key anchor points have been traversed, resulting in the second candidate matching sub-path such as A→D→E→H→J. However, this invention is not limited to this; those skilled in the art can set it according to the actual situation.
[0059] S1213: Perform a synthesis process on the first candidate matching sub-path and the second candidate matching sub-path to obtain a candidate matching path.
[0060] Specifically, after obtaining the first candidate matching sub-path and the second candidate matching sub-path, the first candidate matching sub-path and the second candidate matching sub-path are combined to obtain the candidate matching path.
[0061] Optionally, based on the above embodiments, both the first candidate matching sub-path and the second candidate matching sub-path include multiple matching connection nodes. Based on this, in some embodiments of the present invention, one implementation of S1213 may be: determining the matching connection nodes that are the same in the first candidate matching sub-path and the second candidate matching sub-path, and performing a synthesis process on the first candidate matching sub-path and the second candidate matching sub-path according to the matching connection nodes to obtain a candidate matching path.
[0062] For example, continuing the above example, the first candidate matching sub-path is such as A→B→C→F→G, and the second candidate matching sub-path is such as A→D→E→H→J. The common matching connection node between the first and second candidate matching sub-paths is identified as A. Based on this, and based on the matching connection node A, the first candidate matching sub-path A→B→C→F→G and the second candidate matching sub-path A→D→E→H→J are synthesized to obtain the candidate matching path J←H←E←D←A→B→C→F→G. However, this invention is not limited to this; those skilled in the art can set it according to the actual situation.
[0063] S122: Obtain the cumulative reward value corresponding to multiple candidate matching paths according to the preset cumulative reward function, and determine the candidate matching path corresponding to the largest cumulative reward value as the target optimal matching path based on the multiple cumulative reward values.
[0064] Specifically, the cumulative reward value corresponding to multiple candidate matching paths is obtained according to the pre-set cumulative reward function. The size of the multiple cumulative reward values is compared, and the candidate matching path corresponding to the largest cumulative reward value is determined as the target optimal matching path.
[0065] Optionally, based on the above embodiments, the candidate matching path includes: multiple matching connection nodes, where each matching connection node is a candidate training sample to be matched, or a matching candidate training sample that matches the candidate training sample to be matched. Therefore, in some embodiments of the present invention, one implementation of S122 may be: S1221: For each candidate matching path, obtain the local activation value of each matching connection node based on cosine similarity.
[0066] S1222: Obtain the global incentive value for each candidate matching path.
[0067] S1223: Perform a weighted summation of the local and global incentive values to obtain the path reward value.
[0068] S1224: Substitute the path reward value into the preset cumulative reward function for calculation to obtain the cumulative reward value corresponding to the candidate matching path.
[0069] Specifically, for each candidate matching path, the local incentive value of each matching node in the path is obtained based on cosine similarity. The global incentive value is obtained based on the semantic coherence of each path. After obtaining the local and global incentive values, a weighted sum is performed to obtain the path reward value. Finally, the path reward value is substituted into a preset cumulative reward function to calculate the cumulative reward value corresponding to the candidate matching path.
[0070] Optionally, based on the above embodiments, in some embodiments of the present invention, the preset cumulative reward function may be limited by the following expression:
[0071] in, This represents the discount factor corresponding to the t-th matching connection node. This represents the path reward value corresponding to the t-th matching connection node.
[0072] S13: Update the speech-text matching model according to the cumulative reward value corresponding to the target optimal matching path, and store the relevant matching data of the current multiple candidate training samples to be matched into the experience pool. Return to the execution and construct the semantic feature space according to the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched until the training termination condition is reached, and obtain the trained target speech-text matching model.
[0073] Among them, relevant matching data refers to the data in the process of matching candidate training samples to be matched with matching candidate training samples. The relevant matching data includes: information of semantic feature space, such as: multiple semantic feature candidate spaces, each semantic feature candidate space includes: key anchor points, multiple candidate training samples to be matched and multiple candidate training samples to be matched corresponding to the key anchor points, multiple candidate matching paths corresponding to each candidate training sample to be matched, the target optimal matching path, the cumulative reward value corresponding to the target optimal matching path, etc.
[0074] The target speech-to-text matching model is used to obtain the matching results for the target object. The speech-to-text matching model includes a policy network and a value network.
[0075] Specifically, based on the cumulative reward value corresponding to the obtained target optimal matching path, the speech-text matching model is updated, and the relevant matching data of the current multiple candidate training samples to be matched are stored in the experience pool. Then, the process is returned to execute the construction of the semantic feature space based on the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched, until the training termination condition is reached, and the trained target speech-text matching model is obtained.
[0076] Optionally, based on the above embodiments, in some embodiments of the present invention, updating the speech-text matching model according to the cumulative reward value corresponding to the target optimal matching path can be implemented in the following way: S20: Determine the policy gradient based on the cumulative reward value corresponding to the optimal matching path of the target.
[0077] Among them, the policy gradient is used to characterize the influence of the policy network on the cumulative reward value when adjusting the initial anchor position to obtain the key anchor, as well as the influence of determining the multiple candidate matching paths and the target optimal matching path corresponding to each candidate training sample to be matched on the cumulative reward value.
[0078] S21: Adjust the weight parameters of the policy network according to the policy gradient.
[0079] Specifically, the policy gradient is determined based on the cumulative reward value corresponding to the optimal matching path of the target. After obtaining the policy gradient, the weight parameters of the policy network are adjusted according to the policy gradient.
[0080] S22: Determine the value error based on the cumulative reward value and the preset value error loss function.
[0081] The preset value error loss function can be a difference calculation function.
[0082] S23: Adjust the weight parameters of the value network based on the value error.
[0083] Specifically, the cumulative reward value is substituted into the preset value error loss function for calculation to determine the value error. After obtaining the value error, the weight parameters of the value network are adjusted according to the value error.
[0084] Optionally, based on the above embodiments, in some embodiments of the present invention, the training termination condition may be: setting a preset number of iterations, and when the training number reaches the preset number of iterations, ending the training to obtain the trained target speech-text matching model.
[0085] Optionally, based on the above embodiments, in some embodiments of the present invention, the training termination condition may also be: when the value error is equal to the preset value error, the training ends and the trained target speech-text matching model is obtained.
[0086] Thus, the speech-text matching method based on reinforcement learning provided in this embodiment obtains multiple training samples to be matched and matching training samples corresponding to each training sample to be matched. The training samples to be matched are either speech features to be matched or text features to be matched, and the matching training samples are either speech features to be matched or text features to be matched. There is a one-to-one correspondence between the speech features to be matched and the text features to be matched, and vice versa. Based on the multiple training samples to be matched and the multiple matching training samples, a semantic feature space is constructed. This semantic feature space includes multiple semantic feature candidate spaces, each of which includes a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. Based on the multiple key anchor points, the multiple candidate training samples to be matched corresponding to each key anchor point, the multiple candidate training samples to be matched, and a preset cumulative reward function, a target optimal matching path is determined for each candidate training sample to be matched. The target optimal matching path is used to determine the candidate training samples to be matched corresponding to the candidate training samples to be matched. Based on the cumulative reward value corresponding to the optimal matching path, the speech-text matching model is updated, and the relevant matching data of multiple candidate training samples to be matched are stored in the experience pool. The process then returns to the previous step, constructing a semantic feature space based on the multiple candidate training samples and their corresponding matching training samples, until the training termination condition is met. This yields a trained target speech-text matching model, which is used to obtain the matching results for the target object. In this way, the present invention can improve the ability to obtain semantic alignment information with candidate training samples within the semantic feature space based on key anchor points corresponding to multiple candidate semantic feature spaces. This enhances the diversity and comprehensiveness of matching between candidate training samples and matching candidate training samples, and utilizes the cumulative reward value corresponding to the optimal matching path to optimize the speech-text matching model, thereby improving the accuracy of the model in obtaining the matching results for the target object.
[0087] Optionally, based on the above embodiments, some embodiments of the present invention further include: Obtain the target object to be matched. Input the target object to be matched into the target speech-text matching model, and obtain the matching result of the target object through the policy network in the target speech-text matching model.
[0088] The target object to be matched can be either the speech feature target object or the text feature target object. The matching result is either the matched speech feature target object or the matched text feature target object. There is a one-to-one correspondence between the speech feature target object to be matched and the matched text feature target object, and vice versa.
[0089] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0090] In one embodiment, such as Figure 2 As shown, a speech-text matching device based on reinforcement learning is provided, including: a training sample acquisition module 10, a semantic feature space construction module 11, a target optimal matching path determination module 12, and a target speech-text matching model acquisition module 13.
[0091] The training sample acquisition module 10 is used to acquire multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched. The training sample to be matched is any one of the speech feature to be matched and the text feature to be matched, and the matching training sample is any one of the speech feature to be matched and the text feature to be matched. The speech feature to be matched and the text feature to be matched are in one-to-one correspondence, and the text feature to be matched and the speech feature to be matched are in one-to-one correspondence.
[0092] The semantic feature space construction module 11 is used to construct a semantic feature space based on multiple training samples to be matched and multiple matching training samples. The semantic feature space includes multiple semantic feature candidate spaces, and each semantic feature candidate space includes a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched.
[0093] The target optimal matching path determination module 12 is used to determine the target optimal matching path corresponding to each candidate training sample to be matched based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function. The target optimal matching path is used to determine the candidate training samples to be matched corresponding to the candidate training samples to be matched.
[0094] The target speech-text matching model acquisition module 13 is used to update the speech-text matching model according to the cumulative reward value corresponding to the optimal matching path of the target, and store the relevant matching data of multiple candidate training samples to be matched into the experience pool. It returns to execute the construction of semantic feature space based on multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched until the training termination condition is reached, and obtains the trained target speech-text matching model. The target speech-text matching model is used to obtain the matching result of the target object to be matched.
[0095] In the above embodiments, a training sample acquisition module acquires multiple training samples to be matched and matching training samples corresponding to each training sample to be matched. The training samples to be matched are either speech features to be matched or text features to be matched, and the matching training samples are either speech features to be matched or text features to be matched. There is a one-to-one correspondence between the speech features to be matched and the text features to be matched, and vice versa. A semantic feature space construction module constructs a semantic feature space based on the multiple training samples to be matched and the multiple matching training samples. The semantic feature space includes multiple semantic feature candidate spaces, each including a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. A target optimal matching path determination module determines the target optimal matching path corresponding to each candidate training sample to be matched based on the multiple key anchor points, the multiple candidate training samples to be matched corresponding to each key anchor point, the multiple candidate training samples to be matched, and a preset cumulative reward function. The target optimal matching path is used to determine the candidate training samples to be matched corresponding to the candidate training samples to be matched. The target speech-text matching model acquisition module updates the speech-text matching model based on the cumulative reward value corresponding to the optimal matching path. It also stores the relevant matching data of multiple candidate training samples into an experience pool. The module then returns to the execution phase, constructing a semantic feature space based on the multiple candidate training samples and their corresponding matching training samples until the training termination condition is met. This results in a trained target speech-text matching model, which is used to obtain the matching results for the target object. This approach enhances the ability to acquire semantic alignment information with the candidate training samples within the semantic feature space, based on the key anchor points corresponding to multiple candidate semantic features. This increases the diversity and comprehensiveness of the matching between the candidate training samples and the matching candidate training samples. Furthermore, the cumulative reward value corresponding to the optimal matching path optimizes the speech-text matching model, thereby improving the accuracy of the model in obtaining the matching results for the target object.
[0096] Specific limitations regarding the reinforcement learning-based speech-to-text matching device can be found in the limitations of the reinforcement learning-based speech-to-text matching method above, and will not be repeated here. Each module in the aforementioned server can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independent of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.
[0097] This invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can implement a reinforcement learning-based speech-to-text matching method provided in this invention. For example, when the processor executes the computer program, it can implement... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.
[0098] This invention provides a computer-readable storage medium storing at least one program, which is executed by a processor to implement... Figure 1 The technical solutions of the method embodiments shown are similar in principle and in effect, and will not be described again here.
[0099] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM), etc.
[0100] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0101] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A speech-to-text matching method based on reinforcement learning, characterized in that, include: Obtain multiple training samples to be matched and matching training samples corresponding to each training sample to be matched. The training sample to be matched is either the speech feature to be matched or the text feature to be matched. The matching training sample is either the speech feature to be matched or the text feature to be matched. The speech feature to be matched and the text feature to be matched are in one-to-one correspondence. Based on the multiple training samples to be matched and the multiple training samples to be matched, a semantic feature space is constructed, wherein the semantic feature space includes: multiple semantic feature candidate spaces, and each semantic feature candidate space includes: a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. Based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function, a target optimal matching path is determined for each candidate training sample to be matched, wherein the target optimal matching path is used to determine the candidate training sample to be matched. Based on the cumulative reward value corresponding to the optimal matching path of the target, the speech-text matching model is updated, and the relevant matching data of the current multiple candidate training samples to be matched are stored in the experience pool. Then, the execution is returned to construct the semantic feature space based on the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched, until the training termination condition is reached, and the trained target speech-text matching model is obtained. The target speech-text matching model is used to obtain the matching result of the target object to be matched.
2. The method according to claim 1, characterized in that, The step of constructing a semantic feature space based on the plurality of training samples to be matched and the plurality of matching training samples includes: Based on the multiple training samples to be matched and the multiple matching training samples, an initial semantic feature space is constructed; Based on the distribution density of the multiple training samples to be matched and the multiple matching training samples, the initial semantic feature space is divided to obtain an intermediate semantic feature space. The intermediate semantic feature space includes multiple intermediate semantic feature subspaces, and each intermediate semantic feature subspace includes multiple intermediate training samples to be matched and multiple intermediate matching training samples. Based on the semantic center vector of each intermediate semantic feature subspace, the initial anchor position of each intermediate semantic feature subspace is initialized, wherein the semantic center vector is determined based on multiple intermediate training samples to be matched and multiple intermediate matching training samples. Based on the initial anchor point position, the preset radius, and the matching success rate of the training samples to be matched and the matching training samples in the intermediate semantic feature subspace, the initial anchor point position is adjusted to obtain the key anchor point, and the key anchor point is verified. Based on the verification results, multiple semantic feature candidate spaces are determined in multiple intermediate semantic feature subspaces. Each semantic feature candidate space includes: a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. The semantic feature space is constructed based on multiple semantic feature candidate spaces.
3. The method according to claim 1, characterized in that, The step of determining the target optimal matching path corresponding to each of the multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function includes: Based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, and multiple candidate training samples to be matched, multiple candidate matching paths are determined for each candidate training sample to be matched, wherein the candidate matching paths are used to determine the candidate matching training samples corresponding to the candidate training samples to be matched. The cumulative reward value corresponding to multiple candidate matching paths is obtained according to the preset cumulative reward function, and the candidate matching path corresponding to the largest cumulative reward value is determined as the target optimal matching path based on the multiple cumulative reward values.
4. The method according to claim 3, characterized in that, The step of determining multiple candidate matching paths corresponding to each candidate training sample to be matched based on multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, and multiple candidate training samples to be matched includes: Starting from each key anchor point, the matching is performed on multiple candidate training samples to be matched and multiple candidate training samples to be matched corresponding to each key anchor point to obtain multiple initial candidate matching paths. Each initial candidate matching path includes: a first initial candidate matching sub-path corresponding to the candidate training sample to be matched and a second initial candidate matching sub-path corresponding to the candidate training sample to be matched. For the first initial candidate matching sub-path and the second initial candidate matching sub-path, the path is expanded in multiple candidate training samples to be matched and multiple candidate training samples corresponding to all key anchor points to obtain the first candidate matching sub-path corresponding to the first initial candidate matching sub-path and the second candidate matching sub-path corresponding to the second initial candidate matching sub-path. The first candidate matching sub-path and the second candidate matching sub-path are combined to obtain a candidate matching path.
5. The method according to claim 4, characterized in that, The candidate matching path includes: multiple matching connection nodes, wherein the matching connection node is a candidate training sample to be matched, or a candidate training sample that matches the candidate training sample to be matched, and the step of obtaining the cumulative reward value corresponding to the multiple candidate matching paths according to the preset cumulative reward function includes: For each candidate matching path, the local activation value of each matching connection node is obtained based on the cosine similarity. Obtain the global incentive value for each candidate matching path; The path reward value is obtained by weighted summation of the local incentive value and the global incentive value; The path reward value is substituted into a preset cumulative reward function for calculation to obtain the cumulative reward value corresponding to the candidate matching path.
6. The method according to claim 5, characterized in that, The speech-text matching model includes a policy network and a value network. Updating the speech-text matching model based on the cumulative reward value corresponding to the target optimal matching path includes: The strategy gradient is determined based on the cumulative reward value corresponding to the optimal matching path of the target. Adjust the weight parameters of the policy network according to the policy gradient; The value error is determined based on the cumulative reward value and the preset value error loss function; The weight parameters of the value network are adjusted based on the value error.
7. The method according to claim 6, characterized in that, The method further includes: Obtain the target object to be matched, wherein the target object to be matched is either the speech feature target object to be matched or the text feature target object to be matched; The target object to be matched is input into the target speech-text matching model, and the matching result of the target object is obtained through the policy network in the target speech-text matching model.
8. A speech-text matching device based on reinforcement learning, characterized in that, include: The training sample acquisition module is used to acquire multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched. The training sample to be matched is any one of the speech feature to be matched and the text feature to be matched. The matching training sample is any one of the speech feature to be matched and the text feature to be matched. The speech feature to be matched and the text feature to be matched correspond one-to-one. The semantic feature space construction module is used to construct a semantic feature space based on the multiple training samples to be matched and the multiple matching training samples. The semantic feature space includes multiple semantic feature candidate spaces, and each semantic feature candidate space includes a key anchor point, multiple candidate training samples to be matched corresponding to the key anchor point, and multiple candidate training samples to be matched. The target optimal matching path determination module is used to determine the target optimal matching path corresponding to each of the multiple key anchor points, multiple candidate training samples to be matched corresponding to the multiple key anchor points, multiple candidate training samples to be matched, and a preset cumulative reward function, wherein the target optimal matching path is used to determine the candidate training samples to be matched corresponding to the candidate training samples to be matched. The target speech-text matching model acquisition module is used to update the speech-text matching model according to the cumulative reward value corresponding to the optimal matching path of the target, and store the relevant matching data of multiple candidate training samples to be matched into the experience pool. Then, it returns to execute the construction of semantic feature space according to the multiple training samples to be matched and the matching training samples corresponding to each training sample to be matched, until the training termination condition is reached, and obtains the trained target speech-text matching model. The target speech-text matching model is used to obtain the matching result of the target object to be matched.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the speech-text matching method based on reinforcement learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the speech-text matching method based on reinforcement learning as described in any one of claims 1 to 7.