Video recognition method, model training method, device, equipment and storage medium
By combining speech and text recognition results, using similar words to adjust the recognition results and coding processing, the problem of low accuracy in video subtitles recognition is solved, and higher recognition accuracy and recognition ability of nouns in professional fields is achieved.
Patent Information
- Application Number
- CN202410568795.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-05-09
AI Technical Summary
The prior art has the problem of low accuracy in subtitle recognition in video recognition, especially in professional videos, where subtitle errors are frequent and it is difficult to obtain correct subtitles.
By combining speech recognition and text recognition results, similar words are used to adjust the recognition results, encoding processing is performed to improve the recognition accuracy, and error correction is performed using a large language model.
It improves the recognition accuracy of video subtitles, reduces the recognition error caused by background interference and unfixed subtitles position, and enhances the recognition ability of nouns in professional fields.
Smart Images

Figure CN118314900B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technologies such as deep learning, video recognition, and large models, and can be applied to scenarios such as large language model (LLM) corpora, intelligent question and answer knowledge bases, information retrieval knowledge bases, and knowledge graph corpora. More specifically, the present disclosure provides a video recognition method, a training method for a large language model, a training method for a video recognition model, an apparatus, an electronic device, and a storage medium. Background Art
[0002] With the development of artificial intelligence technologies, the application fields of related technologies such as large language models, intelligent question and answer, knowledge bases, and knowledge graphs are constantly expanding. Summary of the Invention
[0003] The present disclosure provides a video recognition method, a training method for a large language model, a training method for a video recognition model, an apparatus, a device, and a storage medium.
[0004] According to one aspect of the present disclosure, there is provided a video recognition method, the method comprising: determining an adjusted speech recognition result according to at least one first word segmentation of a speech recognition result of a video to be processed and at least one first similar word corresponding to the at least one first word segmentation; determining an adjusted text recognition result according to at least one second word segmentation of a text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation; performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature; performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and decoding the feature to be decoded to obtain a target recognition result.
[0005] According to another aspect of the present disclosure, there is provided a method for training a large language model, the method comprising: training the large language model using a plurality of target recognition results, wherein the target recognition results are obtained by the following operations: determining an adjusted speech recognition result according to at least one first word segment of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segment; determining an adjusted text recognition result according to at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment; performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature; performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature and the second initial speech encoding feature; and decoding the feature to be decoded to obtain the target recognition result.
[0006] According to another aspect of the present disclosure, there is provided a method for training a video recognition model, the method comprising: determining an adjusted speech recognition result according to at least one first word segment of the speech recognition result of the sample video to be processed and at least one first similar word corresponding to the at least one first word segment; determining an adjusted text recognition result according to at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment; inputting the speech recognition result and the adjusted speech recognition result into a first encoding sub-model of the video recognition model to obtain a first initial speech encoding feature and a second initial speech encoding feature; inputting the text recognition result and the adjusted text recognition result into a second encoding sub-model of the video recognition model to obtain a first initial text encoding feature and a second initial text encoding feature; obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature and the second initial speech encoding feature; decoding the feature to be decoded to obtain a sample recognition result; and training the video recognition model according to the sample recognition result and the label of the sample video to be processed.
[0007] According to another aspect of the present disclosure, there is provided a video recognition device, which includes: a first determination module, configured to determine an adjusted speech recognition result according to at least one first word segmentation of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segmentation; a second determination module, configured to determine an adjusted text recognition result according to at least one second word segmentation of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation; a first encoding module, configured to perform first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature; a second encoding module, configured to perform second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; a first obtaining module, configured to obtain a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and a first decoding module, configured to decode the feature to be decoded to obtain a target recognition result.
[0008] According to another aspect of the present disclosure, there is provided a training device for a large language model, which includes: a first training module, configured to train the large language model by using a plurality of target recognition results, where the target recognition results are obtained through the following operations: determining an adjusted speech recognition result according to at least one first word segmentation of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segmentation; determining an adjusted text recognition result according to at least one second word segmentation of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation; performing first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature; performing second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and decoding the feature to be decoded to obtain a target recognition result.
[0009] According to another aspect of the present disclosure, there is provided a training device for a video recognition model, the device including: a third determination module configured to determine an adjusted speech recognition result according to at least one first word segmentation of the speech recognition result of a sample video to be processed and at least one first similar word corresponding to the at least one first word segmentation; a fourth determination module configured to determine an adjusted text recognition result according to at least one second word segmentation of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation; a third encoding module configured to input the speech recognition result and the adjusted speech recognition result into a first encoding sub-model of the video recognition model to obtain a first initial speech encoding feature and a second initial speech encoding feature; a fourth encoding module configured to input the text recognition result and the adjusted text recognition result into a second encoding sub-model of the video recognition model to obtain a first initial text encoding feature and a second initial text encoding feature; a second obtaining module configured to obtain a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; a second decoding module configured to decode the feature to be decoded to obtain a sample recognition result; and a second training module configured to train the video recognition model according to the sample recognition result and the label of the sample video to be processed.
[0010] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided by the present disclosure.
[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided by the present disclosure.
[0012] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method provided by the present disclosure.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0014] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0015] Figure 1It is a schematic diagram of an exemplary system architecture to which a video recognition method and apparatus can be applied according to an embodiment of the present disclosure;
[0016] Figure 2 It is a flowchart of a video recognition method according to an embodiment of the present disclosure;
[0017] Figure 3A It is a schematic diagram of a video frame to be recognized according to an embodiment of the present disclosure;
[0018] Figure 3B It is a schematic diagram of another video frame to be recognized according to an embodiment of the present disclosure;
[0019] Figure 4 It is a schematic diagram of a video recognition method according to an embodiment of the present disclosure;
[0020] Figure 5 It is a schematic flowchart of a training method for a large language model according to an embodiment of the present disclosure;
[0021] Figure 6 It is a schematic diagram of applying a target recognition result according to an embodiment of the present disclosure;
[0022] Figure 7 It is a schematic flowchart of a training method for a video recognition model according to an embodiment of the present disclosure;
[0023] Figure 8 It is a schematic diagram of a video recognition model according to an embodiment of the present disclosure;
[0024] Figure 9 It is a block diagram of a video recognition apparatus according to an embodiment of the present disclosure;
[0025] Figure 10 It is a block diagram of a training apparatus for a large language model according to another embodiment of the present disclosure;
[0026] Figure 11 It is a block diagram of a training apparatus for a video recognition model according to another embodiment of the present disclosure; and
[0027] Figure 12 It is a block diagram of an electronic device to which a video recognition method, a training method for a large language model, and / or a training method for a video recognition model can be applied according to an embodiment of the present disclosure. Detailed implementation manners
[0028] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0029] Videos can contain rich information. The subtitles of videos have high application value. In professional fields such as medicine, law, and finance, the subtitles of relevant videos can contain rich and professional knowledge information, which can help users understand the knowledge in the professional field or solve corresponding professional problems. For example, based on videos in the medical field, users can select the hospital departments to register for. Based on videos in the legal field, users can prepare legal lawsuits by themselves.
[0030] With the rapid development of conversational large language models, corresponding large models have emerged in various professional fields. The training basis of large language models is corpus. In the early training stage, large language models have already used a large amount of existing corpus. It is difficult to obtain new corpus. The subtitles of videos, especially those of videos in professional fields, can be an effective supplement to the training corpus of large language models.
[0031] However, the video can have a complex background. Not only subtitles but also background text can appear in video frames, resulting in a decline in recognition effect. In addition, the appearance position of the subtitles in the video is not fixed, making it difficult to eliminate the interference of background text. In some embodiments, in order to improve the accuracy of subtitle recognition, the text recognition result and speech recognition result of the video can be fused, and the fused result can be input into a large language model for text correction. However, there is an upper limit on the number of input characters of the large language model, and the number of characters in the fused result is easily exceeded, resulting in a decline in the correction effect. In addition, for nouns in professional fields such as medicine and law, they rarely appear in the corpus used for large language model training. During subtitle recognition or correction, synonyms or words with similar shapes are likely to appear in the recognition result, resulting in a low recognition accuracy of professional field nouns.
[0032] In addition, different from common movies and TV dramas, the subtitles of videos in professional fields are very likely to be incorrect themselves, and it is also difficult to obtain correct subtitles from relevant websites.
[0033] Therefore, in order to efficiently obtain subtitles from videos, the present disclosure provides a video recognition method, which will be described below.
[0034] Figure 1 is a schematic diagram of an exemplary system architecture to which the video recognition method and apparatus according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0035] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0036] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0037] The server 105 may be a server that provides various services, such as a background management server (only for example) that supports the websites browsed by users using the terminal devices 101, 102, 103. The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the terminal devices.
[0038] It should be noted that the video recognition method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the video recognition device provided by the embodiments of the present disclosure can generally be set in the server 105. The video recognition method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the video recognition device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0039] It can be understood that the system architecture for applying the method of the present disclosure has been described above, and the method of the present disclosure will be described below.
[0040] Figure 2 is a flowchart of a video recognition method according to an embodiment of the present disclosure.
[0041] As Figure 2 shown, the method 200 may include operation S210 to operation S260.
[0042] In operation S210, based on at least one first word segment of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segment, an adjusted speech recognition result is determined.
[0043] In an embodiment of the present disclosure, the video to be processed may include audio. By performing automatic speech recognition (ASR) on the audio, the speech recognition result of the video to be processed and the time period in which the speech recognition result is located can be obtained.
[0044] In an embodiment of the present disclosure, the speech recognition result may include at least one first word segment.
[0045] In an embodiment of the present disclosure, the first word segment may correspond to one or more first similar words. The pronunciation of the first similar word may be similar to that of the corresponding first word segment. For example, the initial consonant and final sound of the first similar word may be the same as those of the first word segment. In this case, the tone of the first similar word may be the same as or different from that of the corresponding first word segment. Also, for example, the initial consonant and tone of the first similar word may be the same as those of the corresponding first word segment. In this case, the tone of the first similar word may be the same as that of the corresponding first word segment.
[0046] In an embodiment of the present disclosure, the first word segment may be replaced with the first similar word to obtain an adjusted speech recognition result. For example, when there are multiple first word segments or multiple first similar words, multiple speech replacement results may be obtained. The speech replacement result with the highest semantic similarity to the speech recognition result may be determined from the multiple speech replacement results as the adjusted speech recognition result. It can be understood that other methods may also be used to determine the adjusted speech recognition result from the multiple speech replacement results.
[0047] In operation S220, based on at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment, an adjusted text recognition result is determined.
[0048] In an embodiment of the present disclosure, optical character recognition (OCR) may be performed on multiple video frames of the video to be processed to obtain multiple text recognition results and the time periods in which the multiple text recognition results are located.
[0049] In an embodiment of the present disclosure, the text recognition result in the same time period as the speech recognition result may be used as the text recognition result corresponding to the speech recognition result.
[0050] In an embodiment of the present disclosure, the text recognition result may include at least one second word segment.
[0051] In an embodiment of the present disclosure, the second participle may correspond to one or more second similar words. The structure of the second similar words may be similar to that of the corresponding second participle.
[0052] In an embodiment of the present disclosure, the second participle may be replaced with the second similar words to obtain an adjusted text recognition result. For example, in the case where there are multiple second participles or multiple second similar words, multiple text replacement results may be obtained. The text replacement result with the highest semantic similarity to the text recognition result may be determined from the multiple text replacement results as the adjusted text recognition result. It can be understood that other methods may also be used to determine the adjusted text recognition result from the multiple text replacement results.
[0053] In operation S230, the speech recognition result and the adjusted speech recognition result are subjected to a first encoding to obtain a first initial speech encoding feature and a second initial speech encoding feature.
[0054] In an embodiment of the present disclosure, the speech recognition result is subjected to a first encoding to obtain a first initial speech encoding feature. The adjusted speech recognition result is subjected to a first encoding to obtain a second initial speech encoding feature.
[0055] In operation S240, the text recognition result and the adjusted text recognition result are subjected to a second encoding to obtain a first initial text encoding feature and a second initial text encoding feature.
[0056] In an embodiment of the present disclosure, the text recognition result is subjected to a first encoding to obtain a first initial text encoding feature. The adjusted text recognition result is subjected to a first encoding to obtain a second initial text encoding feature.
[0057] In operation S250, a to-be-decoded feature is obtained according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature.
[0058] In an embodiment of the present disclosure, the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature are fused to obtain the to-be-decoded feature. The fusion method may include addition, splicing, etc.
[0059] In operation S260, the to-be-decoded feature is decoded to obtain a target recognition result.
[0060] In an embodiment of the present disclosure, a fully connected process is performed on the to-be-decoded feature to obtain the target recognition result. For example, the target recognition result may be the caption of a video frame of the video to be processed.
[0061] In the embodiments of the present disclosure, the speech recognition result is adjusted by using a first similar word that is phonetically similar to the first word segmentation, which helps to reduce the error of speech recognition. The text recognition result is adjusted by using a second similar word that is structurally similar to the second word segmentation, which helps to reduce the error of text recognition. The same encoding method (the first encoding) is used for the original speech recognition result and the adjusted speech recognition result, which can further eliminate the speech recognition error caused by phonetic similarity. The same encoding method (the second encoding) is used for the original text recognition result and the adjusted text recognition result, which can further reduce the text recognition error caused by character structure similarity. The to-be-decoded feature is obtained based on the two features obtained from the first encoding and the two features obtained from the second encoding, which can make the to-be-decoded feature contain richer information, and then a more accurate target recognition result can be obtained after decoding.
[0062] It can be understood that for operations S210 to S240, when operation S210 is executed first and then operation S230, and operation S220 is executed first and then operation S240, operations S210, S220, S230, and S240 can be executed in a variety of different orders. For example, operations S210, S230, S220, and S240 are executed in sequence. For example, operations S220, S210, S230, and S240 are executed in sequence. For example, operations S220, S240, S210, and S230 are executed in sequence. Another example is that operations S210 and S220 are executed in parallel, operation S230 is executed after operation S210 is completed, and operation S240 is executed after operation S220 is completed.
[0063] It can be understood that the method of the present disclosure has been described above, and the speech recognition method of the present disclosure will be described below.
[0064] In some embodiments, the above method 200 may further include: performing speech recognition on the video to be processed to obtain a first initial recognition result. Performing sentence segmentation on the first initial recognition result to obtain at least one speech recognition result of the video to be processed and at least one time period corresponding to the at least one speech recognition result. For example, the speech recognition result may be a sentence. Another example is that when the video to be processed is about something in the medical field, the speech recognition result may be "Patients with viral cold can take Lianhua Qingwen Capsules and Ganmaoling Granules". It can be understood that this speech recognition result is only an example.
[0065] It can be understood that the speech recognition method of the present disclosure has been described above, and the text recognition method of the present disclosure will be described below.
[0066] In some embodiments, the above method 200 may further include: performing text recognition on the video to be processed to obtain a plurality of initial text boxes. According to the plurality of initial text boxes, a text recognition result is obtained. For example, the video to be processed may be frame-extracted to obtain a plurality of initial video frames. Duplicate removal may be performed on the plurality of initial video frames to obtain a plurality of video frames to be recognized. Text recognition may be performed on the plurality of video frames to be recognized to obtain a plurality of initial text boxes. The following will be described in conjunction with Figure 3A and Figure 3B for illustration.
[0067] Figure 3A FIG. is a schematic diagram of a video frame to be recognized according to an embodiment of the present disclosure.
[0068] As Figure 3A shown, recognizing the video frame to be recognized F31 may obtain a plurality of initial text boxes. The plurality of initial text boxes may include an initial text box b301, an initial text box b311, an initial text box b312, an initial text box b321, and an initial text box b322.
[0069] In the embodiments of the present disclosure, the initial text box corresponds to the text in the video to be processed. For example, the initial text box b301 may correspond to the text "disease". The initial text box b311 may correspond to the text "welcome". The initial text box b312 may correspond to the text "welcome". The initial text box b321 may correspond to the text "medical". The initial text box b322 may correspond to the text "museum".
[0070] In the embodiments of the present disclosure, the initial text box is used to indicate the position of the text in the video to be processed. As Figure 3A shown, the initial text box b301 may indicate the position of the text "disease".
[0071] Figure 3B FIG. is a schematic diagram of another video frame to be recognized according to an embodiment of the present disclosure.
[0072] As Figure 3B shown, recognizing the video frame to be recognized F32 may obtain a plurality of initial text boxes. The plurality of initial text boxes may include an initial text box b302, an initial text box b323, and an initial text box b324. The initial text box b302 may correspond to the text "can". The initial text box b323 may correspond to the text "medical". The initial text box b324 may correspond to the text "museum".
[0073] In the embodiments of the present disclosure, obtaining a text recognition result based on multiple initial text boxes may include: determining multiple position frequencies according to the positions of the multiple initial text boxes respectively. For example, assume that the video to be processed includes a video frame F31 to be recognized and a video frame F32 to be recognized. The initial text box b301 and the initial text box b302 indicate the same position, and the position frequency of this position may be 2. The initial text box b321 and the initial text box b323 indicate the same position, and the position frequency of this position may be 2. The initial text box b322 and the initial text box b324 indicate the same position, and the position frequency of this position may be 2. For the position indicated by the initial text box b311, the position frequency of this position may be 1. For the position indicated by the initial text box b312, the position frequency of this position may be 1. It can be understood that the video to be processed may include a greater number of video frames to be recognized, and correspondingly, the value of the position frequency may be greater than 2.
[0074] In the embodiments of the present disclosure, obtaining a text recognition result based on multiple initial text boxes may further include: determining the text recognition result according to at least one initial text box corresponding to at least one target position frequency. The target position frequency is the position frequency among the multiple position frequencies that is greater than or equal to a preset frequency threshold. For example, assume that the video to be processed includes a video frame F31 to be recognized and a video frame F32 to be recognized. The preset frequency threshold may be 2, for example. In this case, the initial text boxes with position frequencies less than the preset frequency threshold may be deleted. That is, the initial text box b311 and the initial text box b312 are deleted. Next, the second initial recognition result may be determined according to the initial text box b301, the initial text box b302, and the initial text boxes b321 to b322. Through the embodiments of the present disclosure, some text boxes are deleted according to the position frequency, and the text in the moving background can be removed. For example, the non-caption text that occasionally appears on the display screen or the window of a moving vehicle can be removed, which helps to accurately locate the caption position and improve the accuracy of text recognition.
[0075] It can be understood that the above filters multiple initial text boxes based on the frequent pattern. However, the present disclosure is not limited thereto, and the following will be described.
[0076] In the embodiments of the present disclosure, obtaining a text recognition result based on multiple initial text boxes may further include: determining at least one initial text box group according to the multiple initial text boxes. At least one initial text box within the initial text box group indicates the same position. For example, the initial text box b301 and the initial text box b302 indicate the same position. An initial text box group can be determined according to the initial text box b301 and the initial text box b302. The initial text box b321 and the initial text box b323 indicate the same position. An initial text box group can be determined according to the initial text box b321 and the initial text box b323.
[0077] In an embodiment of the present disclosure, obtaining a text recognition result based on multiple initial text boxes may further include: determining the information entropy of an initial text box group according to at least one character corresponding to the initial text box group. For example, the information entropy of the initial text box group may be determined according to at least one target probability of at least one character. For the initial text box group determined according to the initial text box b301 and the initial text box b302, the initial text box group may correspond to the characters "disease" and "possible". For this initial text box group, a total of 2 characters appear. The target probability of the character "disease" may be 0.5, and the appearance probability of the character "possible" may be 0.5. For the initial text box group determined according to the initial text box b321 and the initial text box b323, the initial text box group may correspond to the character "medical". For this initial text box group, a total of 1 character appears. The target probability of the character "medical" may be 1. For another example, the information entropy H(X) of the initial text box group may be determined by the following formula:
[0078] H(X) = -∑p(x)log p(x) (Formula 1)
[0079] p(x) may be the target probability. The base of the log function may be 2.
[0080] The information entropy of the initial text box group determined according to the initial text box b301 and the initial text box b302 may be 1. The information entropy of the initial text box group determined according to the initial text box b321 and the initial text box b323 may be 0. The information entropy of the initial text box group determined according to the initial text box b322 and the initial text box b324 may be 0.
[0081] In an embodiment of the present disclosure, obtaining a text recognition result based on multiple initial text boxes may further include: determining a character recognition result according to at least one initial text box group corresponding to at least one target information entropy. The target information entropy may be an information entropy greater than or equal to a preset entropy threshold. For example, the preset entropy threshold may be 0.8. The information entropy of the initial text box group determined according to the initial text box b301 and the initial text box b302 may be used as the target information entropy. The initial text boxes b321, b322, b323, and b324 may be deleted. According to the initial text box b301 and the initial text box b302, a second initial recognition result is determined. Through the embodiment of the present disclosure, according to the information entropy, the characters in the fixed background can be removed, the interference of the characters in the background on the subtitle recognition is reduced, the position where the subtitle is located can be further accurately located, and the accuracy of text recognition is further improved.
[0082] It can be understood that the above has respectively described some ways of determining the text recognition result in combination with information entropy and frequent patterns. However, the present disclosure is not limited thereto, and the text recognition result can be determined based on the combination of information entropy and frequent patterns, which will be described below.
[0083] In some embodiments, some initial text boxes can be deleted first based on the frequent pattern, and then some initial text boxes can be deleted based on the information entropy. According to the initial text boxes that are not deleted, a second initial recognition result can be determined. For example, based on the frequent pattern, the initial text box b311 and the initial text box b312 can be deleted. Next, based on the information entropy, the initial text box b321, the initial text box b322, the initial text box b323, and the initial text box b324 can be deleted. According to the initial text boxes b301 and b302 that are not deleted, a second initial recognition result can be determined.
[0084] In some other embodiments, some initial text boxes can be deleted first based on the information entropy, and then some initial text boxes can be deleted based on the frequent pattern. According to the initial text boxes that are not deleted, a second initial recognition result can be determined.
[0085] It can be understood that the above has described the way of deleting some initial text boxes, and below will describe some ways of determining the text recognition result.
[0086] In some embodiments, the second initial recognition result can be clause-separated to obtain at least one text recognition result and at least one time period corresponding to the at least one text recognition result. For example, the text recognition result can be a sentence. For another example, when the video to be processed is about a medical matter, one text recognition result can be "Patients with viral cold can take Lianhua Qingwen Capsules and Hanmaoling Granules", and another text recognition result can be "It can relieve various symptoms such as headache, fever, and nasal congestion caused by colds". It can be understood that this text recognition result is only an example.
[0087] It can be understood that the above has described the text recognition method of the present disclosure, and below will describe the alignment method between the speech recognition result and the text recognition result.
[0088] In some embodiments, there can be multiple text recognition results.
[0089] In some embodiments, the above method 200 may further include: determining a plurality of text recognition results to be processed from a plurality of text recognition results according to the time period corresponding to the speech recognition result. For example, the number of text recognition results to be processed may be less than or equal to the number of text recognition results. There may be a delay in the subtitles or the speech. In the same time period, there may be one speech recognition result and two text recognition results. These two text recognition results may be used as the text recognition results to be processed.
[0090] In some embodiments, the above method may further include: determining, according to a plurality of similarities between the speech recognition result and the plurality of text recognition results to be processed, the text recognition result corresponding to the speech recognition result from the plurality of text recognition results to be processed.
[0091] In the embodiments of the present disclosure, a plurality of edit distances between the speech recognition result and the plurality of text recognition results to be processed may be determined. The edit distance may be: the minimum number of operations such as replacing, deleting, and adding characters for the first string to become the second string. That is, the edit distance between the speech recognition result and the text recognition result to be processed may be: the minimum number of operations such as replacing, deleting, and adding characters for the text recognition result to be processed to become the speech recognition result.
[0092] In the embodiments of the present disclosure, the edit distance is processed according to the number of characters in the speech recognition result to obtain a processing result. For example, the processing result may be obtained by dividing the edit distance by the number of characters in the speech recognition result. For another example, the processing result Levenshtein_ratio may be obtained by using the following formula:
[0093] Levenshtein_ratio = Levenshtein_distance / ASR_length (Formula 2)
[0094] Levenshtein_distance may be the edit distance. ASR_length may be the number of characters in the speech recognition result.
[0095] In the embodiments of the present disclosure, according to the processing result, the similarity between the speech recognition result and the text recognition result to be processed may be determined. For example, the reciprocal of the processing result may be used as the similarity. The larger the processing result, the smaller the similarity. The smaller the processing result, the larger the similarity. For another example, the text recognition result corresponding to the speech recognition result "Patients with viral cold can take Lianhua Qingwen Capsules and Ganmaoling Granules" may be "Patients with viral cold can take Lianhua Qingwen Capsules and Hanmaoling Granules". Through the embodiments of the present disclosure, the text recognition result corresponding to the speech recognition result is determined, and the information between the corresponding recognition results can be fully utilized, thereby improving the accuracy of video recognition.
[0096] It can be understood that the alignment method between the speech recognition result and the text recognition result has been described above. Next, some methods for determining the adjusted speech recognition result will be described.
[0097] In some embodiments, at least one first word segment corresponds to M first similar words, where M is an integer greater than or equal to 1. For example, the first similar word can be a homophone corresponding to the first word segment. Suppose the first word segment is "huanbing shaxing", and the corresponding first similar word can be "huanciproxacin". Suppose the first word segment is "lianhu qingwen jiaonang", and the corresponding first similar word can be "lian hua qingwen jiaonang".
[0098] In some embodiments, the above method 200 may further include: segmenting the speech recognition result to obtain N first word segments. N can be an integer greater than or equal to 1. For example, segmenting the speech recognition result "Patients with viral cold can take lianhu qingwen jiaonang and ganmaoling granules" can obtain N first word segments including: "viral cold", "patients", "lianhu qingwen jiaonang", and "ganmaoling granules".
[0099] In some embodiments, the above method may further include: determining, according to the N first word segments, M first similar words corresponding to at least one of the first word segments from a target word library.
[0100] In the embodiments of the present disclosure, the target word library may be a professional word library constructed for a professional field. For example, taking the professional field as the medical field, the target word library may include vocabulary such as disease names and drug names.
[0101] In the embodiments of the present disclosure, based on pronunciation matching, the first similar word corresponding to the first word segment can be retrieved from the target word library. For example, one first word segment can correspond to one or more first similar words. Also, for example, there may be no first similar word corresponding to a first word segment in the target word library. Also, for example, taking the first word segment "lianhu qingwen jiaonang" as an example, "lian hua qingwen jiaonang" can be retrieved from the target word library as the corresponding first similar word. It can be understood that a pronunciation similarity threshold can be set to avoid using a word with a completely different pronunciation from the first word segment as the first similar word.
[0102] In some embodiments, in some implementation manners of the above operation S210, determining the adjusted speech recognition result according to at least one first word segment of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segment may include: using the M first similar words corresponding to the at least one first word segment to replace the at least one first word segment to obtain K replaced speech recognition results.
[0103] In the embodiments of the present disclosure, K may be an integer greater than or equal to 1. The first participle can be replaced with the first similar word in various ways. For example, one first participle can be replaced each time, or multiple first participles can be replaced each time. For another example, by using the first similar word "Lianhua Qingwen Capsule" to replace the first participle "Lotus Qingwen Capsule", the voice recognition result after replacement can be obtained as "Patients with viral colds can take Lianhua Qingwen Capsule and Ganmaoling Granules".
[0104] In some embodiments, determining the adjusted voice recognition result according to at least one first participle of the voice recognition result of the video to be processed and at least one first similar word corresponding to the at least one first participle may further include: determining the adjusted voice recognition result according to the perplexity of K voice recognition results after replacement. For example, the perplexity of the voice recognition result after replacement can be calculated. If there are multiple voice recognition results after replacement, the voice recognition result after replacement with the minimum perplexity can be used as the adjusted voice recognition result. For another example, the above voice recognition result after replacement "Patients with viral colds can take Lianhua Qingwen Capsule and Ganmaoling Granules" can be used as the adjusted voice recognition result. Through the embodiments of the present disclosure, the adjusted voice recognition result is determined based on the first similar word, which can reduce the voice recognition error caused by similar pronunciations, further improve the accuracy of voice recognition, and then improve the accuracy of video recognition.
[0105] It can be understood that some methods for determining the adjusted voice recognition result are described above, and below some methods for determining the adjusted text recognition result will be described.
[0106] In some embodiments, at least one second participle corresponds to J second similar words, and J may be an integer greater than or equal to 1. For example, the second similar word may be a shape-similar word corresponding to the second participle. Suppose the second participle is "Hanmaoling", and the corresponding second similar word may be "Ganmaoling". Suppose the second participle is "Fupaisuan", and the corresponding second similar word may be "Fupai suan".
[0107] In some embodiments, the above method 200 may further include: segmenting the text recognition result to obtain I second participles. I may be an integer greater than or equal to 1. For example, segmenting the text recognition result "Patients with viral colds can take Lianhua Qingwen Capsule and Hanmaoling Granules" may obtain I second participles including: "viral cold", "patient", "Lianhua Qingwen Capsule", and "Hanmaoling Granule".
[0108] In some embodiments, the above method 200 may further include: determining J second similar words corresponding to at least one second participle from the target thesaurus according to the I second participles.
[0109] In the embodiments of the present disclosure, the target vocabulary can be a professional vocabulary constructed for a specific professional field. For example, taking the professional field as the medical field, the target vocabulary can include words such as disease names and drug names.
[0110] In the embodiments of the present disclosure, based on the edit distance, the second similar words corresponding to the second word segmentation can be retrieved from the target vocabulary. For example, one second word segmentation can correspond to one or more second similar words. Also, for example, there may be no second similar words corresponding to a second word segmentation in the target vocabulary. Also, for example, taking the second word segmentation as "Hanmao Ling Granules", "Ganmao Ling Granules" can be retrieved from the target vocabulary as the corresponding second similar word. It can be understood that an edit distance threshold can be set to avoid taking the vocabulary that is exactly the same as the first word segmentation as the first similar word, and also to avoid taking any vocabulary as the first similar word of the first word segmentation.
[0111] In some embodiments, in some implementation manners of the above operation S220, determining the adjusted speech recognition result according to at least one second word segmentation of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation may include: using J second similar words corresponding to the at least one second word segmentation to replace the at least one second word segmentation to obtain L text recognition results after replacement.
[0112] In the embodiments of the present disclosure, L can be an integer greater than or equal to 1. The second similar words can be used to replace the second word segmentation based on various methods. For example, one second word segmentation can be replaced each time, or multiple second word segmentations can be replaced each time. Also, for example, using the second similar word "Ganmao Ling Granules" to replace the second word segmentation "Hanmao Ling Granules", the text recognition result after replacement "Patients with viral cold can take Lianhua Qingwen Capsules and Ganmao Ling Granules" can be obtained.
[0113] In some embodiments, in some implementation manners of the above operation S220, determining the adjusted text recognition result according to at least one second word segmentation of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation may further include: determining the adjusted text recognition result according to the perplexity of the L text recognition results after replacement. For example, the perplexity of the text recognition result after replacement can be calculated. If there are multiple text recognition results after replacement, the text recognition result after replacement with the minimum perplexity can be used as the adjusted text recognition result. Also, for example, the above text recognition result after replacement "Patients with viral cold can take Lianhua Qingwen Capsules and Ganmao Ling Granules" can be used as the adjusted text recognition result. Through the embodiments of the present disclosure, the adjusted speech recognition result is determined based on the first similar word, which can reduce the speech recognition error caused by similar pronunciations, further improve the accuracy of speech recognition, and thus improve the accuracy of video recognition.
[0114] It can be understood that some ways of determining the adjusted recognition result are described above. Next, some ways of obtaining the encoded features will be described in conjunction with Figure 4 some ways of obtaining the encoded features.
[0115] Figure 4 is a schematic diagram of a video recognition method according to an embodiment of the present disclosure.
[0116] In some embodiments, in some implementations of the above operation S230, performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain the first initial speech encoded feature and the second initial speech encoded feature may include: performing an embedding process on the speech recognition result to obtain the first initial speech embedding feature. Performing an embedding process on the adjusted speech recognition result to obtain the second initial speech embedding feature. Performing a first encoding on the first initial speech embedding feature and the second initial speech embedding feature to obtain the first initial speech encoded feature and the second initial speech encoded feature. The first initial speech encoded feature corresponds to the speech recognition result, and the second initial speech encoded feature corresponds to the adjusted speech recognition result. As Figure 4 , performing an embedding process on the speech recognition result x411 may obtain the first initial speech embedding feature e411. Performing an embedding process on the adjusted speech recognition result x412 may obtain the second initial speech embedding feature e412. Performing a first encoding on the first initial speech embedding feature e411 may obtain the first initial speech encoded feature fe411. Performing a first encoding on the second initial speech embedding feature e412 may obtain the second initial speech encoded feature. The first initial speech encoded feature fe411 may correspond to the speech recognition result x411. The second initial speech encoded feature fe412 may correspond to the adjusted speech recognition result x412. It can be understood that the first encoding may be various encoding methods. For example, the first encoding may be a multi-head self-attention encoding. Through the embodiments of the present disclosure, the same encoding method is adopted for the two speech recognition results before and after the adjustment using the first similar word, which can further reduce the speech recognition error caused by similar pronunciations and effectively improve the accuracy of video recognition.
[0117] In some embodiments, performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature may include: performing an embedding process on the text recognition result to obtain a first initial text embedding feature. Performing an embedding process on the adjusted text recognition result to obtain a second initial text embedding feature. Performing a second encoding on the first initial text embedding feature and the second initial text embedding feature to obtain a first initial text encoding feature and a second initial text encoding feature. The first initial text encoding feature corresponds to the text recognition result, and the second initial text encoding feature corresponds to the adjusted text recognition result. As Figure 4 , performing an embedding process on the text recognition result x421 may obtain a first initial text embedding feature e421. Performing an embedding process on the adjusted text recognition result x422 may obtain a second initial text embedding feature e422. Performing a second encoding on the first initial text embedding feature e421 may obtain a first initial text encoding feature fe421. Performing a second encoding on the second initial text embedding feature e422 may obtain a second initial text encoding feature fe422. The first initial text encoding feature fe421 may correspond to the text recognition result x421. The second initial text encoding feature fe422 may correspond to the adjusted text recognition result x422. It can be understood that the second encoding may be various encoding methods. For example, the first encoding may be a multi-head self-attention encoding. Through the embodiments of the present disclosure, the same encoding method is adopted for the two text recognition results before and after being adjusted by the second similar word, which can further reduce the optical character recognition error caused by structural similarity and effectively improve the accuracy of video recognition.
[0118] It can be understood that some methods for obtaining the encoding feature are described above, and below some methods for obtaining the feature to be decoded will be described.
[0119] In some embodiments, in some implementation manners of the above operation S250, obtaining the feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature includes: respectively processing the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature to obtain a first processed text feature, a second processed text feature, a first processed speech feature, and a second processed speech feature. Obtaining the feature to be decoded according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature.
[0120] In the embodiments of the present disclosure, at least one of a fully connected process and a softmax process may be performed on the encoded features. For example, taking the case of performing a fully connected process and a softmax process on the encoded features in sequence, the first initial speech encoded feature fe411 may be processed to obtain a first processed speech feature fp411. The second initial speech encoded feature fe412 may be processed to obtain a second processed speech feature fp412. The first initial text encoded feature fe421 may be processed to obtain a first processed text feature fp421. The second initial speech encoded feature fe422 may be processed to obtain a second processed text feature fp422. Through the embodiments of the present disclosure, different encoded features are processed respectively, which helps to highlight the effective information in different encoded features, so as to improve the efficiency and accuracy of video recognition.
[0121] In the embodiments of the present disclosure, obtaining the feature to be decoded according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature may include: fusing the first initial text embedding feature and the first processed text feature to obtain a first text fusion feature. Fusing the second initial text embedding feature and the second processed text feature to obtain a second text fusion feature. Fusing the first initial speech embedding feature and the first processed speech feature to obtain a first speech fusion feature. Fusing the second initial speech embedding feature and the second processed speech feature to obtain a second speech fusion feature. The fusion method between the embedding feature and the processed feature may be weighted fusion. As Figure 4 shown, the first processed speech feature fp411 may be used as a weight to weight the first initial speech embedding feature e411 to obtain a first speech fusion feature. The second processed speech feature fp412 may be used as a weight to weight the second initial speech embedding feature e412 to obtain a second speech fusion feature. The first processed text feature fp421 may be used as a weight to weight the first initial text embedding feature e421 to obtain a first text fusion feature. The second processed text feature fp422 may be used as a weight to weight the second initial text embedding feature e422 to obtain a second speech fusion feature. Through the embodiments of the present disclosure, weighted fusion is performed on the embedding feature and the processed feature, which can effectively highlight the effective information of different features and further improve the efficiency and accuracy of video recognition.
[0122] In the embodiments of the present disclosure, obtaining the feature to be decoded according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature may include: fusing the first text fusion feature, the second text fusion feature, the first speech fusion feature, and the second speech fusion feature to obtain the feature to be decoded. The fusion method between the fusion features may include concatenation. For example, the feature to be decoded em may be obtained through the following formula:
[0123]
[0124] e1, e2, e3, and e4 may be the first initial speech embedding feature, the second initial speech embedding feature, the first initial text embedding feature, and the second initial text embedding feature respectively. En1() may be the first encoding. En2() may be the second encoding. may be concatenation. fpro1(), fpro2(), fpro3(), and fpro4() respectively represent the processing of the embedding features.
[0125] It can be understood that the above describes the manner of obtaining the feature to be decoded in the present disclosure. Next, the decoding manner of the feature to be decoded will be described.
[0126] In some embodiments, in some implementation manners of the above operation S260, decoding the feature to be decoded to obtain the target recognition result may include: performing a third encoding on the feature to be decoded to obtain the encoded feature to be decoded. Decoding the encoded feature to be decoded to obtain the target recognition result. As Figure 4 shown, performing a third encoding on the feature to be decoded can obtain the encoded feature to be decoded de40. Decoding the encoded feature to be decoded de40 can obtain the target recognition result r40. The target recognition result r40 may be "Patients with viral colds can take Lianhua Qingwen Capsules and Ganmaoling Granules"
[0127] It can be understood that the above describes the method of the present disclosure in combination with a speech recognition result. However, the present disclosure is not limited thereto. For each speech recognition result among multiple speech recognition results, the above method 200 can be used for processing, which will be described below
[0128] In some embodiments, there are multiple speech recognition results and multiple text recognition results, and each speech recognition result corresponds to a text recognition result. For example, after dividing the first initial recognition result, multiple speech recognition results can be obtained. During the time period of each speech recognition result, a text recognition result corresponding to the speech recognition result can be determined.
[0129] In some embodiments, determining an adjusted speech recognition result based on at least one first word segmentation of a speech recognition result of a video to be processed and at least one first similar word corresponding to the at least one first word segmentation may include: determining an adjusted speech recognition result corresponding to each speech recognition result based on at least one first word segmentation of each speech recognition result and at least one first similar word corresponding to the at least one first word segmentation of each speech recognition result.
[0130] In some embodiments, determining an adjusted text recognition result based on at least one second word segmentation of a text recognition result corresponding to a speech recognition result and at least one second similar word corresponding to the at least one second word segmentation may include: determining an adjusted text recognition result corresponding to each text recognition result based on at least one second word segmentation of each text recognition result and at least one second similar word corresponding to the at least one second word segmentation of each text recognition result.
[0131] In some embodiments, a first encoding may be performed on each speech recognition result and the corresponding adjusted speech recognition result to obtain each first initial speech encoding feature and each second initial speech encoding feature. A second encoding may also be performed on each text recognition result and the corresponding adjusted text recognition result to obtain each first initial text encoding feature and each second initial text encoding feature. Based on each first initial speech encoding feature, the corresponding second initial speech encoding feature, the corresponding first initial text encoding feature, and the corresponding second initial text encoding feature, each decoding feature to be decoded can be obtained. Decoding each decoding feature to be decoded can obtain the corresponding target recognition result.
[0132] It can be understood that the video recognition method of the present disclosure has been described above, and the training method of the large language model of the present disclosure will be described below.
[0133] Figure 5 It is a schematic flowchart of a training method of a large language model according to an embodiment of the present disclosure.
[0134] As Figure 5 shown, the method 500 may include operation S510.
[0135] In operation S510, a large language model is trained using multiple target recognition results.
[0136] In an embodiment of the present disclosure, the target recognition result is obtained through the following operations: Determine an adjusted speech recognition result based on at least one first word segment of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segment. Determine an adjusted text recognition result based on at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment. Perform a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature. Perform a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature. Obtain a feature to be decoded based on the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature. Decode the feature to be decoded to obtain the target recognition result. It can be understood that the manner of obtaining the target recognition result is the same as or similar to the above method 200, and the present disclosure will not elaborate herein.
[0137] It can be understood that the above description of the present disclosure is given by taking the target recognition result for training a large language model as an example. However, the present disclosure is not limited thereto, and the following will be described in conjunction with Figure 6 for illustration.
[0138] Figure 6 is a schematic diagram of applying the target recognition result according to an embodiment of the present disclosure.
[0139] As Figure 6 shown, applying the target recognition result may include operation S601 to operation S605.
[0140] In operation S601, obtain the video to be processed. The video to be processed may include videos in a professional field. As Figure 6 shown, the video to be processed may include medical videos, legal videos, financial videos, educational videos, etc.
[0141] In operation S602, data preprocessing may be performed according to the video to be processed. As Figure 6 shown, operations S210 and S220 in the above method 200 may be performed on the video to be processed to perform data preprocessing and obtain a speech recognition result, a text recognition result, an adjusted speech recognition result, and an adjusted text recognition result.
[0142] In operation S603, encoding and decoding processing is performed. As Figure 6As shown, the above operations S230 to S260 can be executed according to the speech recognition result x611, the text recognition result x621, the adjusted speech recognition result x612, and the adjusted text recognition result x622 to perform encoding and decoding processing. Thus, by executing the above video recognition method, the target recognition result can be obtained.
[0143] In operation S604, the recognition result is output. As Figure 6 shown, the target recognition result r60 can be output.
[0144] In operation S605, the recognition result is applied. As Figure 6 shown, based on the target recognition result r60, corpora or knowledge bases such as a large language model corpus, an information retrieval knowledge base, an intelligent question answering knowledge base, and a knowledge graph corpus can be constructed.
[0145] It can be understood that the above describes the application method of the target recognition result of the present disclosure. The above video recognition method can be implemented using a video recognition model, and the training method of the video recognition model will be described below.
[0146] Figure 7 is a schematic flowchart of a training method of a video recognition model according to an embodiment of the present disclosure.
[0147] As Figure 7 shown, the method 700 may include operations S710 to S770.
[0148] In operation S710, based on at least one first word segment of the speech recognition result of the sample video to be processed and at least one first similar word corresponding to the at least one first word segment, the adjusted speech recognition result is determined.
[0149] In operation S720, based on at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment, the adjusted text recognition result is determined.
[0150] In operation S730, the speech recognition result and the adjusted speech recognition result are input into the first encoding sub-model of the video recognition model to obtain the first initial speech encoding feature and the second initial speech encoding feature.
[0151] In operation S740, the text recognition result and the adjusted text recognition result are input into the second encoding sub-model of the video recognition model to obtain the first initial text encoding feature and the second initial text encoding feature.
[0152] In operation S750, based on the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature, the feature to be decoded is obtained.
[0153] In operation S760, the feature to be decoded is decoded to obtain a sample recognition result.
[0154] In the embodiments of the present disclosure, the first encoding sub-model can implement the above-mentioned first encoding. The second encoding sub-model can implement the above-mentioned second encoding operation. The descriptions of operations S710 to S760 are the same as or similar to those of operations S210 to S260, and the present disclosure will not repeat them here.
[0155] In operation S770, the video recognition model is trained according to the sample recognition result and the label of the sample video to be processed.
[0156] In the embodiments of the present disclosure, the sample video to be processed can be manually annotated to obtain the label of the sample video to be processed. The label may include a caption of the sample video to be processed.
[0157] In the embodiments of the present disclosure, according to the difference between the sample recognition result and the label of the sample video to be processed, various loss functions can be used to train the video recognition model.
[0158] It can be understood that for operations S710 to S740, when operation S710 is executed first and then operation S730, and operation S720 is executed first and then operation S740, operations S710, S720, S730, and S740 can be executed in a variety of different orders. For example, operations S710, S730, S720, and S740 are executed in sequence. For example, operations S720, S710, S730, and S740 are executed in sequence. For example, operations S720, S740, S710, and S730 are executed in sequence. For another example, operations S710 and S720 are executed in parallel, operation S730 is executed after operation S710 is completed, and operation S740 is executed after operation S720 is completed.
[0159] It can be understood that the above describes the training method of the video recognition model of the present disclosure, and the following will further describe method 700.
[0160] In some embodiments, the above method 700 further includes: performing text recognition on the video to be processed to obtain a plurality of initial text boxes. The initial text boxes correspond to the text in the video to be processed, and the initial text boxes are used to indicate the positions of the text in the video to be processed. According to the plurality of initial text boxes, a text recognition result is obtained.
[0161] In some embodiments, obtaining a text recognition result based on a plurality of initial text frames includes: determining a plurality of position frequencies according to the positions of the plurality of initial text frames. The position frequencies correspond to at least one initial text frame. Determining a text recognition result according to at least one initial text frame corresponding to at least one target position frequency. The target position frequency is a position frequency greater than or equal to a preset frequency threshold among the plurality of position frequencies.
[0162] In some embodiments, obtaining a text recognition result based on a plurality of initial text frames includes: determining at least one initial text frame group according to the plurality of initial text frames. At least one initial text frame within the initial text frame group indicates the same position. Determining the information entropy of the initial text frame group according to at least one word corresponding to the initial text frame group. Determining a character recognition result according to at least one initial text frame group corresponding to at least one target information entropy. The target information entropy is an information entropy greater than or equal to a preset entropy threshold.
[0163] In some embodiments, there are multiple text recognition results. The method 700 further includes: determining a plurality of text recognition results to be processed from the multiple text recognition results according to the time period corresponding to the speech recognition result. Determining the text recognition result corresponding to the speech recognition result from the multiple text recognition results to be processed according to the multiple similarities between the speech recognition result and the multiple text recognition results to be processed.
[0164] In some embodiments, at least one first word segment corresponds to M first similar words, where M is an integer greater than or equal to 1. In some embodiments of the above operation S710, determining an adjusted speech recognition result according to at least one first word segment of the speech recognition result of the video to be processed and at least one first similar word corresponding to at least one first word segment includes: using the M first similar words corresponding to at least one first word segment to replace at least one first word segment to obtain K replaced speech recognition results, where K is an integer greater than or equal to M. Determining the adjusted speech recognition result according to the perplexity of the K replaced speech recognition results.
[0165] In some embodiments, the method 700 further includes: segmenting the speech recognition result to obtain N first word segments. N is an integer greater than or equal to 1. Determining M first similar words corresponding to at least one first word segment from a target thesaurus according to the N first word segments.
[0166] In some embodiments, at least one second participle corresponds to J second similar words, where J is an integer greater than or equal to 1. In some implementations of the above operation S710, based on at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second participle, determining the adjusted text recognition result includes: using the J second similar words corresponding to the at least one second participle to replace I second participles, obtaining L text recognition results after replacement. L is an integer greater than or equal to J. Based on the perplexity of the L text recognition results after replacement, determine the adjusted text recognition result.
[0167] In some embodiments, method 700 further includes: performing word segmentation on the text recognition result to obtain I second participles. Based on the I first participles, determine J second similar words corresponding to at least one second participle from the target word library.
[0168] It can be understood that the further description of method 700, operation S710, and operation S720 is the same as or similar to the further description of the above method 200, operation S210, and operation S220, and the present disclosure will not repeat it here.
[0169] It can be understood that the above has further described operations S710 and S720 of the present disclosure. Next, operations S730 to S760 of the present disclosure will be further described.
[0170] Figure 8 It is a schematic diagram of a video recognition model according to an embodiment of the present disclosure.
[0171] As Figure 8 shown, the video recognition model may include an embedding layer e80, a first encoding sub-model En81, a second encoding sub-model En82, a third encoding sub-model En83, a first speech processing layer fpro811, a second speech processing layer fpro812, a first text processing layer fpro821, a second text processing layer fpro822, and a decoding processing layer fcrt80. The first encoding sub-model En81 to the third encoding sub-model En83 may include one or more Transformer Encoders.
[0172] In some embodiments, in some implementations of the above operation S730, inputting the speech recognition result and the adjusted speech recognition result into the first encoding sub-model to obtain the first initial speech encoding feature and the second initial speech encoding feature may include: inputting the speech recognition result into the embedding layer of the video recognition model to obtain the first initial speech embedding feature. Inputting the adjusted speech recognition result into the embedding layer to obtain the second initial speech embedding feature. Inputting the first initial speech embedding feature and the second initial speech embedding feature into the first encoding sub-model to obtain the first initial speech encoding feature and the second initial speech encoding feature. The first initial speech encoding feature corresponds to the speech recognition result, and the second initial speech encoding feature corresponds to the adjusted speech recognition result. As Figure 8 shown, inputting the speech recognition result x811 into the embedding layer e80 can obtain the first initial speech embedding feature e811. Inputting the adjusted speech recognition result x811 into the embedding layer e80 can obtain the second initial speech embedding feature e812. Inputting the first initial speech embedding feature e811 and the second initial speech embedding feature e812 into the first encoding sub-model En81 can obtain the first initial speech encoding feature and the second initial speech encoding feature. The first initial speech encoding feature may correspond to the speech recognition result x811. The second initial speech encoding feature may correspond to the adjusted speech recognition result x812.
[0173] In some embodiments, in some implementations of the above operation S740, performing a second encoding on the text recognition result and the adjusted text recognition result to obtain the first initial text encoding feature and the second initial text encoding feature may further include: performing an embedding process on the text recognition result to obtain the first initial text embedding feature. Performing an embedding process on the adjusted text recognition result to obtain the second initial text embedding feature. Performing a second encoding on the first initial text embedding feature and the second initial text embedding feature to obtain the first initial text encoding feature and the second initial text encoding feature. The first initial text encoding feature corresponds to the text recognition result, and the second initial text encoding feature corresponds to the adjusted text recognition result. As Figure 8 shown, inputting the text recognition result x821 into the embedding layer e80 can obtain the first initial text embedding feature e821. Inputting the adjusted text recognition result x821 into the embedding layer e80 can obtain the second initial text embedding feature e822. Inputting the first initial text embedding feature e821 and the second initial text embedding feature e822 into the second encoding sub-model En82 can obtain the first initial text encoding feature and the second initial text encoding feature. The first initial text encoding feature may correspond to the text recognition result x821. The second initial text encoding feature may correspond to the adjusted text recognition result x822.
[0174] In some embodiments, in some implementations of the above operation S750, obtaining the feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature may include: respectively processing the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature to obtain a first processed text feature, a second processed text feature, a first processed speech feature, and a second processed speech feature. Obtaining the feature to be decoded according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature.
[0175] In an embodiment of the present disclosure, the processing layer may perform at least one of a fully connected process and a softmax process on the input feature. For example, taking the processing layer performing a fully connected process and a softmax process on the encoding feature in sequence as an example, inputting the first initial speech encoding feature into the first speech processing layer fpro811 may obtain a first processed speech feature. Inputting the second initial speech encoding feature into the second speech processing layer fpro812 may obtain a second processed speech feature. Inputting the first initial text encoding feature into the first text processing layer fpro821 may obtain a first processed text feature. Inputting the second initial text encoding feature into the second text processing layer fpro822 may obtain a second processed text feature.
[0176] In an embodiment of the present disclosure, obtaining the feature to be decoded according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature may include: fusing the first initial text embedding feature and the first processed text feature to obtain a first text fusion feature. Fusing the second initial text embedding feature and the second processed text feature to obtain a second text fusion feature. Fusing the first initial speech embedding feature and the first processed speech feature to obtain a first speech fusion feature. Fusing the second initial speech embedding feature and the second processed speech feature to obtain a second speech fusion feature. The fusion method between the embedding feature and the processed feature may be weighted fusion. As Figure 8 shown, the first processed speech feature may be used as a weight to weight the first initial speech embedding feature e811 to obtain a first speech fusion feature. The second processed speech feature may be used as a weight to weight the second initial speech embedding feature e812 to obtain a second speech fusion feature. The first processed text feature may be used as a weight to weight the first initial text embedding feature e821 to obtain a first text fusion feature. The second processed text feature may be used as a weight to weight the second initial text embedding feature e822 to obtain a second speech fusion feature.
[0177] In an embodiment of the present disclosure, obtaining the feature to be decoded based on the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature may include: fusing the first text fusion feature, the second text fusion feature, the first speech fusion feature, and the second speech fusion feature to obtain the feature to be decoded. The fusion can be performed through the above formula three.
[0178] In some embodiments, in some implementations of the above operation S760, decoding the feature to be decoded to obtain the target recognition result may include: performing a third encoding on the feature to be decoded to obtain the encoded feature to be decoded. Decoding the encoded feature to be decoded to obtain the target recognition result. As Figure 8 shown, inputting the feature to be decoded into the third encoding sub-model En83 can obtain the encoded feature to be decoded. Inputting the encoded feature to be decoded into the decoding processing layer fcrt80 can obtain the sample recognition result r80.
[0179] It can be understood that the encoding and decoding methods of the present disclosure have been described above, and the method for training the video recognition model will be further described below.
[0180] In some embodiments, a sample video to be processed may be obtained, and the subtitles of the video may be manually proofread to obtain the corrected subtitles, so as to obtain the labels for the sample to be processed for recognition. The sample video to be processed may be a video in a professional field, and there may be problems such as missing subtitles and incorrect professional terms in the video. Through manual proofreading, accurate labels can be obtained, which can improve the training effect and the accuracy of the video recognition model.
[0181] In some embodiments, the video recognition model may be trained according to the labels and the sample recognition results. As Figure 8 shown, according to the sample recognition result r80 and the corresponding label, the loss value may be determined. Adjusting the parameters of at least one of the embedding layer e80, the first encoding sub-model En81, the second encoding sub-model En82, the third encoding sub-model En83, the first speech processing layer fpro811, the second speech processing layer fpro812, the first text processing layer fpro821, the second text processing layer fpro822, and the decoding processing layer fcrt80 according to the loss value to train the video recognition model.
[0182] It can be understood that the present disclosure has been described above by taking the recognition result as Chinese as an example. However, the present disclosure is not limited thereto, and the recognition result may be various languages such as English and German.
[0183] It can be understood that the method of the present disclosure has been described above, and the apparatus of the present disclosure will be described below.
[0184] Figure 9It is a block diagram of a video recognition device according to an embodiment of the present disclosure.
[0185] As Figure 9 shown, the device 900 may include a first determination model 910, a second determination module 920, a first encoding module 930, a second encoding module 940, a first acquisition module 950, and a first decoding module 960.
[0186] The first determination module 910 is configured to determine an adjusted speech recognition result according to at least one first word segmentation of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segmentation.
[0187] The second determination module 920 is configured to determine an adjusted text recognition result according to at least one second word segmentation of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segmentation.
[0188] The first encoding module 930 is configured to perform a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature.
[0189] The second encoding module 940 is configured to perform a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature.
[0190] The first acquisition module 950 is configured to obtain a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature.
[0191] The first decoding module 960 is configured to decode the feature to be decoded to obtain a target recognition result.
[0192] In some embodiments, the device 900 further includes: a first text recognition module, configured to perform text recognition on the video to be processed to obtain a plurality of initial text boxes. The initial text boxes correspond to the text in the video to be processed, and the initial text boxes are used to indicate the positions of the text in the video to be processed. A third acquisition module, configured to obtain a text recognition result according to the plurality of initial text boxes.
[0193] In some embodiments, the third acquisition module includes: a first determination sub-module, configured to determine a plurality of position frequencies according to the positions of the plurality of initial text boxes. The position frequencies correspond to at least one initial text box. A second determination sub-module, configured to determine a text recognition result according to at least one initial text box corresponding to at least one target position frequency. The target position frequency is the position frequency greater than or equal to a preset frequency threshold among the plurality of position frequencies.
[0194] In some embodiments, the third acquisition module includes: a third determination sub-module, configured to determine at least one initial text box group according to a plurality of initial text boxes. At least one initial text box within the initial text box group indicates the same position. A fourth determination sub-module, configured to determine the information entropy of the initial text box group according to at least one text corresponding to the initial text box group. A fifth determination sub-module, configured to determine a text recognition result according to at least one initial text box group corresponding to at least one target information entropy. The target information entropy is an information entropy greater than or equal to a preset entropy threshold.
[0195] In some embodiments, there are multiple text recognition results. The apparatus 900 further includes: a fifth determination module, configured to determine multiple to-be-processed text recognition results from the multiple text recognition results according to the time period corresponding to the speech recognition result. A sixth determination module, configured to determine the text recognition result corresponding to the speech recognition result from the multiple to-be-processed text recognition results according to multiple similarities between the speech recognition result and the multiple to-be-processed text recognition results.
[0196] In some embodiments, at least one first word segmentation corresponds to M first similar words, where M is an integer greater than or equal to 1. The first determination module includes: a first replacement sub-module, configured to replace at least one first word segmentation with the M first similar words corresponding to the at least one first word segmentation to obtain K replaced speech recognition results. K is an integer greater than or equal to M. A sixth determination sub-module, configured to determine an adjusted speech recognition result according to the perplexity of the K replaced speech recognition results.
[0197] In some embodiments, the apparatus 900 further includes: a first word segmentation module, configured to perform word segmentation on the speech recognition result to obtain N first word segmentations. N is an integer greater than or equal to 1. A seventh determination module, configured to determine M first similar words corresponding to at least one first word segmentation from a target word library according to the N first word segmentations.
[0198] In some embodiments, at least one second word segmentation corresponds to J second similar words, where J is an integer greater than or equal to 1. The second determination module includes: a second replacement sub-module, configured to replace I second word segmentations with the J second similar words corresponding to the at least one second word segmentation to obtain L replaced text recognition results, where L is an integer greater than or equal to J. A seventh determination sub-module, configured to determine an adjusted text recognition result according to the perplexity of the L replaced text recognition results.
[0199] In some embodiments, the apparatus 900 further includes: a second word segmentation module, configured to perform word segmentation on the text recognition result to obtain I second word segmentations, where I is an integer greater than or equal to 1. An eighth determination module, configured to determine J second similar words corresponding to at least one second word segmentation from a target word library according to the I first word segmentations.
[0200] In some embodiments, the first encoding module includes: a first embedding sub-module, configured to perform embedding processing on the speech recognition result to obtain a first initial speech embedding feature; a second embedding sub-module, configured to perform embedding processing on the adjusted speech recognition result to obtain a second initial speech embedding feature; and a first encoding sub-module, configured to perform a first encoding on the first initial speech embedding feature and the second initial speech embedding feature to obtain a first initial speech encoding feature and a second initial speech encoding feature. The first initial speech encoding feature corresponds to the speech recognition result, and the second initial speech encoding feature corresponds to the adjusted speech recognition result.
[0201] In some embodiments, the second encoding module includes: a third embedding sub-module, configured to perform embedding processing on the text recognition result to obtain a first initial text embedding feature; a fourth embedding sub-module, configured to perform embedding processing on the adjusted text recognition result to obtain a second initial text embedding feature; and a second encoding sub-module, configured to perform a second encoding on the first initial text embedding feature and the second initial text embedding feature to obtain a first initial text encoding feature and a second initial text encoding feature. The first initial text encoding feature corresponds to the text recognition result, and the second initial text encoding feature corresponds to the adjusted text recognition result.
[0202] In some embodiments, the first obtaining module includes: a first processing sub-module, configured to process the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature respectively to obtain a first processed text feature, a second processed text feature, a first processed speech feature, and a second processed speech feature; and a first obtaining sub-module, configured to obtain a feature to be decoded according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature.
[0203] In some embodiments, the first obtaining sub-module includes: a first fusion unit, configured to fuse the first initial text embedding feature and the first processed text feature to obtain a first text fusion feature; a second fusion unit, configured to fuse the second initial text embedding feature and the second processed text feature to obtain a second text fusion feature; a third fusion unit, configured to fuse the first initial speech embedding feature and the first processed speech feature to obtain a first speech fusion feature; a fourth fusion unit, configured to fuse the second initial speech embedding feature and the second processed speech feature to obtain a second speech fusion feature; and a fifth fusion unit, configured to fuse the first text fusion feature, the second text fusion feature, the first speech fusion feature, and the second speech fusion feature to obtain a feature to be decoded.
[0204] In some embodiments, the first decoding module includes: a third encoding sub-module for performing third encoding on the feature to be decoded to obtain the encoded feature to be decoded; and a first decoding sub-module for decoding the encoded feature to be decoded to obtain the target recognition result.
[0205] In some embodiments, there are multiple speech recognition results and multiple text recognition results, and each speech recognition result corresponds to one text recognition result. The first determination module is further configured to: determine the adjusted speech recognition result corresponding to each speech recognition result according to at least one first word segmentation of each speech recognition result and at least one first similar word corresponding to at least one first word segmentation of each speech recognition result.
[0206] In some embodiments, there are multiple speech recognition results and multiple text recognition results, and each speech recognition result corresponds to one text recognition result. The second determination module is further configured to: determine the adjusted text recognition result corresponding to each text recognition result according to at least one second word segmentation of each text recognition result and at least one second similar word corresponding to at least one second word segmentation of each text recognition result.
[0207] In some embodiments, the first similar word is a homophone of the first word segmentation, and the second similar word is a similar-looking word of the second word segmentation.
[0208] Figure 10 It is a block diagram of a training device for a large language model according to another embodiment of the present disclosure.
[0209] As Figure 10 shown, the device 1000 may include a first training module 1010.
[0210] The first training module 1010 is configured to train the large language model by using multiple target recognition results.
[0211] In the embodiments of the present disclosure, the target recognition result is obtained through the following operations: determining an adjusted speech recognition result according to at least one first word segment of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first word segment; determining an adjusted text recognition result according to at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment; performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature; performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and decoding the feature to be decoded to obtain the target recognition result.
[0212] Figure 11 It is a block diagram of a training device for a video recognition model according to another embodiment of the present disclosure.
[0213] As Figure 11 shown, the device 1100 may include a third determination module 1110, a fourth determination module 1120, a third encoding module 1130, a fourth encoding module 1140, a second obtaining module 1150, a second decoding module 1160, and a second training module 1170.
[0214] The third determination module 1110 is configured to determine an adjusted speech recognition result according to at least one first word segment of the speech recognition result of the sample video to be processed and at least one first similar word corresponding to the at least one first word segment.
[0215] The fourth determination module 1120 is configured to determine an adjusted text recognition result according to at least one second word segment of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to the at least one second word segment.
[0216] The third encoding module 1130 is configured to input the speech recognition result and the adjusted speech recognition result into a first encoding sub-model of the video recognition model to obtain a first initial speech encoding feature and a second initial speech encoding feature.
[0217] The fourth encoding module 1140 is configured to input the text recognition result and the adjusted text recognition result into a second encoding sub-model of the video recognition model to obtain a first initial text encoding feature and a second initial text encoding feature.
[0218] A second obtaining module 1150, configured to obtain a feature to be decoded according to a first initial text encoding feature, a second initial text encoding feature, a first initial speech encoding feature, and a second initial speech encoding feature.
[0219] A second decoding module 1160, configured to decode the feature to be decoded to obtain a sample recognition result.
[0220] A second training module 1170, configured to train a video recognition model according to the sample recognition result and the label of the sample video to be processed.
[0221] In some embodiments, the apparatus 1100 further includes: a second text recognition module, configured to perform text recognition on the video to be processed to obtain a plurality of initial text frames. The initial text frames correspond to the text in the video to be processed, and the initial text frames are used to indicate the positions of the text in the video to be processed. A fourth obtaining module, configured to obtain a text recognition result according to the plurality of initial text frames.
[0222] In some embodiments, the fourth obtaining module includes: a first determining sub-module, configured to determine a plurality of position frequencies according to the positions of the plurality of initial text frames. The position frequencies correspond to at least one initial text frame. An eighth determining sub-module, configured to determine a text recognition result according to at least one initial text frame corresponding to at least one target position frequency. The target position frequency is a position frequency greater than or equal to a preset frequency threshold among the plurality of position frequencies.
[0223] In some embodiments, the fourth obtaining module includes: a ninth determining sub-module, configured to determine at least one group of initial text frames according to the plurality of initial text frames. At least one initial text frame within the group of initial text frames indicates the same position. A tenth determining sub-module, configured to determine the information entropy of the group of initial text frames according to at least one text corresponding to the group of initial text frames. An eleventh determining sub-module, configured to determine a text recognition result according to at least one group of initial text frames corresponding to at least one target information entropy. The target information entropy is an information entropy greater than or equal to a preset entropy threshold.
[0224] In some embodiments, there are multiple text recognition results. The apparatus 1100 further includes: a ninth determining module, configured to determine multiple text recognition results to be processed from the multiple text recognition results according to the time period corresponding to the speech recognition result. A tenth determining module, configured to determine the text recognition result corresponding to the speech recognition result from the multiple text recognition results to be processed according to the multiple similarities between the speech recognition result and the multiple text recognition results to be processed.
[0225] In some embodiments, at least one first word segment corresponds to M first similar words, where M is an integer greater than or equal to 1. The third determination module includes: a third replacement sub-module, configured to use the M first similar words corresponding to at least one first word segment to replace the at least one first word segment, obtaining K replaced speech recognition results, where K is an integer greater than or equal to M. A twelfth determination sub-module, configured to determine an adjusted speech recognition result according to the perplexity of the K replaced speech recognition results.
[0226] In some embodiments, the apparatus 1100 further includes: a third word segmentation module, configured to perform word segmentation on the speech recognition result to obtain N first word segments, where N is an integer greater than or equal to 1. An eleventh determination module, configured to determine, according to the N first word segments, M first similar words corresponding to at least one first word segment from a target word library.
[0227] In some embodiments, at least one second word segment corresponds to J second similar words, where J is an integer greater than or equal to 1. The fourth determination module includes: a fourth replacement sub-module, configured to use the J second similar words corresponding to at least one second word segment to replace I second word segments, obtaining L replaced text recognition results, where L is an integer greater than or equal to J. A twelfth determination sub-module, configured to determine an adjusted text recognition result according to the perplexity of the L replaced text recognition results.
[0228] In some embodiments, the apparatus 900 further includes: a fourth word segmentation module, configured to perform word segmentation on the text recognition result to obtain I second word segments, where I is an integer greater than or equal to 1. A twelfth determination module, configured to determine, according to the I first word segments, J second similar words corresponding to at least one second word segment from the target word library.
[0229] In some embodiments, the third encoding module includes: a fifth embedding sub-module, configured to perform an embedding process on the speech recognition result to obtain a first initial speech embedding feature. A sixth embedding sub-module, configured to perform an embedding process on the adjusted speech recognition result to obtain a second initial speech embedding feature. A third encoding sub-module, configured to perform a first encoding on the first initial speech embedding feature and the second initial speech embedding feature to obtain a first initial speech encoding feature and a second initial speech encoding feature. The first initial speech encoding feature corresponds to the speech recognition result, and the second initial speech encoding feature corresponds to the adjusted speech recognition result.
[0230] In some embodiments, the fourth encoding module includes: a seventh embedding sub-module for performing embedding processing on the text recognition result to obtain a first initial text embedding feature; an eighth embedding sub-module for performing embedding processing on the adjusted text recognition result to obtain a second initial text embedding feature; and a fourth encoding sub-module for performing a second encoding on the first initial text embedding feature and the second initial text embedding feature to obtain a first initial text encoding feature and a second initial text encoding feature. The first initial text encoding feature corresponds to the text recognition result, and the second initial text encoding feature corresponds to the adjusted text recognition result.
[0231] In some embodiments, the second obtaining module includes: a second processing sub-module for respectively processing the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature to obtain a first processed text feature, a second processed text feature, a first processed speech feature, and a second processed speech feature; and a second obtaining sub-module for obtaining a to-be-decoded feature according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature.
[0232] In some embodiments, the second obtaining sub-module includes: a sixth fusion unit for fusing the first initial text embedding feature and the first processed text feature to obtain a first text fusion feature; a seventh fusion unit for fusing the second initial text embedding feature and the second processed text feature to obtain a second text fusion feature; an eighth fusion unit for fusing the first initial speech embedding feature and the first processed speech feature to obtain a first speech fusion feature; a ninth fusion unit for fusing the second initial speech embedding feature and the second processed speech feature to obtain a second speech fusion feature; and a tenth fusion unit for fusing the first text fusion feature, the second text fusion feature, the first speech fusion feature, and the second speech fusion feature to obtain a to-be-decoded feature.
[0233] In some embodiments, the second decoding module includes: a sixth encoding sub-module for performing a third encoding on the to-be-decoded feature to obtain an encoded to-be-decoded feature; and a second decoding sub-module for decoding the encoded to-be-decoded feature to obtain a target recognition result.
[0234] In some embodiments, there are multiple speech recognition results and multiple text recognition results, and each speech recognition result corresponds to a text recognition result. The third determination module is further configured to: determine an adjusted speech recognition result corresponding to each speech recognition result according to at least one first word segmentation of each speech recognition result and at least one first similar word corresponding to at least one first word segmentation of each speech recognition result.
[0235] In some embodiments, there are multiple speech recognition results and multiple text recognition results, and each speech recognition result corresponds to a text recognition result. The fourth determination module is further configured to: determine an adjusted text recognition result corresponding to each text recognition result according to at least one second word segmentation of each text recognition result and at least one second similar word corresponding to at least one second word segmentation of each text recognition result.
[0236] In some embodiments, the first similar word is a homophone of the first word segmentation, and the second similar word is a similar-looking word of the second word segmentation.
[0237] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0238] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0239] Figure 12 FIG. shows a schematic block diagram of an exemplary electronic device 1200 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0240] As Figure 12 shown, the device 1200 includes a computing unit 1201, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0241] Multiple components in device 1200 are connected to I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a disk, an optical disc, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0242] The computing unit 1201 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a Central Processing Unit (CPU), a Graph Processing Unit (GPU), various dedicated Artificial Intelligence (AI) computing chips, various computing units running machine learning model algorithms, a Digital Signal Processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as a video recognition method, a training method for a large language model, and / or a training method for a video recognition model. For example, in some embodiments, the video recognition method, the training method for a large language model, and / or the training method for a video recognition model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the video recognition method, the training method for a large language model, and / or the training method for a video recognition model described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute the video recognition method, the training method for a large language model, and / or the training method for a video recognition model by any other suitable means (e.g., by means of firmware).
[0243] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0244] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0245] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0246] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0247] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0248] A computer system may include a client and a server. The client and the server are generally far apart from each other and typically interact via a communication network. The relationship between the client and the server is generated by computer programs that run on the respective computers and have a client-server relationship with each other.
[0249] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0250] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A video recognition method, comprising: Determine an adjusted speech recognition result according to at least one first participle of a speech recognition result of a video to be processed and at least one first similar word corresponding to at least one first participle, wherein the first similar word is a homophone of the first participle in a target word library, the video to be processed is a video in a professional field, and the target word library is a word library constructed for the professional field; Determining an adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle, wherein the second similar word is a word similar in form to the second participle in the target word library; Performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech coding feature and a second initial speech coding feature; Performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; Obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and Decode the features to be decoded to obtain the target recognition result.
2. The method according to claim 1, further comprising: Performing text recognition on the video to be processed to obtain a plurality of initial text frames, wherein the initial text frames correspond to the text in the video to be processed, and the initial text frames are used to indicate the positions of the text in the video to be processed; and The text recognition result is obtained according to the multiple initial text frames.
3. The method according to claim 2, wherein: The obtaining of the text recognition result according to the plurality of initial text frames comprises: Determining a plurality of position frequencies according to respective positions of the plurality of initial text frames, wherein the position frequency corresponds to at least one of the initial text frames; The text recognition result is determined according to at least one initial text frame corresponding to at least one target position frequency, wherein the target position frequency is a position frequency greater than or equal to a preset frequency threshold among the plurality of position frequencies.
4. The method according to claim 2 or 3, wherein: The obtaining of the text recognition result according to the plurality of initial text frames comprises: According to the plurality of initial text frames, at least one initial text frame group is determined, wherein at least one initial text frame in the initial text frame group indicates the same position; Determining the information entropy of the initial text frame group according to at least one of the characters corresponding to the initial text frame group; The text recognition result is determined according to at least one of the initial text frame groups corresponding to at least one target information entropy, wherein the target information entropy is information entropy greater than or equal to a preset entropy threshold.
5. The method according to claim 2, wherein: The text recognition results are multiple, The method further comprises: Determining a plurality of text recognition results to be processed from the plurality of text recognition results according to the time period corresponding to the speech recognition result; According to a plurality of similarities between the speech recognition result and a plurality of the to-be-processed text recognition results, the text recognition result corresponding to the speech recognition result is determined from the plurality of the to-be-processed text recognition results.
6. The method according to claim 1, wherein: At least one of the first participles corresponds to M of the first similar words, where M is an integer greater than or equal to 1. The step of determining the adjusted speech recognition result according to at least one first participle of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first participle comprises: Using M first similar words corresponding to at least one first participle to replace at least one first participle, to obtain K replaced speech recognition results, where K is an integer greater than or equal to 1; The adjusted speech recognition result is determined according to the perplexities of the K replaced speech recognition results.
7. The method according to claim 6, further comprising: Performing word segmentation on the speech recognition result to obtain N first word segmentations, where N is an integer greater than or equal to 1; According to the N first participles, M first similar words corresponding to at least one of the first participles are determined from a target word library.
8. The method according to claim 1, wherein: At least one of the second participles corresponds to J of the second similar words, where J is an integer greater than or equal to 1. The determining of the adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle comprises: Using J second similar words corresponding to at least one second participle to replace at least one second participle, to obtain L replaced text recognition results, where L is an integer greater than or equal to 1; The adjusted text recognition result is determined according to the perplexities of the L replaced text recognition results.
9. The method according to claim 8, further comprising: Performing word segmentation on the text recognition result to obtain I second word segmentations, where I is an integer greater than or equal to 1; According to the I first participles, J second similar words corresponding to at least one second participle are determined from a target word library.
10. The method according to claim 1, wherein: The performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech coding feature and a second initial speech coding feature comprises: Performing embedding processing on the speech recognition result to obtain a first initial speech embedding feature; Performing embedding processing on the adjusted speech recognition result to obtain a second initial speech embedding feature; The first encoding is performed on the first initial speech embedding feature and the second initial speech embedding feature to obtain the first initial speech encoding feature and the second initial speech encoding feature, wherein the first initial speech encoding feature corresponds to the speech recognition result, and the second initial speech encoding feature corresponds to the adjusted speech recognition result.
11. The method according to claim 1, wherein: The performing second encoding on the text recognition result and the adjusted text recognition result to obtain the first initial text encoding feature and the second initial text encoding feature comprises: Performing embedding processing on the text recognition result to obtain a first initial text embedding feature; Performing embedding processing on the adjusted text recognition result to obtain a second initial text embedding feature; The first initial text embedding feature and the second initial text embedding feature are subjected to the second encoding to obtain the first initial text encoding feature and the second initial text encoding feature, wherein the first initial text encoding feature corresponds to the text recognition result, and the second initial text encoding feature corresponds to the adjusted text recognition result.
12. The method according to claim 1, wherein: The obtaining of the to-be-decoded feature according to the first initial text coding feature, the second initial text coding feature, the first initial speech coding feature, and the second initial speech coding feature comprises: Respectively processing the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature to obtain a first processed text feature, a second processed text feature, a first processed speech feature, and a second processed speech feature; The to-be-decoded feature is obtained according to the first processed text feature, the second processed text feature, the first processed speech feature and the second processed speech feature.
13. The method according to claim 12, wherein: The obtaining the to-be-decoded feature according to the first processed text feature, the second processed text feature, the first processed speech feature, and the second processed speech feature comprises: fusing the first initial text embedding feature and the first processed text feature to obtain a first text fusion feature; fusing the second initial text embedding feature and the second processed text feature to obtain a second text fusion feature; fusing the first initial speech embedding feature and the first processed speech feature to obtain a first speech fusion feature; Fusing the second initial speech embedding feature and the second processed speech feature to obtain a second speech fusion feature; The first text fusion feature, the second text fusion feature, the first speech fusion feature and the second speech fusion feature are fused to obtain the feature to be decoded.
14. The method according to claim 1, wherein: Decoding the feature to be decoded to obtain the target recognition result includes: Performing a third encoding on the feature to be decoded to obtain an encoded feature to be decoded; The encoded feature to be decoded is decoded to obtain the target recognition result.
15. The method according to any one of claims 1 to 3, 5 to 14, wherein: There are multiple speech recognition results, and there are multiple text recognition results, and each speech recognition result corresponds to one text recognition result. The step of determining the adjusted speech recognition result according to at least one first participle of the speech recognition result of the video to be processed and at least one first similar word corresponding to the at least one first participle comprises: According to at least one first participle of each of the speech recognition results and at least one first similar word corresponding to at least one first participle of each of the speech recognition results, an adjusted speech recognition result corresponding to each of the speech recognition results is determined.
16. The method according to any one of claims 1 to 3, 5 to 14, wherein: There are multiple speech recognition results, and there are multiple text recognition results, and each speech recognition result corresponds to one text recognition result. The determining of the adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle comprises: An adjusted text recognition result corresponding to each of the text recognition results is determined according to at least one of the second participles of each of the text recognition results and at least one second similar word corresponding to at least one of the second participles of each of the text recognition results.
17. A method for training a large language model, comprising: The large language model is trained using multiple target recognition results, wherein the target recognition results are obtained by the following operations: Determine an adjusted speech recognition result according to at least one first participle of a speech recognition result of a video to be processed and at least one first similar word corresponding to at least one first participle, wherein the first similar word is a homophone of the first participle in a target word library, the video to be processed is a video in a professional field, and the target word library is a word library constructed for the professional field; Determining an adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle, wherein the second similar word is a word similar in form to the second participle in the target word library; Performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech coding feature and a second initial speech coding feature; Performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; Obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and The feature to be decoded is decoded to obtain the target recognition result.
18. A method for training a video recognition model, comprising: Determine an adjusted speech recognition result according to at least one first participle of a speech recognition result of a sample video to be processed and at least one first similar word corresponding to at least one first participle, wherein the first similar word is a homophone of the first participle in a target word library, the video to be processed is a video in a professional field, and the target word library is a word library constructed for the professional field; Determining an adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle, wherein the second similar word is a word similar in form to the second participle in the target word library; Inputting the speech recognition result and the adjusted speech recognition result into a first coding sub-model of a video recognition model to obtain a first initial speech coding feature and a second initial speech coding feature; Inputting the text recognition result and the adjusted text recognition result into a second encoding sub-model of the video recognition model to obtain a first initial text encoding feature and a second initial text encoding feature; Obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; Decoding the features to be decoded to obtain sample recognition results; and The video recognition model is trained according to the sample recognition result and the label of the sample video to be processed.
19. A video recognition device, comprising: A first determination module is used to determine an adjusted speech recognition result according to at least one first segmented word of the speech recognition result of the video to be processed and at least one first similar word corresponding to at least one first segmented word, wherein the first similar word is a homophone of the first segmented word in a target word library, the video to be processed is a video in a professional field, and the target word library is a word library constructed for the professional field; a second determination module, configured to determine an adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle, wherein the second similar word is a word similar in form to the second participle in a target word library; A first encoding module, configured to perform a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech encoding feature and a second initial speech encoding feature; A second encoding module, used for performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; a first obtaining module, configured to obtain a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and The first decoding module is used to decode the features to be decoded to obtain a target recognition result.
20. A large language model training device, comprising: The first training module is used to train a large language model using multiple target recognition results, wherein the target recognition results are obtained by the following operations: Determine an adjusted speech recognition result according to at least one first participle of a speech recognition result of a video to be processed and at least one first similar word corresponding to at least one first participle, wherein the first similar word is a homophone of the first participle in a target word library, the video to be processed is a video in a professional field, and the target word library is a word library constructed for the professional field; Determining an adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle, wherein the second similar word is a word similar in form to the second participle in the target word library; Performing a first encoding on the speech recognition result and the adjusted speech recognition result to obtain a first initial speech coding feature and a second initial speech coding feature; Performing a second encoding on the text recognition result and the adjusted text recognition result to obtain a first initial text encoding feature and a second initial text encoding feature; Obtaining a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; and The feature to be decoded is decoded to obtain the target recognition result.
21. A training device for a video recognition model, comprising: a third determination module, configured to determine an adjusted speech recognition result according to at least one first segmented word of the speech recognition result of the sample video to be processed and at least one first similar word corresponding to at least one first segmented word, wherein the first similar word is a homophone of the first segmented word in a target word library, the video to be processed is a video in a professional field, and the target word library is a word library constructed for the professional field; a fourth determination module, configured to determine an adjusted text recognition result according to at least one second participle of the text recognition result corresponding to the speech recognition result and at least one second similar word corresponding to at least one second participle, wherein the second similar word is a word similar in form to the second participle in the target word library; A third encoding module, configured to input the speech recognition result and the adjusted speech recognition result into a first encoding sub-model of a video recognition model to obtain a first initial speech encoding feature and a second initial speech encoding feature; a fourth encoding module, configured to input the text recognition result and the adjusted text recognition result into a second encoding sub-model of the video recognition model to obtain a first initial text encoding feature and a second initial text encoding feature; A second obtaining module, configured to obtain a feature to be decoded according to the first initial text encoding feature, the second initial text encoding feature, the first initial speech encoding feature, and the second initial speech encoding feature; A second decoding module is used to decode the feature to be decoded to obtain a sample recognition result; and The second training module is used to train the video recognition model according to the sample recognition result and the label of the sample video to be processed.
22. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 18.
23. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 18.
24. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Speech recognition method, device and equipment based on multi-system fusion and readable storage medium
CN116168706A
Subtitle recognition method and device, equipment, storage medium and program product
CN117218635A