Speech recognition method and system for synthesizing video, and storage medium
By extracting the synthetic videos multimodal feature and analyzing them using edge smoothing index and texture repeating index, the problem of low speech recognition accuracy in synthetic videos in professional fields is solved, and higher speech recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510395651.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-31
AI Technical Summary
The existing speech recognition technology is not very accurate when facing overly professional synthetic videos in the field.
By extracting the synthetic videos multimodal feature, analyzing them using edge smoothing index and texture repeating index, multimodal feature vectors are established, domain recognition is performed and speech recognition is optimized.
Improves the accuracy of speech recognition in specific professional fields and reduces the rate of error.
Smart Images

Figure CN120260544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech recognition, and particularly relates to a speech recognition method, system and storage medium for synthesizing videos. Background Art
[0002] A synthesized video refers to dynamic video content generated by fusing static images, dynamic images, virtual scenes or real data through technologies such as artificial intelligence. Synthesized videos can not only reduce production costs but also achieve infinite extension of creative content. In the fields of film and television entertainment, education and training, virtual anchors, etc., synthesized videos have been widely used. For synthesized videos, speech recognition is an essential part, especially playing a core role in aspects such as automatic subtitle generation and speech-driven animation.
[0003] However, when the synthesized video involves overly specialized fields (such as medicine, law, professional technology, etc.), industry terms and complex sentence patterns may lead to recognition errors, thereby affecting subtitle accuracy and video credibility.
[0004] Therefore, people need a speech recognition method for synthesizing videos that can improve the speech recognition accuracy in specialized fields. Summary of the Invention
[0005] The purpose of the present invention is to provide a speech recognition method, system and storage medium for synthesizing videos, and solve the following technical problems:
[0006] Existing speech recognition technologies have low accuracy when facing synthesized videos in overly specialized fields.
[0007] The purpose of the present invention can be achieved through the following technical solutions:
[0008] A speech recognition method for synthesizing videos includes the following steps:
[0009] Obtain a target synthesized video, and perform multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, where the multi-modal feature data includes an edge smoothing index and a texture repetition index. The edge smoothing index is used to characterize the smoothness of the edges of the picture content in the target synthesized video, and the texture repetition index is used to characterize the repetition degree of the texture of the picture content in the target synthesized video;
[0010] According to the multi-modal feature data, establish a multi-modal feature vector;
[0011] According to the multi-modal feature vector, perform field recognition on the content of the target synthesized video to obtain content field classification data of the target synthesized video;
[0012] Based on the content field classification data, perform speech recognition on the target synthesized video to obtain a speech recognition result.
[0013] As a further solution of the present invention: Obtain a target synthesized video, and perform multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, including:
[0014] Obtain key images according to the target synthesized video;
[0015] Perform image feature extraction on the key images to obtain an edge smoothness index;
[0016] Perform image feature extraction on the key images to obtain a texture repetition index.
[0017] As a further solution of the present invention: Obtain key images according to the target synthesized video, including:
[0018] Extract key video frames from the target synthesized video;
[0019] Perform object detection on the key video frames to obtain the position and size of the target object;
[0020] Crop the key video frames according to the position and size of the target object to obtain target object images;
[0021] Both the key video frames and the target object images are used as key images.
[0022] As a further solution of the present invention: Perform image feature extraction on the key images to obtain an edge smoothness index, including:
[0023] Obtain the image size of the key images;
[0024] Statistically analyze the pixel differences between adjacent pixels in the key images;
[0025] Obtain the edge smoothness index according to the image size and the pixel differences.
[0026] As a further solution of the present invention: Perform image feature extraction on the key images to obtain a texture repetition index, including:
[0027] Analyze the pixel changes of the key images in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction;
[0028] Obtain the texture repetition index according to the pixel repetition periods corresponding to the multiple preset directions.
[0029] As a further solution of the present invention: Analyze the pixel changes of the key images in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction, including:
[0030] Based on multiple preset intervals, compare the consistency of different pixels in the key images at equal intervals along the preset direction;
[0031] Obtain a pixel repetition period according to the preset interval with the highest consistency.
[0032] As a further solution of the present invention: the multi-modal feature data further includes a preset feature frequency; obtain a target synthesized video, and perform multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, including:
[0033] Obtain a target synthesized video, a preset image feature, and a preset sound feature;
[0034] Detect the occurrence times of the preset image feature and the preset sound feature in the target synthesized video to obtain the preset feature frequency, which is used as the multi-modal feature data.
[0035] As a further solution of the present invention: perform domain recognition on the content of the target synthesized video according to the multi-modal feature vector to obtain content domain classification data of the target synthesized video, including:
[0036] Obtain a multi-modal feature vector;
[0037] Input the multi-modal feature vector into a preset neural network model to obtain the content domain classification data output by the preset neural network model.
[0038] A speech recognition system for a synthesized video, including:
[0039] A memory for storing programs;
[0040] A processor for executing the steps in any one of the above speech recognition methods for synthesized videos when executing the programs in the memory.
[0041] A computer-readable storage medium for storing computer-readable programs or instructions, which, when executed by a processor, can execute the steps in any one of the above speech recognition methods for synthesized videos.
[0042] The beneficial effects of the present invention:
[0043] The present invention provides a speech recognition method for synthetic videos. First, a target synthetic video is obtained, and multi-modal feature extraction is performed on the target synthetic video to obtain multi-modal feature data. Then, based on the multi-modal feature data, a multi-modal feature vector is established. After that, the content domain of the target synthetic video is identified according to the multi-modal feature vector to obtain content domain classification data of the target synthetic video. Finally, speech recognition is performed on the target synthetic video based on the content domain classification data to obtain a speech recognition result. Compared with the prior art, the invention identifies the professional field of the video by extracting multi-modal features from the synthetic video, and then optimizes speech recognition according to the specific field of the video to improve the accuracy of speech recognition in a specific professional field. It should be noted that when performing multi-modal feature extraction in the present invention, the edge smoothing index and the texture repetition index are used for analysis, making the most of the characteristics of the synthetic video itself, improving the accuracy of domain recognition, and thus reducing the error rate of speech recognition, solving the problem that the existing speech recognition technology has low accuracy when facing synthetic videos with overly specialized fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The present invention will be further described below with reference to the accompanying drawings.
[0045] Figure 1 is a schematic flowchart of the speech recognition method for synthetic videos of the present invention;
[0046] Figure 2 is a schematic structural diagram of the speech recognition system for synthetic videos of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Please refer to Figure 1 as shown, the present invention is a speech recognition method for synthetic videos, which is characterized by including the following steps:
[0049] S101. Obtain a target synthetic video, and perform multi-modal feature extraction on the target synthetic video to obtain multi-modal feature data, where the multi-modal feature data includes an edge smoothing index and a texture repetition index. The edge smoothing index is used to characterize the smoothness of the edges of the picture content in the target synthetic video, and the texture repetition index is used to characterize the repetition degree of the texture of the picture content in the target synthetic video;
[0050] S102. Establish a multi-modal feature vector based on the multi-modal feature data;
[0051] S103. Perform domain recognition on the target synthesized video content according to the multi-modal feature vector to obtain the content domain classification data of the target synthesized video;
[0052] S104. Perform speech recognition on the target synthesized video based on the content domain classification data to obtain the speech recognition result.
[0053] The invention identifies the professional field of the video by extracting multi-modal features from the synthesized video, and then optimizes the speech recognition according to the specific field of the video to improve the accuracy of speech recognition in a specific professional field. It should be noted that when the invention performs multi-modal feature extraction, it uses the edge smoothing index and the texture repetition index for analysis, making the most of the characteristics of the synthesized video itself, improving the accuracy of domain recognition, and thus reducing the error rate of speech recognition, solving the problem that the existing speech recognition technology has low accuracy when facing synthesized videos with overly professional fields.
[0054] It can be understood that the multi-modal feature extraction in the above process can be implemented by using any existing technology. Specifically, in a preferred embodiment, the multi-modal feature data further includes a preset feature frequency, where the preset feature refers to any feature that can be used to classify the field of the target synthesized video. On this basis, the above step S101, obtaining the target synthesized video and performing multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, specifically includes:
[0055] Obtain the target synthesized video, preset image features, and preset sound features;
[0056] Detect the occurrence times of the preset image features and preset sound features in the target synthesized video to obtain the preset feature frequency, which is used as the multi-modal feature data.
[0057] In the above process, the preset image features refer to image features that are helpful for video domain classification. For example, image features such as masks and syringes can be regarded as related to the medical field, features such as pens and books can be regarded as related to the education field, and ball objects can be regarded as related to the sports field. Similarly, the preset sound features refer to sound features that are helpful for video domain classification. For example, current sounds and keyboard tapping sounds can be regarded as related to the technology field, and engine roars and tire noises can be regarded as related to the automotive field.
[0058] By counting the occurrence frequencies of the above features, the field to which the target synthetic video belongs can be roughly determined. It can be understood that in practice, which specific features are used as the above prediction features and which specific method is used for feature extraction are all prior arts that can be understood by those skilled in the art, so no further description will be given in this article.
[0059] Furthermore, the present invention provides a more preferred embodiment. Specifically, multimodal feature data of an edge smoothing index and a texture repetition index are used. Its advantage lies in making the most of the characteristics of the synthetic video itself. Through the edge smoothing index and the texture repetition index, the image and content features in the target field can be captured more accurately. For example, in synthetic videos in the medical field, irregular stacked textures often appear with relatively smooth edges (such as CT scans and blood vessel models), while videos in the industrial field usually contain sharp edges of geometric shapes (such as gears and screws) and virtual backgrounds with high texture repetition and strong regularity (coordinate grids, etc.). Based on this characteristic, the present invention effectively fills the defects of traditional multimodal feature extraction techniques through the organic combination of the edge smoothing index and the texture repetition index, and more effectively extracts the unique feature characteristics of the synthetic video, realizing intelligent field classification based on video content understanding.
[0060] Specifically, in a preferred solution, on the basis of using multimodal feature data of an edge smoothing index and a texture repetition index, the above step S101, obtaining the target synthetic video and performing multimodal feature extraction on the target synthetic video to obtain multimodal feature data, specifically includes:
[0061] Obtaining a key image according to the target synthetic video;
[0062] Performing image feature extraction on the key image to obtain an edge smoothing index;
[0063] Performing image feature extraction on the key image to obtain a texture repetition index.
[0064] In the above process, the key image refers to the image in the target synthetic video used for feature extraction to obtain the edge smoothing index and the texture repetition index, which can be obtained by random selection or manual designation. It can be understood that the key image can be one or multiple. When there are multiple key images, the edge smoothing index and the texture repetition index can be obtained by summarizing multiple key images (for example, calculating the edge smoothing index and the texture repetition index of each key image respectively, and then taking the average as the multimodal feature data of the target synthetic video).
[0065] And in a preferred solution, the above step: obtaining a key image according to the target synthetic video, specifically includes the following steps:
[0066] Extract key video frames from the target synthesized video;
[0067] Perform object detection on the key video frames to obtain the position and size of the target object;
[0068] Crop the key video frames according to the position and size of the target object to obtain the target object image;
[0069] Both the key video frames and the target object images are used as key images.
[0070] In the above process, the target object refers to specific content objects such as people and objects in the target synthesized video. It can be understood that the number of objects contained in the video image will have a certain impact on the calculation results when calculating and analyzing the edge smoothness index and the texture repetition index. For example, when there are many and dense objects in the key image, the system may misjudge the edge of one object as the edge of another object, which may lead to the analysis results reflected by the obtained texture repetition index and edge smoothness index being more irregular than the actual texture distribution and the image edge being smoother and blurrier. Therefore, in this embodiment, on the basis of using the key video frames as key images, object detection is further performed on the key video frames. An image representing only the target object is separately extracted from the key video frames, that is, the target object image, and then redundant analysis is performed on these key images to obtain a more scientific and reasonable edge smoothness index and texture repetition index.
[0071] It can be understood that how to specifically calculate the edge smoothness index and the texture repetition index can be implemented by using any existing technology.
[0072] The present invention provides a preferred method. In a preferred solution, the above step: extracting image features from the key images to obtain the edge smoothness index specifically includes:
[0073] Obtain the image size of the key image;
[0074] Statistically calculate the pixel difference between adjacent pixels in the key image;
[0075] Obtain the edge smoothness index according to the image size and the pixel difference.
[0076] Specifically, the above process can be reflected by the following formula:
[0077]
[0078] Among them, ES represents the edge smoothness index, W represents the width of the key image, H represents the height of the key image, and ΔP represents the pixel difference between a pixel and its adjacent pixel.
[0079] The present invention provides a preferred method. In a preferred embodiment, the above steps of extracting image features from the key image to obtain the texture repetition index specifically include:
[0080] Analyze the pixel changes of the key image in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction;
[0081] Obtain the texture repetition index according to the pixel repetition periods corresponding to multiple preset directions.
[0082] In the above process, the pixel repetition period refers to the width of texture repetition in the image, which is used to characterize the degree of texture repetition in the image. The pixel repetition period can be narrowly understood as the distance between two identical pixels in the image. In a specific direction (i.e., the preset direction), the number of identical pixels obtained by counting this distance for pixels is the largest.
[0083] The present invention also provides a specific method for calculating the pixel repetition period. In this embodiment, the above steps of analyzing the pixel changes of the key image in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction specifically include:
[0084] Based on multiple preset intervals, compare the consistency of different pixels in the key image at equal intervals along the preset direction;
[0085] Obtain the pixel repetition period according to the preset interval with the highest consistency.
[0086] Specifically, the above process can be reflected by the following formula:
[0087]
[0088] Wherein, TR refers to the pixel repetition period, N refers to the preset normalization coefficient (which can be the width of the image in the preset direction), k refers to the preset interval, p i refers to the pixel value of the i-th pixel in the preset direction, and p i+nk refers to the pixel value of the pixel after the n-th preset interval after the i-th pixel in the preset direction. The meaning of the above formula is to calculate the sum of the differences between each pixel and the pixel after multiple different preset intervals along the preset direction, perform a cosine operation and sum, and normalize by taking the preset interval that can make the sum of the cosine operation results the largest to obtain the pixel repetition period.
[0089] It can be understood that after obtaining the above multi-modal feature data, a multi-modal feature vector can be established according to the corresponding values. Then, any existing method (such as conditional judgment) can be used to perform domain classification based on the multi-modal feature vector.
[0090] Further, the present invention also provides a preferred solution, wherein, in step S103 above, domain recognition is performed on the target synthesized video content according to the multi-modal feature vector to obtain content domain classification data of the target synthesized video, which specifically includes:
[0091] Obtain the multi-modal feature vector;
[0092] Input the multi-modal feature vector into a preset neural network model to obtain content domain classification data output by the preset neural network model.
[0093] The advantage of this embodiment is that it can combine the powerful non-linear fitting ability of the neural network model, accurately capture the potential correlations of multi-dimensional information such as vision, texture, and frequency in the video, and significantly improve the accuracy and robustness of domain classification.
[0094] After obtaining the domain data, step S104 can be performed: perform speech recognition on the target synthesized video based on the content domain classification data to obtain a speech recognition result. For example, for a video in the medical field, the system can automatically load a medical-specific speech recognition model and preferentially determine medical-specific terms such as "myocardial infarction" and "heart rate monitoring" to improve the accuracy of speech recognition; similarly, in the financial field, the system can preferentially determine industry-specific terms such as "price-earnings ratio" and "K-line trend" to avoid misidentifying them as general terms, thereby significantly improving the performance and applicability of speech recognition in professional field scenarios.
[0095] Reference Figure 2 , which shows a schematic structural diagram of an electronic device provided by an embodiment of the present invention. In this embodiment, the electronic device includes:
[0096] A memory 210 for storing programs;
[0097] A processor 220 for executing the steps in any of the above speech recognition methods for synthesized videos when executing the programs in the memory.
[0098] This embodiment also provides a computer-readable storage medium, on which a speech recognition program for synthesized videos is stored. When the speech recognition program for synthesized videos is executed by a processor, the steps in the above embodiments can be implemented.
[0099] The present invention provides a speech recognition method for synthetic videos. First, a target synthetic video is obtained, and multi-modal feature extraction is performed on the target synthetic video to obtain multi-modal feature data. Then, based on the multi-modal feature data, a multi-modal feature vector is established. After that, the content domain of the target synthetic video is identified according to the multi-modal feature vector to obtain the content domain classification data of the target synthetic video. Finally, speech recognition is performed on the target synthetic video based on the content domain classification data to obtain the speech recognition result. Compared with the prior art, the invention identifies the professional field of the video by performing multi-modal feature extraction on the synthetic video, and then optimizes the speech recognition according to the specific field of the video to improve the accuracy of speech recognition in a specific professional field. It should be noted that when performing multi-modal feature extraction in the present invention, the edge smoothing index and the texture repetition index are used for analysis, making the most of the characteristics of the synthetic video itself, improving the accuracy of domain recognition, and thus reducing the error rate of speech recognition, solving the problem that the existing speech recognition technology has low accuracy when facing synthetic videos with overly professional fields.
[0100] The above has described in detail one embodiment of the present invention, but the described content is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the scope covered by the patent of the present invention.
Claims
1. A speech recognition method for synthetic video, characterized in that, It includes the following steps: Obtain a target synthesized video, and perform multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, where the multi-modal feature data includes an edge smoothness index and a texture repetition index. The edge smoothness index is used to characterize the smoothness of the edges of the picture content in the target synthesized video, and the texture repetition index is used to characterize the repetition degree of the texture of the picture content in the target synthesized video; Establish a multi-modal feature vector according to the multi-modal feature data; Perform domain recognition on the content of the target synthesized video according to the multi-modal feature vector to obtain content domain classification data of the target synthesized video; Perform speech recognition on the target synthesized video based on the content domain classification data to obtain a speech recognition result.
2. The speech recognition method for synthesizing a video according to claim 1, wherein Obtain a target synthesized video, and perform multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, including: Obtain key images according to the target synthesized video; Perform image feature extraction on the key images to obtain an edge smoothness index; Perform image feature extraction on the key images to obtain a texture repetition index.
3. The speech recognition method for synthesizing a video according to claim 2, wherein Obtain key images according to the target synthesized video, including: Extract key video frames from the target synthesized video; Perform object detection on the key video frames to obtain the position and size of the target object; Crop the key video frames according to the position and size of the target object to obtain target object images; Both the key video frames and the target object images are used as key images.
4. The speech recognition method for synthesizing a video according to claim 2, wherein Perform image feature extraction on the key images to obtain an edge smoothness index, including: Obtain the image size of the key images; Statistically analyze the pixel differences between adjacent pixels in the key images; Obtain the edge smoothness index according to the image size and the pixel differences.
5. The speech recognition method for synthesizing a video according to claim 2, wherein Perform image feature extraction on the key images to obtain a texture repetition index, including: Analyze the pixel changes in the key images in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction; Obtain the texture repetition index according to the pixel repetition periods corresponding to the multiple preset directions.
6. The speech recognition method for synthesizing video according to claim 5, characterized in that, Analyze the pixel changes in the key images in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction, including: Based on multiple preset intervals, compare the consistency of different pixels in the key images at equal intervals along the preset direction; Obtain the pixel repetition period according to the preset interval with the highest consistency.
7. The speech recognition method for synthesizing video according to claim 1, characterized in that, The multi-modal feature data further includes a preset feature frequency; Obtain a target synthesized video, and perform multi-modal feature extraction on the target synthesized video to obtain multi-modal feature data, including: Obtain the target synthesized video, preset image features, and preset sound features; Detect the occurrence times of the preset image features and the preset sound features in the target synthesized video to obtain the preset feature frequency as the multi-modal feature data.
8. The speech recognition method for synthesizing a video according to claim 1, wherein, Perform domain recognition on the content of the target synthesized video according to the multi-modal feature vector to obtain content domain classification data of the target synthesized video, including: Obtain the multi-modal feature vector; Input the multi-modal feature vector into a preset neural network model to obtain the content domain classification data output by the preset neural network model.
9. A speech recognition system for synthesizing videos, characterized in that, It includes: A memory for storing programs; A processor for executing the steps in any one of claims 1-8 for the speech recognition method of synthesized videos when executing the programs in the memory.
10. A computer-readable storage medium, characterized in that, For storing computer-readable programs or instructions, when the programs or instructions are executed by a processor, they can implement the steps in the speech recognition method for synthesizing video according to any one of claims 1-8.
Citation Information
Patent Citations
Video classification method and device, model training method and device, medium and electronic equipment
CN115311599A
Subtitle content display method and device, equipment, medium and program product
CN116962600A
Video classification model training method and system and machine generated video recognition method and system
CN117789086A
Speech recognition method of video data, server and storage medium
CN117953898A
Video processing method and device, equipment, storage medium and program product
CN119418704A