A speech recognition method, system and storage medium for synthesizing a video

By extracting multimodal features from synthetic videos and analyzing them using edge smoothness index and texture repetition index, the problem of low speech recognition accuracy in professional fields of synthetic videos is solved, and higher speech recognition accuracy is achieved.

CN120260544BActive Publication Date: 2026-04-21SHENZHEN SMART INSURANCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SMART INSURANCE TECH CO LTD
Filing Date
2025-03-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing speech recognition technology is not accurate enough when dealing with synthetic videos that are too specialized in a particular field.

Method used

Multimodal feature extraction is performed by acquiring target synthetic videos, and analysis is conducted using edge smoothness index and texture repetition index to establish multimodal feature vectors for domain recognition and speech recognition optimization.

Benefits of technology

It improves the accuracy of speech recognition in specific professional fields and reduces the error rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260544B_ABST
    Figure CN120260544B_ABST
Patent Text Reader

Abstract

This invention relates to the field of speech recognition technology, specifically disclosing a speech recognition method for synthesized videos. The method first acquires a target synthesized video and extracts multimodal features from it to obtain multimodal feature data. Then, based on this multimodal feature data, a multimodal feature vector is established. Next, the multimodal feature vector is used to identify the domain of the target synthesized video content, resulting in content domain classification data. Finally, speech recognition is performed on the target synthesized video based on this content domain classification data to obtain the speech recognition result. Compared to existing technologies, this invention identifies the professional domain of the synthesized video through multimodal feature extraction and then optimizes speech recognition based on the specific domain to improve accuracy in that domain. This solves the problem of low accuracy in existing speech recognition technologies when dealing with synthesized videos with highly specialized domains.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech recognition technology, and specifically to a speech recognition method, system, and storage medium for synthesized video. Background Technology

[0002] Composite video refers to dynamic video content generated by fusing still images, moving images, virtual scenes, or real data using technologies such as artificial intelligence. Composite video reduces production costs and allows for the unlimited expansion of creative content, finding widespread application in film and entertainment, education and training, and virtual broadcasting. Speech recognition is an indispensable component of composite video, playing a crucial role, especially in automated subtitle generation and voice-driven animation.

[0003] However, when the synthesized video involves highly specialized fields (such as medicine, law, and technical expertise), industry terminology and complex sentence structures can lead to recognition errors, thereby affecting the accuracy of the subtitles and the credibility of the video.

[0004] Therefore, there is a need for a speech recognition method for synthesized videos that can improve the accuracy of speech recognition in specialized fields. Summary of the Invention

[0005] The purpose of this invention is to provide a speech recognition method, system, and storage medium for synthesized video, and to solve the following technical problems:

[0006] Existing speech recognition technology is not very accurate when dealing with synthetic videos that are too specialized in a particular field.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] A speech recognition method for synthesized video includes the following steps:

[0009] The target synthesized video is acquired, and multimodal features are extracted from the target synthesized video to obtain multimodal feature data. The multimodal feature data includes edge smoothness index and texture repetition index. Edge smoothness index is used to characterize the smoothness of the edges of the image content in the target synthesized video, and texture repetition index is used to characterize the repetition of the texture of the image content in the target synthesized video.

[0010] Based on the multimodal feature data, establish multimodal feature vectors;

[0011] Domain identification is performed on the target synthesized video content based on multimodal feature vectors to obtain content domain classification data of the target synthesized video;

[0012] Speech recognition is performed on the target synthetic video based on content domain classification data to obtain speech recognition results.

[0013] As a further aspect of the present invention: acquiring a target synthesized video, and extracting multimodal features from the target synthesized video to obtain multimodal feature data, including:

[0014] Based on the target video, key images are obtained;

[0015] Image features are extracted from key images to obtain the edge smoothness index;

[0016] Image features are extracted from key images to obtain the texture repetition index.

[0017] As a further aspect of the present invention: obtaining key images based on the target synthesized video, including:

[0018] Extract key video frames from the target synthesized video;

[0019] Target detection is performed on key video frames to obtain the position and size of the target object;

[0020] The key video frames are cropped based on the location and size of the target object to obtain the target object image;

[0021] Both key video frames and target object images are used as key images.

[0022] As a further aspect of the present invention: image feature extraction is performed on key images to obtain an edge smoothness index, including:

[0023] Obtain the image dimensions of the key image;

[0024] Statistical analysis of pixel differences between adjacent pixels in key images;

[0025] The edge smoothness index is obtained based on the image size and pixel difference.

[0026] As a further aspect of the present invention: image feature extraction is performed on key images to obtain a texture repetition index, including:

[0027] The pixel changes of the key image in multiple preset directions are analyzed to obtain the pixel repetition period corresponding to each preset direction;

[0028] The texture repetition index is obtained based on the pixel repetition period corresponding to multiple preset directions.

[0029] As a further aspect of the present invention: analyzing the pixel changes of the key image in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction, including:

[0030] Based on multiple preset intervals, the consistency of different pixels in the key image is compared at equal intervals along a preset direction;

[0031] The pixel repetition period is obtained based on the preset interval with the highest consistency.

[0032] As a further aspect of the present invention: the multimodal feature data further includes preset feature frequencies; acquiring the target synthesized video, and performing multimodal feature extraction on the target synthesized video to obtain multimodal feature data, including:

[0033] Acquire the target synthesized video, preset image features, and preset sound features;

[0034] The frequency of occurrence of preset image features and preset sound features in the target synthetic video is detected to obtain the preset feature frequency, which is used as multimodal feature data.

[0035] As a further aspect of the present invention: Domain identification is performed on the target synthesized video content based on multimodal feature vectors to obtain content domain classification data for the target synthesized video, including:

[0036] Obtain multimodal feature vectors;

[0037] The multimodal feature vectors are input into a preset neural network model to obtain the content domain classification data output by the preset neural network model.

[0038] A speech recognition system for synthesized video, comprising:

[0039] Memory, used to store programs;

[0040] A processor for executing any of the steps in the above-described speech recognition method for synthesizing video while executing a program in memory.

[0041] A computer-readable storage medium for storing a computer-readable program or instructions, which, when executed by a processor, enable the steps of any of the above-described speech recognition methods for synthesizing video.

[0042] The beneficial effects of this invention are:

[0043] This invention provides a speech recognition method for synthesized videos. First, a target synthesized video is acquired, and multimodal feature extraction is performed on the video to obtain multimodal feature data. Then, a multimodal feature vector is established based on this data. Next, domain identification is performed on the target synthesized video content based on the multimodal feature vector, resulting in content domain classification data. Finally, speech recognition is performed on the target synthesized video based on this classification data to obtain the speech recognition result. Compared to existing technologies, this invention identifies the professional domain of the video through multimodal feature extraction and then optimizes speech recognition based on the specific domain to improve accuracy in that domain. Notably, this invention uses edge smoothness index and texture repetition index for analysis during multimodal feature extraction, maximizing the utilization of the synthesized video's inherent characteristics, improving the accuracy of domain identification, and thus reducing the error rate of speech recognition. This solves the problem of low accuracy in existing speech recognition technologies when dealing with highly specialized synthesized videos. Attached Figure Description

[0044] The invention will now be further described with reference to the accompanying drawings.

[0045] Figure 1 This is a flowchart illustrating the speech recognition method for synthesized video according to the present invention;

[0046] Figure 2 This is a schematic diagram of the speech recognition system for synthesized video according to the present invention. Detailed Implementation

[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] Please see Figure 1 As shown, the present invention is a speech recognition method for synthesized video, characterized by comprising the following steps:

[0049] S101. Obtain the target synthesized video and perform multimodal feature extraction on the target synthesized video to obtain multimodal feature data. The multimodal feature data includes edge smoothness index and texture repetition index. The edge smoothness index is used to characterize the smoothness of the edges of the image content in the target synthesized video, and the texture repetition index is used to characterize the repetition of the texture of the image content in the target synthesized video.

[0050] S102. Based on the multimodal feature data, establish a multimodal feature vector;

[0051] S103. Based on the multimodal feature vector, perform domain identification on the target synthesized video content to obtain the content domain classification data of the target synthesized video;

[0052] S104. Perform speech recognition on the target synthetic video based on content domain classification data to obtain speech recognition results.

[0053] This invention identifies the professional domain of a synthesized video by performing multimodal feature extraction. It then optimizes speech recognition based on this domain to improve accuracy in specific professional fields. Notably, the invention employs edge smoothness and texture repetition indices during multimodal feature extraction, maximizing the utilization of the synthesized video's inherent characteristics. This improves domain identification accuracy, reduces speech recognition error rates, and addresses the problem of low accuracy in existing speech recognition technologies when dealing with highly specialized synthesized videos.

[0054] It is understood that the multimodal feature extraction in the above process can be implemented using any existing technology. Specifically, in a preferred embodiment, the multimodal feature data further includes preset feature frequencies, where preset features refer to any features that can be used to classify the domain of the target synthesized video. Based on this, the above step S101, obtaining the target synthesized video and performing multimodal feature extraction on the target synthesized video to obtain multimodal feature data, specifically includes:

[0055] Acquire the target synthesized video, preset image features, and preset sound features;

[0056] The frequency of occurrence of preset image features and preset sound features in the target synthetic video is detected to obtain the preset feature frequency, which is used as multimodal feature data.

[0057] In the above process, the preset image features refer to image features that are helpful for classifying video domains. For example, image features such as masks and syringes can be considered related to the medical field, features such as pens and books can be considered related to the education field, and ball-shaped objects can be considered related to the sports field. Similarly, the preset sound features refer to sound features that are helpful for classifying video domains. For example, electrical noise and keyboard typing sounds can be considered related to the technology field, and engine roaring and tire noise can be considered related to the automotive field.

[0058] By statistically analyzing the frequency of occurrence of the aforementioned features, the domain of the target synthetic video can be roughly determined. It is understood that the specific features used as the predictive features and the methods employed for feature extraction are existing technologies that are readily understood by those skilled in the art, and therefore will not be elaborated upon in this paper.

[0059] Furthermore, this invention provides a more preferred embodiment, specifically employing multimodal feature data of edge smoothness index and texture repetition index. Its advantage lies in maximizing the utilization of the characteristics of the synthesized video itself. Through edge smoothness index and texture repetition index, image and content features in the target domain can be captured more accurately. For example, synthesized videos in the medical field often exhibit irregular, layered textures with relatively smooth edges (such as CT scans and vascular models), while videos in the industrial field typically contain sharp edges of geometric shapes (such as gears and screws) and virtual backgrounds (such as coordinate grids) with high texture repetition and strong regularity. Based on these characteristics, this invention effectively fills the gaps in traditional multimodal feature extraction techniques by organically combining edge smoothness index and texture repetition index, more effectively extracting the unique features of synthesized videos and realizing intelligent domain classification based on video content understanding.

[0060] Specifically, in a preferred embodiment, based on multimodal feature data using edge smoothness index and texture repetition index, the above step S101—acquiring the target synthesized video and extracting multimodal features from the target synthesized video to obtain multimodal feature data—specifically includes:

[0061] Based on the target video, key images are obtained;

[0062] Image features are extracted from key images to obtain the edge smoothness index;

[0063] Image features are extracted from key images to obtain the texture repetition index.

[0064] In the above process, key images refer to the images in the target synthetic video used for feature extraction to obtain the edge smoothness index and texture repetition index. These images can be obtained through random selection or manual specification. It is understood that there can be one or multiple key images. When there are multiple key images, the edge smoothness index and texture repetition index can be obtained by summing up the values ​​from multiple key images (e.g., calculating the edge smoothness index and texture repetition index of each key image separately, and then taking the average as the multimodal feature data of the target synthetic video).

[0065] In a preferred embodiment, the above steps—obtaining key images from the target synthesized video—specifically include the following steps:

[0066] Extract key video frames from the target synthesized video;

[0067] Target detection is performed on key video frames to obtain the position and size of the target object;

[0068] The key video frames are cropped based on the location and size of the target object to obtain the target object image;

[0069] Both key video frames and target object images are used as key images.

[0070] In the above process, the target object refers to the specific people, objects, and other content objects in the target synthesized video. It is understandable that the number of objects in the video image will affect the calculation results when calculating and analyzing the edge smoothness index and texture repetition index. For example, when there are many and densely packed objects in the key image, the system may misidentify one object as the edge of another, which may lead to the analysis results of the texture repetition index and edge smoothness index being more irregular and the image edges smoother and blurrier than the actual texture distribution. Therefore, this embodiment, based on using key video frames as key images, further performs target detection on the key video frames. Images representing only the target objects, i.e., target object images, are extracted separately from the key video frames. Then, redundant analysis is performed on these key images to obtain more scientifically reasonable edge smoothness indices and texture repetition indices.

[0071] Understandably, any existing technology can be used to specifically calculate the edge smoothness index and texture repeatability index.

[0072] The present invention provides a preferred embodiment in which the above steps, namely: extracting image features from the key image to obtain the edge smoothness index, specifically include:

[0073] Obtain the image dimensions of the key image;

[0074] Statistical analysis of pixel differences between adjacent pixels in key images;

[0075] The edge smoothness index is obtained based on the image size and pixel difference.

[0076] Specifically, the above process can be represented by the following formula:

[0077]

[0078] Where ES represents the edge smoothing index, W represents the width of the key image, H represents the height of the key image, and ΔP represents the pixel difference between a pixel and its neighboring pixel.

[0079] The present invention provides a preferred embodiment in which the above steps, namely: extracting image features from the key image to obtain the texture repetition index, specifically include:

[0080] The pixel changes of the key image in multiple preset directions are analyzed to obtain the pixel repetition period corresponding to each preset direction;

[0081] The texture repetition index is obtained based on the pixel repetition period corresponding to multiple preset directions.

[0082] In the above process, the pixel repetition period refers to the width of texture repetition in the image, used to characterize the degree of texture repetition in the image. The pixel repetition period can be narrowly understood as the distance between two identical pixels in the image. In a specific direction (i.e., a preset direction), this distance is used to count the pixels, and the number of identical pixels obtained is the largest.

[0083] This invention also provides a specific method for calculating pixel repetition period. In this embodiment, the above steps—analyzing pixel changes in a key image in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction—specifically include:

[0084] Based on multiple preset intervals, the consistency of different pixels in the key image is compared at equal intervals along a preset direction;

[0085] The pixel repetition period is obtained based on the preset interval with the highest consistency.

[0086] Specifically, the above process can be represented by the following formula:

[0087]

[0088] Where TR refers to the pixel repetition period, N refers to the preset normalization coefficient (which can be the width of the image in the preset direction), k refers to the preset interval, and p i It refers to the pixel value of the i-th pixel in the preset direction, p i+nk This refers to the pixel value after the i-th pixel and the n-th preset interval along a preset direction. The meaning of the above formula is to calculate the sum of the differences between each pixel and the pixels after multiple different preset intervals along the preset direction, perform a cosine operation and sum them, and then normalize by taking the preset interval that maximizes the sum of the cosine operation results to obtain the pixel repetition period.

[0089] Understandably, after obtaining the aforementioned multimodal feature data, a multimodal feature vector can be constructed based on the corresponding values. Then, any existing method (such as conditional judgment) can be used to perform domain classification based on the multimodal feature vector.

[0090] Furthermore, the present invention also provides a preferred embodiment, wherein step S103, performing domain identification on the target synthesized video content based on multimodal feature vectors to obtain content domain classification data of the target synthesized video, specifically includes:

[0091] Obtain multimodal feature vectors;

[0092] The multimodal feature vectors are input into a preset neural network model to obtain the content domain classification data output by the preset neural network model.

[0093] The advantage of this embodiment is that it can combine the powerful nonlinear fitting capability of neural network models to accurately capture the potential correlations of multi-dimensional information such as visual, texture and frequency in videos, and significantly improve the accuracy and robustness of domain classification.

[0094] After obtaining the domain data, step S104 can be performed to perform speech recognition on the target synthesized video based on the content domain classification data, and obtain the speech recognition results. For example, for videos in the medical field, the system can automatically load a medical-specific speech recognition model and prioritize the identification of medical terms such as "myocardial infarction" and "heart rate monitoring" to improve the accuracy of speech recognition. Similarly, in the financial field, the system can prioritize the identification of industry-specific terms such as "price-to-earnings ratio" and "K-line trend" to avoid misidentifying them as general terms, thereby significantly improving the performance and applicability of speech recognition in professional domain scenarios.

[0095] refer to Figure 2 The diagram illustrates a structural schematic of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device includes:

[0096] Memory 210 is used to store programs;

[0097] The processor 220 is configured to perform any of the steps in the above-described speech recognition method for synthesizing video when executing a program in memory.

[0098] This embodiment also provides a computer-readable storage medium storing a speech recognition program for synthesizing video, which, when executed by a processor, can implement the steps in the above embodiments.

[0099] This invention provides a speech recognition method for synthesized videos. First, a target synthesized video is acquired, and multimodal feature extraction is performed on the video to obtain multimodal feature data. Then, a multimodal feature vector is established based on this data. Next, domain identification is performed on the target synthesized video content based on the multimodal feature vector, resulting in content domain classification data. Finally, speech recognition is performed on the target synthesized video based on this classification data to obtain the speech recognition result. Compared to existing technologies, this invention identifies the professional domain of the video through multimodal feature extraction and then optimizes speech recognition based on the specific domain to improve accuracy in that domain. Notably, this invention uses edge smoothness index and texture repetition index for analysis during multimodal feature extraction, maximizing the utilization of the synthesized video's inherent characteristics, improving the accuracy of domain identification, and thus reducing the error rate of speech recognition. This solves the problem of low accuracy in existing speech recognition technologies when dealing with highly specialized synthesized videos.

[0100] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A speech recognition method for synthesized video, characterized in that, Includes the following steps: The target synthesized video is acquired, and multimodal features are extracted from the target synthesized video to obtain multimodal feature data. The multimodal feature data includes edge smoothness index and texture repetition index. Edge smoothness index is used to characterize the smoothness of the edges of the image content in the target synthesized video, and texture repetition index is used to characterize the repetition of the texture of the image content in the target synthesized video. Based on the multimodal feature data, establish multimodal feature vectors; Domain identification is performed on the target synthesized video content based on multimodal feature vectors to obtain content domain classification data of the target synthesized video; Speech recognition is performed on the target synthetic video based on content domain classification data to obtain speech recognition results; The target synthesized video is acquired, and multimodal feature extraction is performed on the target synthesized video to obtain multimodal feature data, including: Based on the target video, key images are obtained; Image features are extracted from key images to obtain the edge smoothness index; Image features are extracted from key images to obtain the texture repetition index; Based on the target synthetic video, key images are obtained, including: Extract key video frames from the target synthesized video; Target detection is performed on key video frames to obtain the position and size of the target object; The key video frames are cropped based on the location and size of the target object to obtain the target object image; Both key video frames and target object images are used as key images.

2. The speech recognition method for synthesizing video according to claim 1, characterized in that, Image feature extraction is performed on key images to obtain the edge smoothness index, including: Obtain the image dimensions of the key image; Statistical analysis of pixel differences between adjacent pixels in key images; The edge smoothness index is obtained based on the image size and pixel difference.

3. The speech recognition method for synthesized video according to claim 1, characterized in that, Image feature extraction is performed on key images to obtain the texture repetition index, including: The pixel changes of the key image in multiple preset directions are analyzed to obtain the pixel repetition period corresponding to each preset direction; The texture repetition index is obtained based on the pixel repetition period corresponding to multiple preset directions.

4. The speech recognition method for synthesized video according to claim 3, characterized in that, Analyze the pixel changes of the key image in multiple preset directions to obtain the pixel repetition period corresponding to each preset direction, including: Based on multiple preset intervals, the consistency of different pixels in the key image is compared at equal intervals along a preset direction; The pixel repetition period is obtained based on the preset interval with the highest consistency.

5. The speech recognition method for synthesizing video according to claim 1, characterized in that, Multimodal feature data also includes preset feature frequencies; The target synthesized video is acquired, and multimodal feature extraction is performed on the target synthesized video to obtain multimodal feature data, including: Acquire the target synthesized video, preset image features, and preset sound features; The frequency of occurrence of preset image features and preset sound features in the target synthetic video is detected to obtain the preset feature frequency, which is used as multimodal feature data.

6. The speech recognition method for synthesizing video according to claim 1, characterized in that, Domain identification is performed on the target synthesized video content based on multimodal feature vectors to obtain content domain classification data for the target synthesized video, including: Obtain multimodal feature vectors; The multimodal feature vectors are input into a preset neural network model to obtain the content domain classification data output by the preset neural network model.

7. A speech recognition system for synthesized video, characterized in that, include: Memory, used to store programs; A processor for performing the steps of the speech recognition method for synthesizing video according to any one of claims 1-6 when executing a program in memory.

8. A computer-readable storage medium, characterized in that, Used to store computer-readable programs or instructions, which, when executed by a processor, are capable of implementing the steps in the speech recognition method for synthesizing video according to any one of claims 1-6.

Citation Information

Patent Citations

  • Video classification method and device, model training method and device, medium and electronic equipment

    CN115311599A

  • Multimodal speech recognition for real-time video audio-based display indicia application

    US20170169827A1