Subtitle matching display method and system based on intelligent image processing
Through intelligent image processing technology, the pixel value and display size of subtitles are adjusted using lip and expression information, which solves the problem that subtitles cannot highlight the focus of the video and improves the viewing effect.
Patent Information
- Application Number
- CN202510076610.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In the prior art, the display of subtitles cannot highlight the focus of images or videos, resulting in poor viewing effects.
The lip and expression information of the target object is obtained through the image information processing model, and combined with the text information of the audio file, the pixel values and display size of the subtitles are determined to highlight the key text.
It matches the display effect of subtitles with the video content, highlights key text, facilitates viewers' understanding and attention, and improves the viewing experience.
Smart Images

Figure CN119992530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a subtitle matching display method and system based on intelligent image processing. Background Art
[0002] In the related art, CN116612365B relates to the field of image subtitle technology, and specifically discloses a method for generating image subtitles based on target detection and natural language processing, including: obtaining a subtitle image to be generated, and performing vector processing and target detection on the subtitle image to be generated to obtain two sets of identical vector image features; inputting one set of vector image features into the encoder for feature extraction processing to obtain image processing features; inputting another set of vector image features into the decoder to perform a first information interaction with the image description text to obtain a first interaction result; inputting the image processing features into the decoder to perform a second information interaction with the first interaction result to obtain a second interaction result; converting the second interaction result to obtain image subtitles, and outputting the image subtitles. The image subtitle generation method based on target detection and natural language processing provided by the scheme solves the problem of deviation between image subtitles and the actual content expression of the image.
[0003] CN116665012B relates to the technical field of natural language processing, and specifically discloses an automatic generation method of image subtitles, an automatic generation device of image subtitles, and a computer storage medium, including: obtaining a subtitle image to be generated, and processing the subtitle image to be generated to obtain vector image features; inputting the vector image features into an encoder to construct prior knowledge and obtain effective image features; inputting the effective image features into a decoder to allow multimodal interaction between the effective image features and the image description text to obtain an interaction result; generating a text sequence according to the interaction result; converting the text sequence to obtain image subtitles, and outputting the image subtitles. The automatic generation method of image subtitles provided by the scheme can reduce the deviation between image subtitles and the actual content expression of the image.
[0004] CN117115516A discloses a multimodal sentiment analysis method and system integrating image captions and BERT, which relates to the technical field of multimodal sentiment analysis, including extracting images from multimodal datasets to generate text describing the images; obtaining image features through ResNet, using BERT encoder to calculate the hidden layer representation of the target, and obtaining the final visual representation based on the target image matching layer; inputting the image description and text of the multimodal data into the BERT encoder to calculate the feature representation of the image description and the feature representation of the text, and obtaining the final text feature representation; calculating the multimodal hidden layer representation, and obtaining the final sentiment classification through the pooling layer, fully connected layer and Softmax. The scheme extracts images from multimodal datasets, combines more than two modalities to realize cross-modal sentiment analysis, increases the ability to understand and recognize image content, effectively solves the limitations of single modality, converts image information into a more expressive and semantically informative visual representation, and improves the stability of the multimodal system.
[0005] Therefore, the related technology can improve the matching degree between the semantics of subtitles and images or videos and improve the accuracy of subtitles. However, the related technology does not set the display effect of subtitles, resulting in that the display of subtitles cannot highlight the key points of the image or video.
[0006] The information disclosed in the background technology section of this application is only intended to deepen the understanding of the general background technology of this application, and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to those skilled in the art. Summary of the invention
[0007] The present invention provides a subtitle matching display method and system based on intelligent image processing, which can solve the technical problem in the related art that the display of subtitles cannot highlight the key points of an image or video.
[0008] According to a first aspect of the present invention, there is provided a subtitle matching display method based on intelligent image processing, comprising:
[0009] Parsing the video to be processed to obtain multiple video images;
[0010] Processing the target object region in the video image through an image information processing model to obtain the lip shape information and expression information of the target object;
[0011] The audio file corresponding to the video to be processed is processed through the text recognition model to determine the text information corresponding to the audio file;
[0012] Processing the audio file corresponding to the video to be processed, obtaining an audio sub-file corresponding to the pronunciation of each text in the text information, and determining a video image corresponding to each audio sub-file;
[0013] Determine the pixel value of the text in the subtitle corresponding to the audio subfile according to the audio subfile, the lip shape information and the expression information of the target object in the video image;
[0014] Determining the display size of text in the subtitles corresponding to the audio subfile according to the expression information of the target object in the audio subfile and the video image;
[0015] Obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0016] According to the display information of each text, the subtitles of the video to be processed are obtained.
[0017] According to a second aspect of the present invention, there is provided a subtitle matching display system based on intelligent image processing, comprising:
[0018] A parsing module, used for parsing the video to be processed to obtain multiple video images;
[0019] An information module, used to process the area where the target object is located in the video image through an image information processing model to obtain the lip shape information and expression information of the target object;
[0020] A text module is used to process the audio file corresponding to the video to be processed through a text recognition model to determine the text information corresponding to the audio file;
[0021] The video image module is used to process the audio file corresponding to the video to be processed, obtain the audio sub-file corresponding to the pronunciation of each text in the text information, and determine the video image corresponding to each audio sub-file;
[0022] A pixel value module, used to determine the pixel value of the text corresponding to the audio subfile in the subtitle according to the audio subfile, the lip shape information and the expression information of the target object in the video image;
[0023] A display size module, used to determine the display size of the text in the subtitle corresponding to the audio subfile according to the expression information of the target object in the audio subfile and the video image;
[0024] A display information module, used for obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0025] The subtitle module is used to obtain the subtitles of the video to be processed according to the display information of each text.
[0026] According to a third aspect of the present invention, there is provided a subtitle matching display device based on intelligent image processing, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the subtitle matching display method based on intelligent image processing.
[0027] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the subtitle matching display method based on intelligent image processing.
[0028] By adopting the above technical solution, the present invention can achieve the following technical effects:
[0029] According to the present invention, the lip shape information and expression information of the target object can be obtained through the image information processing model, so as to determine the key points in the text information of the subtitle based on the lip shape information and expression information, so as to set a specific pixel value and display size for the subtitle to highlight the key text in the subtitle, so as to facilitate the viewer to watch and understand, and improve the display effect. When training the audio processing model, the lip shape pronunciation prediction model and the image information processing model, the competitive training of the audio processing model, the lip shape pronunciation prediction model and the image information processing model can be realized through the conditional function, and the accuracy of the audio processing model is continuously improved, and the supervision standard of the lip shape pronunciation prediction model and the image information processing model is improved, so that the accuracy of the lip shape pronunciation prediction model and the image information processing model is close to and reaches the accuracy of the audio processing model, so as to achieve the common improvement of the accuracy of the three models, and improve the training accuracy and efficiency. When determining the pixel adjustment coefficient, the frequency spectrum information of the audio clip, the difference between the text pronunciation information and the lip shape pronunciation information, and the expression information can be used to determine the possibility of encountering a situation that needs to be focused on, so as to obtain the pixel adjustment coefficient, which can objectively and comprehensively reflect the possibility of encountering a situation that needs to be focused on, and improve the accuracy of adjusting the pixel value. When determining the size adjustment coefficient, the volume information and expression information can be used to determine whether the target object has its volume amplified, and the size adjustment coefficient can be determined through a conditional function, so that when the volume is amplified, the size of the text in the subtitles is enlarged, and when the volume is not amplified, the size of the text is maintained, thereby comprehensively and objectively reflecting the relationship between the display size of the text in the subtitles and the volume, and effectively prompting the viewer's attention.
[0030] It should be understood that the above general description and the following detailed description are exemplary and explanatory only and do not limit the present invention. Other features and aspects of the present invention will become more apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other embodiments can be obtained based on these drawings without creative work.
[0032] Figure 1 A schematic diagram exemplarily shows a flow chart of a subtitle matching and display method based on intelligent image processing according to an embodiment of the present invention;
[0033] Figure 2 A block diagram of a subtitle matching and display system based on intelligent image processing according to an embodiment of the present invention is exemplarily shown. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0036] Figure 1 A schematic flow chart of a subtitle matching and display method based on intelligent image processing according to an embodiment of the present invention is exemplarily shown, and the method includes:
[0037] Step S101, parsing the video to be processed to obtain multiple video images;
[0038] Step S102, processing the area where the target object in the video image is located by using an image information processing model to obtain the lip shape information and expression information of the target object;
[0039] Step S103, processing the audio file corresponding to the video to be processed by a text recognition model to determine the text information corresponding to the audio file;
[0040] Step S104, processing the audio file corresponding to the video to be processed, obtaining an audio sub-file corresponding to the pronunciation of each text in the text information, and determining a video image corresponding to each audio sub-file;
[0041] Step S105, determining the pixel value of the text in the subtitle corresponding to the audio subfile according to the audio subfile, the lip shape information and the expression information of the target object in the video image;
[0042] Step S106, determining the display size of the text in the subtitle corresponding to the audio subfile according to the expression information of the target object in the audio subfile and the video image;
[0043] Step S107, obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0044] Step S108, obtaining subtitles of the video to be processed according to the display information of each text.
[0045] According to the subtitle matching display method based on intelligent image processing of an embodiment of the present invention, the lip shape information and expression information of the target object can be obtained through the image information processing model, so as to determine the key points in the text information of the subtitles based on the lip shape information and expression information, and thus set specific pixel values and display sizes for the subtitles to highlight the key text in the subtitles, so as to facilitate viewers to watch and understand and improve the display effect.
[0046] According to an embodiment of the present invention, in step S101, the video to be processed may be parsed, and the video to be processed may be a music video, a speech video, a course video, etc., and the video screen includes a target object that makes a sound, such as a singing target object, a speech target object, a lecture target object, etc., and the video to be processed has a corresponding audio file. The video to be processed may be parsed to obtain multiple video images, that is, multiple video frames.
[0047] According to one embodiment of the present invention, in step S102, the image information processing model is a deep learning neural network model, such as a convolutional neural network model, which can process the area where the target object in the video image is located to obtain the lip shape information and expression information of the target object, and the lip shape information and expression information are both information in the form of vectors. Among them, each component in the vector of the lip shape information can represent the relative position relationship between multiple key points on the lips, such as the distance and angle between multiple key points, and the lip shape of the target object can be reflected in combination with the relative position relationship between multiple key points. The multiple components of the vector of the expression information can represent the probability that the expression of the target object belongs to multiple expression types, such as the probability of belonging to a serious expression, the probability of belonging to an excited expression, the probability of belonging to an angry expression, the probability of belonging to a plain expression, etc.
[0048] According to one embodiment of the present invention, in step S103, the audio file corresponding to the video to be processed can be processed by a text recognition model to determine the text information corresponding to the audio file. The text recognition model is a deep learning neural network model, such as a recursive neural network model. The audio file can be converted into text information.
[0049] According to an embodiment of the present invention, in step S104, the audio file can be segmented to obtain audio sub-files corresponding to the pronunciation of each text, for example, a whole audio segment is segmented into multiple audio sub-files, each audio sub-file only includes the pronunciation of one word. Further, the video image has a timestamp, and the timestamp corresponds to the moment in the audio file. Therefore, the start time and end time of the audio sub-file can be used to find the video image with the corresponding timestamp as the video image corresponding to the start time and end time of the audio sub-file, so that the multiple video images between the two video images are determined as the video images corresponding to the audio sub-file.
[0050] According to an embodiment of the present invention, in step S105, the pixel value of the text corresponding to the audio subfile in the subtitle is determined according to the lip shape information and expression information of the target object in the audio subfile and the video image, including: determining the timestamps of multiple video images corresponding to the audio subfile; determining the audio segment in the time period between the timestamp of the i-th video image and the timestamp of the (i+1)-th video image in the audio subfile according to the timestamp; determining the text pronunciation information of the audio segment through an audio processing model, wherein the text pronunciation information is used to represent the pronunciation characteristics of the text corresponding to the audio segment; determining the lip shape pronunciation information of the lip shape information through a lip shape pronunciation prediction model, wherein the lip shape pronunciation information is used to represent the pronunciation characteristics of the sound that can be emitted based on the lip shape information; performing spectrum analysis on the audio segment to determine the spectrum information of the audio segment; and determining the pixel value of the text corresponding to the audio subfile in the subtitle of the i-th video image according to the text pronunciation information, the lip shape pronunciation information, the spectrum information and the expression information.
[0051] According to one embodiment of the present invention, each audio sub-file includes only the pronunciation of one word, but the pronunciation duration of each word is different, which makes the duration of each audio sub-file different, so the number of video images corresponding to the audio sub-files is different. And within the pronunciation duration of a word, the volume, tone and other parameters of the word may also change, so the audio sub-file can be divided again. In order to facilitate the combined analysis with the video image, the timestamp of the video image can be used as the segmentation point of the audio sub-file, and the audio sub-file can be divided into multiple audio segments, so that the audio segment in the time period between the timestamp of the i-th video image and the timestamp of the i+1-th video image can be analyzed in combination with the i-th video image.
[0052] According to one embodiment of the present invention, the audio processing model may be a BP neural network model, which may analyze the audio segment to obtain text pronunciation information of the audio segment. The text pronunciation information is information in vector form, which may be used to represent the pronunciation features of the text corresponding to the audio segment. For example, the multiple components of the text pronunciation information may include parameters representing the pronunciation type of the text (for example, the pronunciation is "a", "o", "e", etc.), pronunciation pitch (for example, the pronunciation pitch can be described by features such as the frequency of the pronunciation), volume and other information.
[0053] According to one embodiment of the present invention, the lip shape pronunciation prediction model can be a BP neural network model, which can process the lip shape information of the video image to predict the pronunciation characteristics of the sound that can be emitted by the lip shape information, and obtain the lip shape pronunciation information. The lip shape pronunciation information is of the same type as the text pronunciation information, both are information in vector form, and both can include parameters representing pronunciation type, pronunciation pitch, volume and other information.
[0054] According to an embodiment of the present invention, spectrum analysis may include obtaining spectrum information of the audio segment by Fourier transform or other methods, and obtaining frequency characteristics of the audio segment from the spectrum information. Furthermore, the pixel value of the text corresponding to the audio subfile may be determined based on the text pronunciation information, the lip pronunciation information, the spectrum information and the expression information.
[0055] According to an embodiment of the present invention, the above-mentioned mouth shape pronunciation prediction model can be trained before use, and accurate mouth shape pronunciation information can be obtained after training, and then the above-mentioned process of determining the pixel value of the text can be performed. The training step of the lip pronunciation prediction model includes: obtaining first sample videos of multiple test persons performing multiple pronunciations; obtaining multiple first images of the first sample video, and obtaining first sample lip information of the test persons in the first image through an image information processing model; segmenting the audio file corresponding to the first sample video according to the timestamp of the first image to obtain sample audio clips, and processing the sample audio clips through an audio processing model to obtain sample text pronunciation information; processing the first sample lip information according to the lip pronunciation prediction model to obtain sample lip pronunciation information; performing spectrum analysis on the sample audio clip to obtain first sample spectrum information; obtaining reference pronunciation information according to the first sample spectrum information; determining a first comprehensive loss function of the lip pronunciation prediction model, the audio processing model and the image information processing model according to the sample text pronunciation information, the sample lip pronunciation information and the reference pronunciation information; training the lip pronunciation prediction model, the audio processing model and the image information processing model according to the first comprehensive loss function to obtain a trained lip pronunciation prediction model, a trained audio processing model and a trained image information processing model.
[0056] According to one embodiment of the present invention, scenes in which multiple test subjects pronounce words can be filmed to obtain first sample videos in which multiple test subjects perform multiple pronunciations. The first sample videos can be analyzed to obtain multiple first images, i.e., video frames of the first sample videos. The first sample lip shape information of the test subjects in the first image can also be obtained through the above-mentioned image information processing model. The method of obtaining the information is the same as the method of obtaining the above-mentioned lip shape information, which will not be repeated here.
[0057] According to one embodiment of the present invention, the audio file corresponding to the first sample video is segmented by the timestamp of the first image to obtain sample audio segments between the timestamps of adjacent first images, for example, the sample audio segment between the timestamp of the ith first image and the timestamp of the (i+1)th first image. The sample audio segment can be processed by the audio processing model to obtain sample text pronunciation information, and the processing process is the same as the process of obtaining the text pronunciation information above, which will not be repeated here.
[0058] According to one embodiment of the present invention, the first sample mouth information can be processed by the mouth shape pronunciation prediction model to obtain the sample mouth shape pronunciation information. The sample text pronunciation information and the sample mouth shape pronunciation information are all information obtained by the untrained model (audio processing model and mouth shape pronunciation prediction model), and there may be errors. Therefore, by performing spectrum analysis on the sample audio segment, and obtaining error-free reference pronunciation information based on the obtained first sample spectrum information, the error between the sample text pronunciation information and the sample mouth shape pronunciation information and the reference pronunciation information can be determined. The reference pronunciation information is information in the form of vectors, and may include parameters for describing the pronunciation type, pronunciation tone, and volume of the text of the sample audio segment, that is, the pronunciation type can be manually marked, and the frequency characteristics of the pronunciation can be determined by spectrum analysis, and then the pronunciation tone can be described. The volume can also be determined by the analysis of sound waves. The parameters obtained in these ways are accurate parameters, so the reference pronunciation information is error-free vector information.
[0059] According to an embodiment of the present invention, according to the sample text pronunciation information, the sample mouth shape pronunciation information and the reference pronunciation information, determining a first comprehensive loss function of the mouth shape pronunciation prediction model, the audio processing model and the image information processing model, including: determining a first comprehensive loss function LOSS1 of the mouth shape pronunciation prediction model, the audio processing model and the image information processing model according to formula (1),
[0060]
[0061] Among them, P j,k,text is the sample text pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test subject, P j,k,ms The sample mouth shape pronunciation information corresponding to the first sample mouth shape information of the kth first image of the jth tester, P j,k,R is the reference pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test subject, (P j,k,text )T P j,k,text , w1 and w2 are preset weights, if is a conditional function, m is the number of first images in the first sample video, n is the number of test personnel, k≤m, j≤n, and k, j, m and n are all positive integers.
[0062] According to one embodiment of the present invention, in formula (1), the conditional function if{|P j,k,text -P j,k,R |<|P j,k,ms -P j,k,R |, w1|P j,k,text-P j,k,R |+w2|P j,k,ms -P j,k,R |} means in |P j,k,text -P j,k,R |<|P j,k,ms -P j,k,R |, the conditional function value is Otherwise, the conditional function value is w1|P j,k,text -P j,k,R |+w2|P j,k,ms -P j,k,R |. |P j,k,text -P j,k,R |<|P j,k,ms -P j,k,R | indicates that the error between the sample text pronunciation information and the reference pronunciation information is smaller than the error between the sample lip pronunciation information and the reference pronunciation information, wherein the sample text pronunciation information is obtained through an audio processing model, and the sample lip pronunciation information is obtained through an image information processing model and a lip pronunciation prediction model. Therefore, this condition of the conditional function also indicates that the error of the audio processing model is smaller than the error of the image information processing model and the lip pronunciation prediction model.
[0063] According to one embodiment of the present invention, based on the above analysis, when the condition of the conditional function is met, the error between the image information processing model and the mouth shape pronunciation prediction model is large, and the two models can be trained based on the error between the sample mouth shape pronunciation information and the reference pronunciation information. In addition, due to the large error, the model can also be used. As the denominator, the error is amplified, thereby increasing the training intensity of the image information processing model and the lip pronunciation prediction model, so that the error between the two can be quickly reduced. is the cosine similarity between the sample mouth shape pronunciation information and the sample text pronunciation information. Using it as the denominator can not only amplify the above error and improve the training intensity, but also can be used in the training process. The overall image is reduced, thereby improving the cosine similarity, thereby improving the consistency between the sample mouth pronunciation information and the sample text pronunciation information, and making the results obtained by the image information processing model and the mouth pronunciation prediction model close to the results obtained by the audio processing model, that is, making the accuracy of the image information processing model and the mouth pronunciation prediction model close to the accuracy of the audio processing model, thereby improving the accuracy of the image information processing model and the mouth pronunciation prediction model.
[0064] According to one embodiment of the present invention, if the condition of the conditional function is not met, that is, the error of the audio processing model is large and the accuracy is low, the weighted sum of the error between the sample text pronunciation information and the reference pronunciation information, and the error between the sample mouth shape pronunciation information and the reference pronunciation information is used as the conditional function value, so that the mouth shape pronunciation prediction model, the audio processing model and the image information processing model are trained respectively through the two errors, that is, the audio processing model is trained through the error between the sample text pronunciation information and the reference pronunciation information, and the image information processing model and the mouth shape pronunciation prediction model are trained through the error between the sample mouth shape pronunciation information and the reference pronunciation information, thereby improving the accuracy of the three models together.
[0065] Therefore, the above conditional function can realize the competitive training of the audio processing model, the lip pronunciation prediction model and the image information processing model, that is, the accuracy of the audio processing model is used to supervise the training of the lip pronunciation prediction model and the image information processing model. When the accuracy of the audio processing model is low, the accuracy of the three models can be improved at the same time. As the training progresses, if the accuracy of the audio processing model exceeds the accuracy of the lip pronunciation prediction model and the image information processing model, the training intensity of the lip pronunciation prediction model and the image information processing model is improved, so that the accuracy of the lip pronunciation prediction model and the image information processing model is close to and reaches the accuracy of the audio processing model. Moreover, as the accuracy of the audio processing model is improved, a higher standard can be used to supervise the training of the lip pronunciation prediction model and the image information processing model, so as to quickly improve the three models and improve the training efficiency.
[0066] According to one embodiment of the present invention, the conditional function values corresponding to each first image of each test person can be summed to obtain a first comprehensive loss function, and the gradient descent method can be used for feedback propagation to adjust the parameters of the above three models. After multiple trainings, the training can be completed to obtain a trained lip pronunciation prediction model, a trained audio processing model and a trained image information processing model.
[0067] In this way, competitive training of the audio processing model, the lip pronunciation prediction model and the image information processing model can be achieved through conditional functions. By continuously improving the accuracy of the audio processing model, the supervision standard of the lip pronunciation prediction model and the image information processing model can be improved, so that the accuracy of the lip pronunciation prediction model and the image information processing model can be close to and reach the accuracy of the audio processing model, so as to achieve a joint improvement in the accuracy of the three models and improve the training accuracy and efficiency.
[0068] According to one embodiment of the present invention, the lip shape pronunciation information obtained by processing the lip shape information of the video image based on the lip shape pronunciation prediction model trained in the above manner can be used to represent the pronunciation characteristics of the sound that the lip shape can theoretically emit. The pronunciation characteristics are theoretically consistent with the pronunciation characteristics described by the text pronunciation characteristics, but if encountering special circumstances, such as emotional speeches, difficult singing, etc., it may cause deformation of facial expressions and lip shapes, making the lip shape pronunciation information inconsistent with the text pronunciation information. In other words, due to the deformation of facial expressions and lip shapes, the lip shape of the text corresponding to the audio sub-file deviates from the lip shape during normal pronunciation. Therefore, when the difference between the lip shape pronunciation information and the text pronunciation information is large, there may be special circumstances, and the special circumstances are the key circumstances in the video to be processed, such as the key parts of the speech or singing. On the other hand, the spectral information of the audio clip can also reflect to a certain extent whether there are special circumstances. For example, when a speech or singing is relatively dull, the voice is relatively low and the frequency of the sound waves is relatively low. When the speaker is emotional or the singer sings high notes, the frequency of the sound waves is higher. Therefore, the frequency of the sound waves can be analyzed through the spectral information to determine whether there are special circumstances that require special attention, so that more eye-catching pixel values can be set for the subtitles to prompt the viewer to pay attention.
[0069] According to an embodiment of the present invention, determining the pixel value of the text corresponding to the audio subfile in the subtitle of the i-th video image according to the text pronunciation information, the lip pronunciation information, the spectrum information and the expression information includes: obtaining the maximum pronunciation frequency of the audio segment according to the spectrum information; determining the pixel adjustment coefficient εi of the text corresponding to the audio subfile in the subtitle of the i-th video image according to formula (2): ,
[0070]
[0071] Among them, f i,max is the maximum pronunciation frequency of the audio segment in the time period between the i-th video image and the i+1-th video image, f p P is the preset pronunciation frequency. i,text is the text pronunciation information of the audio segment in the time period between the i-th video image and the i+1-th video image, P i,ms is the lip shape pronunciation information of the ith video image, (P i,text ) T P i,text The transposed vector, E i is the expression information of the i-th video image, E 1,p is the preset first expression information, (E i ) T For E i; determining the pixel value of the text corresponding to the audio sub-file in the subtitle of the i-th video image according to the pixel adjustment coefficient and the basic pixel value.
[0072] According to one embodiment of the present invention, in formula (2), the preset pronunciation frequency may be the average pronunciation frequency of the audio file corresponding to the video to be processed, It is the ratio of the maximum pronunciation frequency of the audio segment in the time period between the i-th video image and the i+1-th video image to the preset pronunciation frequency. The higher the ratio, the higher the possibility of encountering a situation that requires special attention, for example, the higher the possibility of encountering the high-pitched part of singing or the key part of a speech.
[0073] According to one embodiment of the present invention, in formula (2), in, is the cosine similarity between text pronunciation information and mouth pronunciation information, so, is the difference between the text pronunciation information and the mouth pronunciation information. As mentioned above, when the difference between the two is large, there is a high possibility of encountering a situation that requires special attention. Therefore, As the magnification factor, the higher the possibility of encountering a situation that requires special attention, the larger the magnification factor.
[0074] According to one embodiment of the present invention, in formula (2), the preset first expression information can indicate that the expression of the target object is a specific expression, for example, indicating that the expression of the target object is an excited expression. For example, among the components of the first expression information, the probability of a serious expression is 0, the probability of an excited expression is 1, the probability of an angry expression is 0, and the probability of a neutral expression is 0. It is the cosine similarity between the expression information of the video image and the preset first expression information. The higher the cosine similarity, the higher the possibility that the target object is in an excited state and the higher the possibility of encountering a situation that requires special attention. Therefore, the cosine similarity can be used as another amplification factor, so that the higher the possibility of encountering a situation that requires special attention, the larger the amplification factor.
[0075] According to an embodiment of the present invention, the above three items can be multiplied to obtain the pixel adjustment coefficient of the text corresponding to the audio subfile in the subtitle of the i-th video image, and the pixel value of the text corresponding to the audio subfile in the subtitle of the i-th video image can be obtained by multiplying the adjustment coefficient with the basic pixel value, that is, if the text is the key content, the pixel value of the text can be increased and the color of the text can be deepened to prompt the viewer to pay attention. In addition, since the audio subfile can correspond to multiple audio clips and multiple video images, each audio clip and video image can determine a pixel adjustment coefficient, and the audio subfile only corresponds to one text, therefore, the text may also produce a color change effect to prompt the viewer to pay attention. For example, during the singing process, a word in the lyrics is sung in different tones, then the word can correspond to multiple pixel adjustment coefficients during the singing period, so that the pixel value of the word can change during the video playback process, thereby prompting the viewer to pay attention.
[0076] In this way, the possibility of encountering a situation that requires special attention can be determined through the spectral information of the audio clip, the difference between the text pronunciation information and the lip pronunciation information, and the expression information, so as to obtain the pixel adjustment coefficient, which can objectively and comprehensively reflect the possibility of encountering a situation that requires special attention and improve the accuracy of adjusting the pixel value.
[0077] According to an embodiment of the present invention, in step S106, the display size of the text corresponding to the audio subfile can also be used to prompt the viewer to pay attention to the situation that needs to be focused on. According to the expression information of the target object in the audio subfile and the video image, the display size of the text corresponding to the audio subfile in the subtitle is determined, including: determining the maximum volume information of the audio segment; determining the average volume information of the audio file; setting the second expression information related to the pronunciation volume; according to the maximum volume information, the average volume information and the second expression information, determining the display size of the text corresponding to the audio subfile in the subtitle.
[0078] According to an embodiment of the present invention, determining the display size of the text corresponding to the audio subfile in the subtitle according to the maximum volume information, the average volume information and the second expression information includes: determining the size adjustment coefficient σ of the text corresponding to the audio subfile in the subtitle of the i-th video image according to formula (3): i ,
[0079]
[0080] Among them, V i,max is the maximum volume information of the audio clip in the time period between the i-th video image and the i+1-th video image, V ave is the average volume information, Ei is the expression information of the i-th video image, (E i ) T For E i The transposed vector, E 2,p is the second expression information, if is a conditional function; according to the size adjustment coefficient and the basic size value, the display size of the text corresponding to the audio subfile in the subtitle of the i-th video image is determined.
[0081] According to one embodiment of the present invention, in formula (3), the conditional function Indicated in In the case of Otherwise 1. It is the ratio of the maximum volume information to the average volume information of the audio clip. The larger the ratio is, the louder the target object's pronunciation volume of the text is, and the more likely it is to encounter a situation that requires special attention. On the other hand, the second expression information is similar to the first expression information, and can also be preset expression information. Moreover, the second expression information can be consistent with the first expression information, or it can also be other preset expression information. For example, among the components of the second expression information, the probability of a serious expression is 0, the probability of an excited expression is 0, the probability of an angry expression is 1, and the probability of a dull expression is 0. Then the second expression information can indicate that the target object's expression is an angry expression. Therefore, It can be used as a volume amplification factor, that is, the amplification factor that indicates the situation where the volume of the target object is amplified. If the result of the multiplication of the two is greater than or equal to 1, it means that the target object has amplified the volume. As a resizing factor, the size of the text in the subtitle is enlarged when the target object increases the volume. If the result of multiplying the two is less than 1, the resizing factor is equal to 1, that is, the size of the text will not be reduced.
[0082] According to one embodiment of the present invention, the display size of the text corresponding to the audio sub-file can be determined by multiplying the size adjustment value by the basic size value, so that when the target object that makes the sound increases the volume, the size of the text in the letters is enlarged. Moreover, since the audio sub-file corresponds to multiple video images and multiple audio clips, each video image and audio clip can obtain a size adjustment coefficient. Therefore, the size of the text corresponding to the audio sub-file is variable. For example, when the target object sings a word in the lyrics, its volume keeps changing, and the size of the word in the subtitles can also change, thereby increasing the viewer's attention.
[0083] In this way, it is possible to determine whether the target object should increase the volume through volume information and expression information, and determine the size adjustment coefficient through a conditional function, so that when the volume is increased, the size of the text in the subtitles is increased, and when the volume is not increased, the size of the text is maintained, thereby comprehensively and objectively reflecting the relationship between the display size of the text in the subtitles and the volume, and effectively prompting the viewer's attention.
[0084] According to one embodiment of the present invention, in step S107, the display information of the text corresponding to the audio sub-file in the video image can be obtained based on the pixel value of the text and the display size, that is, the color and size of the text at the moment corresponding to each video image are determined, so as to display the text in the video image.
[0085] According to an embodiment of the present invention, in step S108, based on the above method, the display information of each text can be determined to obtain the subtitles of the video to be processed, that is, the color and size of each text in each video image can be determined, so as to display the text in the video image. For example, if the 1st to 10th video images are the video images corresponding to the 1st text, the 1st text can be displayed based on the display information of the 1st to 10th video images, and the other texts keep the basic pixel value and basic size value; if the 11th to 20th video images are the video images corresponding to the 2nd text, the 2nd text can be displayed based on the display information of the 11th to 20th video images, and the other texts keep the basic pixel value and basic size value... Based on this method, appropriate subtitles can be displayed in the video to be processed, so as to prompt the viewer to pay attention to the key content.
[0086] According to the subtitle matching display method based on intelligent image processing of the embodiment of the present invention, the lip shape information and expression information of the target object can be obtained through the image information processing model, so as to determine the key points in the text information of the subtitle based on the lip shape information and expression information, so as to set a specific pixel value and display size for the subtitle to highlight the key text in the subtitle, so as to facilitate the viewer to watch and understand, and improve the display effect. When training the audio processing model, the lip shape pronunciation prediction model and the image information processing model, the competitive training of the audio processing model, the lip shape pronunciation prediction model and the image information processing model can be realized through the conditional function, and the accuracy of the audio processing model is continuously improved, and the supervision standard of the lip shape pronunciation prediction model and the image information processing model is improved, so that the accuracy of the lip shape pronunciation prediction model and the image information processing model is close to and reaches the accuracy of the audio processing model, so as to achieve the common improvement of the accuracy of the three models, and improve the training accuracy and efficiency. When determining the pixel adjustment coefficient, the frequency spectrum information of the audio clip, the difference between the text pronunciation information and the lip shape pronunciation information, and the expression information can be used to determine the possibility of encountering a situation that needs to be focused on, so as to obtain the pixel adjustment coefficient, which can objectively and comprehensively reflect the possibility of encountering a situation that needs to be focused on, and improve the accuracy of adjusting the pixel value. When determining the size adjustment coefficient, the volume information and expression information can be used to determine whether the target object has its volume amplified, and the size adjustment coefficient can be determined through a conditional function, so that when the volume is amplified, the size of the text in the subtitles is enlarged, and when the volume is not amplified, the size of the text is maintained, thereby comprehensively and objectively reflecting the relationship between the display size of the text in the subtitles and the volume, and effectively prompting the viewer's attention.
[0087] Figure 2 A block diagram of a subtitle matching and display system based on intelligent image processing according to an embodiment of the present invention is exemplarily shown, wherein the system comprises:
[0088] A parsing module, used for parsing the video to be processed to obtain multiple video images;
[0089] An information module, used to process the area where the target object is located in the video image through an image information processing model to obtain the lip shape information and expression information of the target object;
[0090] A text module is used to process the audio file corresponding to the video to be processed through a text recognition model to determine the text information corresponding to the audio file;
[0091] The video image module is used to process the audio file corresponding to the video to be processed, obtain the audio sub-file corresponding to the pronunciation of each text in the text information, and determine the video image corresponding to each audio sub-file;
[0092] A pixel value module, used to determine the pixel value of the text corresponding to the audio subfile in the subtitle according to the audio subfile, the lip shape information and the expression information of the target object in the video image;
[0093] A display size module, used to determine the display size of the text in the subtitle corresponding to the audio subfile according to the expression information of the target object in the audio subfile and the video image;
[0094] A display information module, used for obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0095] The subtitle module is used to obtain the subtitles of the video to be processed according to the display information of each text.
[0096] According to one embodiment of the present invention, a subtitle matching display device based on intelligent image processing is provided, comprising: a processor; a memory for storing processor executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the subtitle matching display method based on intelligent image processing.
[0097] According to an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the subtitle matching display method based on intelligent image processing is implemented.
[0098] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0099] It should be understood by those skilled in the art that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The purpose of the present invention has been fully and effectively achieved. The functional and structural principles of the present invention have been demonstrated and explained in the embodiments, and the embodiments of the present invention may be deformed or modified in any way without departing from the principles.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A subtitle matching display method based on intelligent image processing, characterized in that: include: Parsing the video to be processed to obtain multiple video images; Processing the target object region in the video image through an image information processing model to obtain the lip shape information and expression information of the target object; The audio file corresponding to the video to be processed is processed through the text recognition model to determine the text information corresponding to the audio file; Processing the audio file corresponding to the video to be processed, obtaining an audio sub-file corresponding to the pronunciation of each text in the text information, and determining a video image corresponding to each audio sub-file; Determine the pixel value of the text in the subtitle corresponding to the audio subfile according to the audio subfile, the lip shape information and the expression information of the target object in the video image; Determining the display size of text in the subtitles corresponding to the audio subfile according to the expression information of the target object in the audio subfile and the video image; Obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size; According to the display information of each text, the subtitles of the video to be processed are obtained.
2. The subtitle matching display method based on intelligent image processing according to claim 1 is characterized in that: Determining pixel values of text corresponding to the audio subfile in the subtitles according to the audio subfile, the lip shape information and the expression information of the target object in the video image, comprises: Determine timestamps of multiple video images corresponding to the audio sub-files; According to the timestamp, determine the audio segment in the time period between the timestamp of the i-th video image and the timestamp of the (i+1)-th video image in the audio subfile; Determining text pronunciation information of the audio segment through an audio processing model, wherein the text pronunciation information is used to represent pronunciation features of the text corresponding to the audio segment; Determining the mouth shape pronunciation information of the mouth shape information by using a mouth shape pronunciation prediction model, wherein the mouth shape pronunciation information is used to represent the pronunciation characteristics of the sound that can be produced based on the mouth shape information; Performing spectrum analysis on the audio clip to determine spectrum information of the audio clip; The pixel value of the text corresponding to the audio sub-file in the subtitle of the i-th video image is determined according to the text pronunciation information, the lip pronunciation information, the frequency spectrum information and the expression information.
3. The subtitle matching display method based on intelligent image processing according to claim 2 is characterized in that: The training steps of the mouth shape pronunciation prediction model include: Obtaining first sample videos of multiple test persons performing multiple pronunciations; Acquire multiple first images of the first sample video, and acquire first sample lip shape information of the test person in the first image through the image information processing model; Segmenting the audio file corresponding to the first sample video according to the timestamp of the first image to obtain sample audio segments, and processing the sample audio segments through an audio processing model to obtain sample text pronunciation information; Processing the first sample mouth shape information according to the mouth shape pronunciation prediction model to obtain sample mouth shape pronunciation information; Performing spectrum analysis on the sample audio segment to obtain first sample spectrum information; Obtaining reference pronunciation information according to the first sample spectrum information; Determine a first comprehensive loss function of a lip pronunciation prediction model, an audio processing model, and an image information processing model according to the sample text pronunciation information, the sample lip pronunciation information, and the reference pronunciation information; According to the first comprehensive loss function, the lip pronunciation prediction model, the audio processing model and the image information processing model are trained to obtain the trained lip pronunciation prediction model, the trained audio processing model and the trained image information processing model.
4. The subtitle matching display method based on intelligent image processing according to claim 3 is characterized in that: Determining a first comprehensive loss function of a lip pronunciation prediction model, an audio processing model, and an image information processing model according to the sample text pronunciation information, the sample mouth shape pronunciation information, and the reference pronunciation information, including: According to the formula Determine the first comprehensive loss function LOSS1 of the mouth shape pronunciation prediction model, the audio processing model and the image information processing model, where P j,k,text is the sample text pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test subject, P j,k,ms The sample mouth shape pronunciation information corresponding to the first sample mouth shape information of the kth first image of the jth tester, P j,k,R is the reference pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test person, (P j,k,text ) T P j,k,text , w1 and w2 are preset weights, if is a conditional function, m is the number of first images in the first sample video, n is the number of test personnel, k≤m, j≤n, and k, j, m and n are all positive integers.
5. The subtitle matching display method based on intelligent image processing according to claim 2 is characterized in that: Determining the pixel value of the text corresponding to the audio sub-file in the subtitle of the i-th video image according to the text pronunciation information, the lip pronunciation information, the spectrum information and the expression information, comprises: According to the spectrum information, obtain the maximum pronunciation frequency of the audio clip; According to the formula Determine the pixel adjustment coefficient ε of the text corresponding to the audio subfile in the subtitle of the i-th video image i , where f i,max is the maximum pronunciation frequency of the audio segment in the time period between the i-th video image and the i+1-th video image, f p P is the preset pronunciation frequency. i,text is the text pronunciation information of the audio segment in the time period between the i-th video image and the i+1-th video image, P i,ms is the lip shape pronunciation information of the ith video image, (P i,text ) T P i,text The transposed vector, E i is the expression information of the i-th video image, E 1,p is the preset first expression information, (E i ) T For E i The transposed vector of ; The pixel value of the text corresponding to the audio sub-file in the subtitle of the i-th video image is determined according to the pixel adjustment coefficient and the basic pixel value.
6. The subtitle matching display method based on intelligent image processing according to claim 2 is characterized in that: Determining the display size of text corresponding to the audio subfile in the subtitle according to the expression information of the target object in the audio subfile and the video image includes: Determine the maximum volume information of the audio clip; Determine the average volume information of an audio file; Setting the second expression information related to the pronunciation volume; The display size of the text corresponding to the audio subfile in the subtitle is determined according to the maximum volume information, the average volume information and the second expression information.
7. The subtitle matching display method based on intelligent image processing according to claim 6 is characterized in that: Determining a display size of text corresponding to the audio subfile in a subtitle according to the maximum volume information, the average volume information, and the second expression information includes: According to the formula Determine the size adjustment coefficient σ of the text corresponding to the audio subfile in the subtitle of the i-th video image i , where V i,max is the maximum volume information of the audio clip in the time period between the i-th video image and the i+1-th video image, V ave is the average volume information, E i is the expression information of the i-th video image, (E i ) T For E i The transposed vector, E 2,p is the second expression information, if is a conditional function; The display size of the text corresponding to the audio sub-file in the subtitle of the i-th video image is determined according to the size adjustment coefficient and the basic size value.
8. A subtitle matching display system based on intelligent image processing, characterized in that: include: A parsing module, used for parsing the video to be processed to obtain multiple video images; An information module, used to process the area where the target object is located in the video image through an image information processing model to obtain the lip shape information and expression information of the target object; A text module is used to process the audio file corresponding to the video to be processed through a text recognition model to determine the text information corresponding to the audio file; The video image module is used to process the audio file corresponding to the video to be processed, obtain the audio sub-file corresponding to the pronunciation of each text in the text information, and determine the video image corresponding to each audio sub-file; A pixel value module, used to determine the pixel value of the text corresponding to the audio subfile in the subtitle according to the audio subfile, the lip shape information and the expression information of the target object in the video image; A display size module, used to determine the display size of the text in the subtitle corresponding to the audio subfile according to the expression information of the target object in the audio subfile and the video image; A display information module, used for obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size; The subtitle module is used to obtain the subtitles of the video to be processed according to the display information of each text.
9. A subtitle matching display device based on intelligent image processing, characterized in that: include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that: Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-modal sentiment analysis method and system fusing image subtitles and BERT
CN117115516A
Learning video caption adding method and device, terminal equipment and storage medium
CN111639233A
Video display and processing method, device and system, equipment and medium
CN112579826A
Live video stream processing method and device, electronic equipment and storage medium
CN115086753A
Video display method and device, equipment and storage medium
CN116055792A
Cited By
Batch video generation method, electronic equipment, storage medium and product
CN120166267A