Subtitle matching display method and system based on intelligent image processing
Through the intelligent image processing model, the lip and expression information of the target object is obtained, combined with audio text information, the pixel value and display size of the subtitles are determined, which solves the problem of not outstanding subtitles display effect, realizes the highlighting of key text of subtitles, and improves the viewer's understanding effect.
Patent Information
- Application Number
- CN202510076610.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-01-17
AI Technical Summary
In the prior art, the display effect of subtitles cannot highlight the focus of images or videos, making it difficult for viewers to accurately understand the key content.
The lip and expression information of the target object is obtained through the image information processing model, and combined with the text information of the audio file, the pixel value and display size of the subtitles are determined to highlight the key text.
The highlighting of key text of subtitles is achieved, which improves the viewer's understanding effect, and improves the accuracy and efficiency of the display effect through model training.
Smart Images

Figure CN119992530B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a subtitle matching display method and system based on intelligent image processing. Background Art
[0002] In related technologies, CN116612365B relates to the field of image subtitle technology, specifically disclosing a method for generating image subtitles based on target detection and natural language processing, including: obtaining a subtitle image to be generated, and performing vector processing and target detection on the subtitle image to obtain two sets of identical vector image features; inputting one set of vector image features into an encoder for feature extraction processing to obtain image processing features; inputting another set of vector image features into a decoder to perform a first information interaction with the image description text to obtain a first interaction result; inputting the image processing features into a decoder to perform a second information interaction with the first interaction result to obtain a second interaction result; converting the second interaction result to obtain image subtitles, and outputting the image subtitles. The image subtitle generation method based on target detection and natural language processing provided by this solution solves the problem of deviation between image subtitles and the actual content expression of the image.
[0003] CN116665012B relates to the field of natural language processing technology and specifically discloses a method for automatically generating image subtitles, an image subtitle automatic generation device, and a computer storage medium. The method comprises: obtaining an image to be subtitled and processing the image to be subtitled to obtain vector image features; inputting the vector image features into an encoder to construct prior knowledge and obtain effective image features; inputting the effective image features into a decoder to enable multimodal interaction between the effective image features and image description text to obtain interaction results; generating a text sequence based on the interaction results; converting the text sequence to obtain image subtitles, and outputting the image subtitles. The method for automatically generating image subtitles provided by this solution can reduce the deviation between image subtitles and the actual content of the image.
[0004] CN117115516A discloses a multimodal sentiment analysis method and system that integrates image captioning and BERT. This relates to the technical field of multimodal sentiment analysis and includes extracting images from a multimodal dataset to generate text describing the image; obtaining image features through ResNet, using a BERT encoder to calculate the target's hidden layer representation, and obtaining the final visual representation based on the target image matching layer; inputting the image description and text of the multimodal data into the BERT encoder to calculate the feature representation of the image description and the feature representation of the text, obtaining the final text feature representation; calculating the multimodal hidden layer representation, and obtaining the final sentiment classification through a pooling layer, a fully connected layer, and a softmax. This solution extracts images from a multimodal dataset and combines two or more modalities to achieve cross-modal sentiment analysis, enhancing the understanding and recognition capabilities of image content, effectively addressing the limitations of a single modality, converting image information into a more expressive and semantically informative visual representation, and improving the stability of the multimodal system.
[0005] Therefore, the relevant technology can improve the matching degree between the semantics of subtitles and images or videos and improve the accuracy of subtitles. However, the relevant technology does not set the display effect of subtitles, resulting in the display of subtitles failing to highlight the key points of the image or video.
[0006] The information disclosed in the background technology section of this application is only intended to deepen the understanding of the general background technology of this application, and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to those skilled in the art. Summary of the Invention
[0007] The present invention provides a subtitle matching display method and system based on intelligent image processing, which can solve the technical problem in the related art that the display of subtitles cannot highlight the key points of an image or video.
[0008] According to a first aspect of the present invention, a subtitle matching and display method based on intelligent image processing is provided, comprising:
[0009] Parsing the video to be processed to obtain multiple video images;
[0010] Processing the target object region in the video image using an image information processing model to obtain lip shape information and facial expression information of the target object;
[0011] The audio file corresponding to the video to be processed is processed through the text recognition model to determine the text information corresponding to the audio file;
[0012] Processing the audio file corresponding to the video to be processed to obtain an audio subfile corresponding to the pronunciation of each text in the text information, and determining a video image corresponding to each audio subfile;
[0013] Determining pixel values of text corresponding to the audio subfile in the subtitles based on the audio subfile and the lip shape information and facial expression information of the target object in the video image;
[0014] Determining a display size of text in the subtitles corresponding to the audio subfile based on the expression information of the target object in the audio subfile and the video image;
[0015] Obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0016] Obtain subtitles of the video to be processed based on the display information of each text.
[0017] According to a second aspect of the present invention, there is provided a subtitle matching and display system based on intelligent image processing, comprising:
[0018] A parsing module, used to parse the video to be processed to obtain multiple video images;
[0019] An information module, configured to process the target object region in the video image using an image information processing model to obtain lip shape information and facial expression information of the target object;
[0020] The text module is used to process the audio file corresponding to the video to be processed through a text recognition model to determine the text information corresponding to the audio file;
[0021] The video image module is used to process the audio file corresponding to the video to be processed, obtain the audio sub-file corresponding to the pronunciation of each text in the text information, and determine the video image corresponding to each audio sub-file;
[0022] a pixel value module for determining pixel values of text corresponding to the audio subfile in the subtitles based on the audio subfile and the lip shape information and facial expression information of the target object in the video image;
[0023] A display size module, configured to determine the display size of text in the subtitles corresponding to the audio subfile based on the expression information of the target object in the audio subfile and the video image;
[0024] a display information module, configured to obtain display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0025] The subtitle module is used to obtain the subtitles of the video to be processed based on the display information of each text.
[0026] According to a third aspect of the present invention, a subtitle matching display device based on intelligent image processing is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the subtitle matching display method based on intelligent image processing.
[0027] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the subtitle matching and display method based on intelligent image processing is implemented.
[0028] By adopting the above technical solution, the present invention can achieve the following technical effects:
[0029] According to the present invention, the lip shape information and facial expression information of the target object can be obtained through the image information processing model, and the key points in the text information of the subtitles can be determined based on the lip shape information and facial expression information, so that the subtitles are set with specific pixel values and display sizes to highlight the key text in the subtitles, making it easier for viewers to watch and understand, and improving the display effect. When training the audio processing model, the lip shape pronunciation prediction model and the image information processing model, the competitive training of the audio processing model, the lip shape pronunciation prediction model and the image information processing model can be achieved through the conditional function. By continuously improving the accuracy of the audio processing model, the supervision standard of the lip shape pronunciation prediction model and the image information processing model is improved, so that the accuracy of the lip shape pronunciation prediction model and the image information processing model is close to and reaches the accuracy of the audio processing model, so as to achieve a common improvement in the accuracy of the three models and improve the training accuracy and efficiency. When determining the pixel adjustment coefficient, the possibility of encountering a situation that needs to be focused on can be determined by the spectrum information of the audio clip, the difference between the text pronunciation information and the lip shape pronunciation information, and the facial expression information, so as to obtain the pixel adjustment coefficient, which can objectively and comprehensively reflect the possibility of encountering a situation that needs to be focused on, and improve the accuracy of adjusting the pixel value. When determining the size adjustment coefficient, the volume information and expression information can be used to determine whether the target object has its volume amplified, and the size adjustment coefficient can be determined through a conditional function. Thus, when the volume is amplified, the size of the text in the subtitles is amplified, and when the volume is not amplified, the size of the text is maintained. This comprehensively and objectively reflects the relationship between the display size and volume of the text in the subtitles, effectively prompting the viewer's attention.
[0030] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and not limiting of the present invention. Other features and aspects of the present invention will become more apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can derive other embodiments based on these drawings without inventive efforts.
[0032] Figure 1 A schematic diagram exemplarily illustrates a flow chart of a subtitle matching and display method based on intelligent image processing according to an embodiment of the present invention;
[0033] Figure 2 A block diagram of a subtitle matching and display system based on intelligent image processing according to an embodiment of the present invention is exemplarily shown. DETAILED DESCRIPTION
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0035] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0036] Figure 1 A flowchart of a subtitle matching and display method based on intelligent image processing according to an embodiment of the present invention is exemplarily shown. The method includes:
[0037] Step S101: parsing the video to be processed to obtain multiple video images;
[0038] Step S102, processing the target object region in the video image using an image information processing model to obtain lip shape information and facial expression information of the target object;
[0039] Step S103: Processing the audio file corresponding to the video to be processed by a text recognition model to determine the text information corresponding to the audio file;
[0040] Step S104, processing the audio file corresponding to the video to be processed, obtaining an audio sub-file corresponding to the pronunciation of each text in the text information, and determining a video image corresponding to each audio sub-file;
[0041] Step S105, determining the pixel values of the text corresponding to the audio subfile in the subtitles based on the audio subfile and the lip shape information and facial expression information of the target object in the video image;
[0042] Step S106, determining the display size of the text corresponding to the audio subfile in the subtitles based on the expression information of the target object in the audio subfile and the video image;
[0043] Step S107, obtaining display information of the text corresponding to the audio sub-file in the video image according to the pixel value of the text and the display size;
[0044] Step S108: Obtain subtitles of the video to be processed according to the display information of each text.
[0045] According to the subtitle matching display method based on intelligent image processing in an embodiment of the present invention, the lip shape information and facial expression information of the target object can be obtained through an image information processing model, thereby determining the key points in the text information of the subtitles based on the lip shape information and facial expression information, and setting specific pixel values and display sizes for the subtitles to highlight the key text in the subtitles, making it easier for viewers to watch and understand, and improving the display effect.
[0046] According to one embodiment of the present invention, in step S101, a video to be processed may be parsed. The video to be processed may be a music video, a lecture video, a course video, or the like. The video includes a target object emitting sound, such as a singing target object, a speaking target object, a lecturing target object, and the like, and the video to be processed has a corresponding audio file. The video to be processed may be parsed to obtain multiple video images, i.e., multiple video frames.
[0047] According to one embodiment of the present invention, in step S102, the image information processing model is a deep learning neural network model, such as a convolutional neural network model, which can process the area where the target object in the video image is located to obtain the target object's lip shape information and expression information, both of which are information in the form of vectors. Each component of the lip shape information vector can represent the relative positional relationship between multiple key points on the lips, such as the distance and angle between the multiple key points. Combined with the relative positional relationship between the multiple key points, it can reflect the target object's lip shape. The multiple components of the expression information vector can represent the probability that the target object's expression belongs to various expression types, such as the probability of belonging to a serious expression, the probability of belonging to an excited expression, the probability of belonging to an angry expression, the probability of belonging to a neutral expression, etc.
[0048] According to one embodiment of the present invention, in step S103, the audio file corresponding to the video to be processed may be processed using a text recognition model to determine the text information corresponding to the audio file. The text recognition model is a deep learning neural network model, such as a recursive neural network model, which can convert the audio file into text information.
[0049] According to one embodiment of the present invention, in step S104, the audio file can be segmented to obtain audio subfiles corresponding to the pronunciation of each text. For example, an entire audio segment can be segmented into multiple audio subfiles, each of which contains only the pronunciation of a single word. Furthermore, video images have timestamps that correspond to moments in the audio file. Therefore, the start and end times of the audio subfiles can be used to find the video image with the corresponding timestamp, which serves as the video image corresponding to the start and end times of the audio subfile. Multiple video images between these two video images can then be identified as the video images corresponding to the audio subfile.
[0050] According to one embodiment of the present invention, in step S105, pixel values of text corresponding to the audio subfile in the subtitles are determined based on the audio subfile, the lip shape information, and the facial expression information of the target object in the video image, including: determining the timestamps of multiple video images corresponding to the audio subfile; determining, based on the timestamps, an audio segment within a time period between the timestamp of the i-th video image and the timestamp of the (i+1)-th video image in the audio subfile; determining text pronunciation information of the audio segment through an audio processing model, wherein the text pronunciation information is used to represent pronunciation features of the text corresponding to the audio segment; determining lip shape pronunciation information of the lip shape information through a lip shape pronunciation prediction model, wherein the lip shape pronunciation information is used to represent pronunciation features of sounds that can be produced based on the lip shape information; performing spectral analysis on the audio segment to determine spectral information of the audio segment; and determining, based on the text pronunciation information, the lip shape pronunciation information, the spectral information, and the facial expression information, pixel values of the text corresponding to the audio subfile in the subtitles of the i-th video image.
[0051] According to one embodiment of the present invention, each audio subfile contains only the pronunciation of a single word, but the pronunciation duration of each word varies. This results in different durations for each audio subfile, and therefore, different numbers of video images corresponding to each audio subfile. Furthermore, within the duration of a word's pronunciation, parameters such as the volume and pitch of the word may also change. Therefore, the audio subfiles can be further divided. To facilitate analysis in conjunction with the video images, the video image timestamps can be used as segmentation points for the audio subfiles, dividing the audio subfiles into multiple audio segments. This allows analysis of the audio segments between the timestamp of the i-th video image and the timestamp of the (i+1)-th video image in conjunction with the i-th video image.
[0052] According to one embodiment of the present invention, the audio processing model can be a BP neural network model, which can analyze the audio segment to obtain text pronunciation information of the audio segment. The text pronunciation information is information in vector form, which can be used to represent the pronunciation characteristics of the text corresponding to the audio segment. For example, the multiple components of the text pronunciation information may include parameters representing the pronunciation type of the text (for example, the pronunciation is "a", "o", "e", etc.), pronunciation tone (for example, the pronunciation tone can be described by characteristics such as the frequency of the pronunciation), volume and other information.
[0053] According to one embodiment of the present invention, the lip shape pronunciation prediction model can be a BP neural network model, which can process the lip shape information of the video image to predict the pronunciation characteristics of the sound that can be emitted by the lip shape information, and obtain the lip shape pronunciation information. The lip shape pronunciation information is of the same type as the text pronunciation information, both are information in vector form, and both can include parameters representing information such as pronunciation type, pronunciation tone, volume, etc.
[0054] According to one embodiment of the present invention, spectrum analysis may include obtaining spectrum information of an audio clip through a Fourier transform or other methods, and obtaining frequency characteristics of the audio clip from the spectrum information. Furthermore, pixel values of the text corresponding to the audio subfile may be determined based on the text pronunciation information, lip pronunciation information, spectrum information, and expression information.
[0055] According to one embodiment of the present invention, the above-mentioned lip-pronunciation prediction model may be trained before use, and accurate lip-pronunciation information may be obtained after training, and then the above-mentioned process of determining the pixel value of the text may be performed. The training steps of the lip pronunciation prediction model include: obtaining first sample videos of multiple test subjects performing multiple pronunciations; obtaining multiple first images of the first sample video, and obtaining first sample lip shape information of the test subjects in the first image through an image information processing model; segmenting the audio file corresponding to the first sample video according to the timestamp of the first image to obtain sample audio clips, and processing the sample audio clips through an audio processing model to obtain sample text pronunciation information; processing the first sample lip shape information according to the lip pronunciation prediction model to obtain sample lip pronunciation information; performing spectrum analysis on the sample audio clip to obtain first sample spectrum information; obtaining reference pronunciation information based on the first sample spectrum information; determining a first comprehensive loss function of the lip pronunciation prediction model, the audio processing model and the image information processing model based on the sample text pronunciation information, the sample lip pronunciation information and the reference pronunciation information; training the lip pronunciation prediction model, the audio processing model and the image information processing model according to the first comprehensive loss function to obtain a trained lip pronunciation prediction model, a trained audio processing model and a trained image information processing model.
[0056] According to one embodiment of the present invention, scenes in which multiple test subjects are pronouncing words can be filmed to obtain first sample videos of the multiple test subjects performing various pronunciations, and the first sample videos can be analyzed to obtain multiple first images, i.e., video frames of the first sample videos. The first sample lip shape information of the test subjects in the first image can also be obtained through the above-mentioned image information processing model. The acquisition method is the same as the acquisition method of the above-mentioned lip shape information, and will not be repeated here.
[0057] According to one embodiment of the present invention, the audio file corresponding to the first sample video is segmented based on the timestamps of the first images to obtain sample audio segments between the timestamps of adjacent first images, for example, the sample audio segment between the timestamps of the i-th first image and the timestamps of the (i+1)-th first image. The sample audio segments can be processed using an audio processing model to obtain sample text pronunciation information. This processing process is the same as the process for obtaining text pronunciation information described above and is not further described here.
[0058] According to one embodiment of the present invention, the first sample mouth information can be processed by the mouth shape pronunciation prediction model to obtain sample mouth shape pronunciation information. The sample text pronunciation information and the sample mouth shape pronunciation information are information obtained by the untrained model (audio processing model and mouth shape pronunciation prediction model), and there may be errors. Therefore, by performing spectrum analysis on the sample audio segment and obtaining error-free reference pronunciation information based on the obtained first sample spectrum information, the error between the sample text pronunciation information and the sample mouth shape pronunciation information and the reference pronunciation information can be determined. The reference pronunciation information is information in the form of vectors, which may include parameters for information such as the pronunciation type, pronunciation pitch and volume of the text used to describe the sample audio segment, that is, the pronunciation type can be manually marked, and the frequency characteristics of the pronunciation can be determined by spectrum analysis, and then the pronunciation pitch can be described. The volume can also be determined by analysis of sound waves. The parameters obtained in these ways are accurate parameters. Therefore, the reference pronunciation information is error-free vector information.
[0059] According to one embodiment of the present invention, determining a first comprehensive loss function of a lip pronunciation prediction model, an audio processing model, and an image information processing model based on the sample text pronunciation information, the sample lip pronunciation information, and the reference pronunciation information includes: determining a first comprehensive loss function LOSS1 of the lip pronunciation prediction model, the audio processing model, and the image information processing model according to formula (1),
[0060]
[0061] Among them, P j,k,text is the sample text pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test person, P j,k,ms The sample mouth shape pronunciation information corresponding to the first sample mouth shape information of the kth first image of the jth test person, P j,k,R is the reference pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test person, (P j,k,text )T P j,k,text , w1 and w2 are preset weights, if is a conditional function, m is the number of first images in the first sample video, n is the number of test personnel, k≤m, j≤n, and k, j, m and n are all positive integers.
[0062] According to one embodiment of the present invention, in formula (1), the conditional function if{|P j,k,text -P j,k,R |<|P j,k,ms -P j,k,R |, w1|P j,k,text-P j,k,R |+w2|P j,k,ms -P j,k,R |} means in |P j,k,text -P j,k,R |<|P j,k,ms -P j,k,R |, the conditional function value is Otherwise, the conditional function value is w1|P j,k,text -P j,k,R |+w2|P j,k,ms -P j,k,R |. |P j,k,text -P j,k,R |<|P j,k,ms -P j,k,R | indicates that the error between the sample text pronunciation information and the reference pronunciation information is smaller than the error between the sample lip pronunciation information and the reference pronunciation information, wherein the sample text pronunciation information is obtained through an audio processing model, and the sample lip pronunciation information is obtained through an image information processing model and a lip pronunciation prediction model. Therefore, this condition of the conditional function also indicates that the error of the audio processing model is smaller than the error of the image information processing model and the lip pronunciation prediction model.
[0063] According to one embodiment of the present invention, based on the above analysis, when the condition of the conditional function is met, the error between the image information processing model and the mouth shape pronunciation prediction model is large, and the two models can be trained based on the error between the sample mouth shape pronunciation information and the reference pronunciation information. In addition, due to the large error, the error can also be used. As the denominator, the error is amplified, thereby improving the training intensity of the image information processing model and the mouth pronunciation prediction model, so that the error between the two is quickly reduced. The cosine similarity between the sample mouth pronunciation information and the sample text pronunciation information is used as the denominator, which can not only amplify the above error and improve the training intensity, but also can be used in the training process. The overall image is reduced, thereby improving the cosine similarity, thereby improving the consistency between the sample mouth pronunciation information and the sample text pronunciation information, and making the results obtained by the image information processing model and the mouth pronunciation prediction model close to the results obtained by the audio processing model, that is, making the accuracy of the image information processing model and the mouth pronunciation prediction model close to the accuracy of the audio processing model, thereby improving the accuracy of the image information processing model and the mouth pronunciation prediction model.
[0064] According to one embodiment of the present invention, if the condition of the conditional function is not met, that is, the error of the audio processing model is large and the accuracy is low, the weighted sum of the error between the sample text pronunciation information and the reference pronunciation information, and the error between the sample mouth shape pronunciation information and the reference pronunciation information is used as the conditional function value, so that the mouth shape pronunciation prediction model, the audio processing model and the image information processing model are trained respectively through the two errors, that is, the audio processing model is trained by the error between the sample text pronunciation information and the reference pronunciation information, and the image information processing model and the mouth shape pronunciation prediction model are trained by the error between the sample text pronunciation information and the reference pronunciation information, so that the accuracy of the three models are improved together.
[0065] Therefore, the above conditional function can realize the competitive training of the audio processing model, the lip pronunciation prediction model and the image information processing model, that is, the accuracy of the audio processing model is used to supervise the training of the lip pronunciation prediction model and the image information processing model. When the accuracy of the audio processing model is low, the accuracy of the three models can be improved at the same time. As the training progresses, if the accuracy of the audio processing model exceeds the accuracy of the lip pronunciation prediction model and the image information processing model, the training intensity of the lip pronunciation prediction model and the image information processing model is increased, so that the accuracy of the lip pronunciation prediction model and the image information processing model is close to and reaches the accuracy of the audio processing model. Moreover, as the accuracy of the audio processing model is improved, higher standards can be used to supervise the training of the lip pronunciation prediction model and the image information processing model, so that the three models can be quickly improved and the training efficiency is improved.
[0066] According to one embodiment of the present invention, the conditional function values corresponding to each first image of each test person can be summed to obtain a first comprehensive loss function, and the gradient descent method can be used for feedback propagation to adjust the parameters of the above three models. After multiple trainings, the training can be completed to obtain a trained lip pronunciation prediction model, a trained audio processing model, and a trained image information processing model.
[0067] In this way, competitive training of the audio processing model, the lip pronunciation prediction model, and the image information processing model can be achieved through conditional functions. By continuously improving the accuracy of the audio processing model and improving the supervision standards of the lip pronunciation prediction model and the image information processing model, the accuracy of the lip pronunciation prediction model and the image information processing model can be close to and reach the accuracy of the audio processing model, so as to achieve a joint improvement in the accuracy of the three models and improve training accuracy and efficiency.
[0068] According to one embodiment of the present invention, the lip shape pronunciation information obtained by processing the lip shape information of the video image based on the lip shape pronunciation prediction model trained in the above manner can be used to represent the pronunciation characteristics of the sound that the lip shape can theoretically produce. The pronunciation characteristics are theoretically consistent with the pronunciation characteristics described by the text pronunciation characteristics, but if special circumstances are encountered, such as an emotional speech, a difficult singing performance, etc., it may cause deformation of the facial expression and the lip shape, making the lip shape pronunciation information inconsistent with the text pronunciation information. In other words, due to the deformation of the facial expression and the lip shape, the lip shape deviates from the lip shape during the normal pronunciation of the text corresponding to the audio sub-file. Therefore, when the difference between the lip shape pronunciation information and the text pronunciation information is large, there may be a special case, and the special case is a key case in the video to be processed, such as the key part of the speech or singing. On the other hand, the spectral information of the audio clip can also reflect to a certain extent whether there are special circumstances. For example, when the speech or singing is relatively dull, the voice is relatively low and the frequency of the sound waves is low. When the speaker is emotional or the singer sings high notes, the frequency of the sound waves is higher. Therefore, the frequency of the sound waves can be analyzed through the spectral information to determine whether there are special circumstances that require special attention, so that more eye-catching pixel values can be set for the subtitles to prompt the viewer's attention.
[0069] According to one embodiment of the present invention, determining the pixel value of the text corresponding to the audio subfile in the subtitle of the i-th video image based on the text pronunciation information, the lip pronunciation information, the spectrum information and the expression information includes: obtaining the maximum pronunciation frequency of the audio segment based on the spectrum information; determining the pixel adjustment coefficient εi of the text corresponding to the audio subfile in the subtitle of the i-th video image based on formula (2) ,
[0070]
[0071] Among them, f i,max is the maximum pronunciation frequency of the audio segment in the time period between the i-th video image and the i+1-th video image, f p P is the preset pronunciation frequency. i,text is the text pronunciation information of the audio segment in the time period between the i-th video image and the i+1-th video image, P i,ms is the lip shape pronunciation information of the ith video image, (P i,text ) T P i,text The transposed vector, E i is the expression information of the i-th video image, E 1,p is the preset first expression information, (E i ) T For E iThe transposed vector of ; determining the pixel value of the text corresponding to the audio sub-file in the subtitle of the i-th video image according to the pixel adjustment coefficient and the basic pixel value.
[0072] According to one embodiment of the present invention, in formula (2), the preset pronunciation frequency may be the average pronunciation frequency of the audio file corresponding to the video to be processed, It is the ratio of the maximum pronunciation frequency of the audio segment in the time period between the i-th video image and the i+1-th video image to the preset pronunciation frequency. The higher the ratio, the higher the possibility of encountering a situation that requires special attention, for example, the higher the possibility of encountering the high-pitched part of singing or the key part of a speech.
[0073] According to one embodiment of the present invention, in formula (2), in, is the cosine similarity between text pronunciation information and mouth pronunciation information, so, The difference between the text pronunciation information and the mouth pronunciation information is as mentioned above. When the difference between the two is large, there is a high possibility of encountering a situation that requires special attention. Therefore, As the amplification factor, the higher the possibility of encountering a situation that requires special attention, the larger the amplification factor.
[0074] According to one embodiment of the present invention, in formula (2), the preset first expression information may indicate that the expression of the target object is a specific expression, for example, indicating that the expression of the target object is an excited expression. For example, among the various components of the first expression information, the probability of belonging to a serious expression is 0, the probability of belonging to an excited expression is 1, the probability of belonging to an angry expression is 0, and the probability of belonging to a dull expression is 0. It is the cosine similarity between the facial expression information of the video image and the preset first facial expression information. The higher the cosine similarity, the higher the possibility that the target object is in an excited state and the higher the possibility of encountering a situation that requires special attention. Therefore, the cosine similarity can be used as another amplification factor, so that the higher the possibility of encountering a situation that requires special attention, the larger the amplification factor.
[0075] According to one embodiment of the present invention, the above three items can be multiplied to obtain a pixel adjustment coefficient for the text corresponding to the audio subfile in the subtitles of the i-th video image. The pixel value of the text corresponding to the audio subfile in the subtitles of the i-th video image can be obtained by multiplying the adjustment coefficient with the base pixel value. Specifically, if the text is highlighted, the pixel value of the text can be increased, and the color of the text can be deepened to draw the viewer's attention. Furthermore, since an audio subfile can correspond to multiple audio clips and multiple video images, each audio clip and video image can be assigned a pixel adjustment coefficient, while an audio subfile only corresponds to one text. Therefore, the text can also produce a color-changing effect to draw the viewer's attention. For example, if a word in the lyrics is sung in different tones during a performance, the word can be assigned multiple pixel adjustment coefficients during the performance period, so that the pixel value of the word can change during video playback, thereby drawing the viewer's attention.
[0076] In this way, the possibility of encountering a situation that requires special attention can be determined through the spectral information of the audio clip, the difference between the text pronunciation information and the lip pronunciation information, and the facial expression information, so as to obtain the pixel adjustment coefficient, which can objectively and comprehensively reflect the possibility of encountering a situation that requires special attention and improve the accuracy of adjusting the pixel value.
[0077] According to one embodiment of the present invention, in step S106, the display size of the text corresponding to the audio subfile can also be used to prompt the viewer to pay attention to situations that require special attention. The display size of the text corresponding to the audio subfile in the subtitles is determined based on the facial expression information of the target object in the audio subfile and the video image, including: determining the maximum volume information of the audio clip; determining the average volume information of the audio file; setting second facial expression information related to the pronunciation volume; and determining the display size of the text corresponding to the audio subfile in the subtitles based on the maximum volume information, the average volume information, and the second facial expression information.
[0078] According to one embodiment of the present invention, determining the display size of the text corresponding to the audio subfile in the subtitles according to the maximum volume information, the average volume information, and the second expression information includes: determining the size adjustment coefficient σ of the text corresponding to the audio subfile in the subtitles of the i-th video image according to formula (3) i ,
[0079]
[0080] Among them, V i,max is the maximum volume information of the audio clip in the time period between the i-th video image and the i+1-th video image, V ave is the average volume information, Ei is the facial expression information of the i-th video image, (E i ) T For E i The transposed vector, E 2,p is the second expression information, if is a conditional function; according to the size adjustment coefficient and the basic size value, the display size of the text corresponding to the audio subfile in the subtitle of the i-th video image is determined.
[0081] According to one embodiment of the present invention, in formula (3), the conditional function Indicates In the case of Otherwise 1. It is the ratio of the maximum volume information to the average volume information of the audio clip. The larger the ratio is, the louder the target object's pronunciation volume of the text is, and the more likely it is that the target object will encounter a situation that requires special attention. On the other hand, the second expression information is similar to the first expression information, and it can also be preset expression information. Moreover, the second expression information can be consistent with the first expression information, or it can be other preset expression information. For example, among the various components of the second expression information, the probability of a serious expression is 0, the probability of an excited expression is 0, the probability of an angry expression is 1, and the probability of a dull expression is 0. Then the second expression information can indicate that the target object's expression is an angry expression. Therefore, It can be used as a volume amplification factor, that is, the amplification factor that indicates the situation where the target object's volume is amplified. If the result of the two multiplication is greater than or equal to 1, it means that the target object has amplified its volume. As a resizing factor, the size of the text in the subtitles is enlarged when the target object increases the volume. If the result of multiplying the two is less than 1, the resizing factor is equal to 1, that is, the size of the text will not be reduced.
[0082] According to one embodiment of the present invention, the display size of the text corresponding to the audio sub-file can be determined by multiplying the size adjustment coefficient and the basic size value, so that when the target object that makes the sound increases the volume, the size of the text in the letters is enlarged. Moreover, since the audio sub-file corresponds to multiple video images and multiple audio clips, each video image and audio clip can obtain a size adjustment coefficient. Therefore, the size of the text corresponding to the audio sub-file is variable. For example, when the target object sings a word in the lyrics, its volume keeps changing, and the size of the word in the subtitles can also change, thereby attracting the viewer's attention.
[0083] In this way, whether the target object should amplify the volume can be determined through both volume information and expression information, and the size adjustment coefficient can be determined through a conditional function. Therefore, when the volume is amplified, the size of the text in the subtitles is enlarged, and when the volume is not amplified, the size of the text is maintained. This comprehensively and objectively reflects the relationship between the display size of the text in the subtitles and the volume, effectively prompting the viewer's attention.
[0084] According to one embodiment of the present invention, in step S107, the display information of the text corresponding to the audio sub-file in the video image can be obtained based on the pixel value of the text and the display size, that is, the color and size of the text at the corresponding moment of each video image are determined, so that the text is displayed in the video image.
[0085] According to one embodiment of the present invention, in step S108, based on the above method, the display information of each text can be determined to obtain the subtitles of the video to be processed, that is, the color and size of each text in each video image can be determined, so that the text can be displayed in the video image. For example, if the 1st to 10th video images are the video images corresponding to the first text, the first text can be displayed based on the display information of the 1st to 10th video images, and the other texts maintain the basic pixel value and basic size value. If the 11th to 20th video images are the video images corresponding to the second text, the second text can be displayed based on the display information of the 11th to 20th video images, and the other texts maintain the basic pixel value and basic size value... Based on this method, appropriate subtitles can be displayed in the video to be processed, thereby prompting the viewer to pay attention to the key content.
[0086] According to an embodiment of the present invention, a subtitle matching and display method based on intelligent image processing can obtain the lip shape information and facial expression information of the target object through an image information processing model, thereby determining the key points in the text information of the subtitle based on the lip shape information and facial expression information, thereby setting specific pixel values and display sizes for the subtitles to highlight the key text in the subtitles, making it easier for viewers to view and understand, and improving the display effect. When training the audio processing model, the lip shape pronunciation prediction model, and the image information processing model, competitive training of the audio processing model, the lip shape pronunciation prediction model, and the image information processing model can be achieved through a conditional function. By continuously improving the accuracy of the audio processing model, the supervision standard of the lip shape pronunciation prediction model and the image information processing model is improved, so that the accuracy of the lip shape pronunciation prediction model and the image information processing model approaches and reaches the accuracy of the audio processing model, thereby achieving a joint improvement in the accuracy of the three models and improving training accuracy and efficiency. When determining the pixel adjustment coefficient, the probability of encountering a situation that requires special attention can be determined by using the spectrum information of the audio clip, the difference between the text pronunciation information and the lip shape pronunciation information, and the facial expression information, thereby obtaining the pixel adjustment coefficient, which can objectively and comprehensively reflect the probability of encountering a situation that requires special attention, and improve the accuracy of adjusting the pixel value. When determining the size adjustment coefficient, the volume information and expression information can be used to determine whether the target object has its volume amplified, and the size adjustment coefficient can be determined through a conditional function. Thus, when the volume is amplified, the size of the text in the subtitles is amplified, and when the volume is not amplified, the size of the text is maintained. This comprehensively and objectively reflects the relationship between the display size and volume of the text in the subtitles, effectively prompting the viewer's attention.
[0087] Figure 2 A block diagram of a subtitle matching and display system based on intelligent image processing according to an embodiment of the present invention is exemplarily shown. The system includes:
[0088] A parsing module, used to parse the video to be processed to obtain multiple video images;
[0089] An information module, configured to process the target object region in the video image using an image information processing model to obtain lip shape information and facial expression information of the target object;
[0090] The text module is used to process the audio file corresponding to the video to be processed through a text recognition model to determine the text information corresponding to the audio file;
[0091] The video image module is used to process the audio file corresponding to the video to be processed, obtain the audio sub-file corresponding to the pronunciation of each text in the text information, and determine the video image corresponding to each audio sub-file;
[0092] a pixel value module for determining pixel values of text corresponding to the audio subfile in the subtitles based on the audio subfile and the lip shape information and facial expression information of the target object in the video image;
[0093] A display size module, configured to determine the display size of text in the subtitles corresponding to the audio subfile based on the expression information of the target object in the audio subfile and the video image;
[0094] a display information module, configured to obtain display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size;
[0095] The subtitle module is used to obtain the subtitles of the video to be processed based on the display information of each text.
[0096] According to one embodiment of the present invention, a subtitle matching display device based on intelligent image processing is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the subtitle matching display method based on intelligent image processing.
[0097] According to one embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the subtitle matching and display method based on intelligent image processing is implemented.
[0098] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0099] Those skilled in the art will appreciate that the embodiments of the present invention described above and shown in the accompanying drawings are intended to be illustrative only and are not intended to limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functional and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Any variations or modifications may be made to the embodiments of the present invention without departing from the principles described.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A subtitle matching display method based on intelligent image processing, characterized in that: include: Parsing the video to be processed to obtain multiple video images; Processing the target object region in the video image using an image information processing model to obtain lip shape information and facial expression information of the target object; The audio file corresponding to the video to be processed is processed through the text recognition model to determine the text information corresponding to the audio file; Processing the audio file corresponding to the video to be processed to obtain an audio subfile corresponding to the pronunciation of each text in the text information, and determining a video image corresponding to each audio subfile; Determining pixel values of text corresponding to the audio subfile in the subtitles based on the audio subfile and the lip shape information and facial expression information of the target object in the video image; Determining a display size of text in the subtitles corresponding to the audio subfile based on the expression information of the target object in the audio subfile and the video image; Obtaining display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size; Obtain subtitles of the video to be processed according to the display information of each text; Determining pixel values of text corresponding to the audio subfile in the subtitles according to the audio subfile and the lip shape information and facial expression information of the target object in the video image includes: Determining timestamps of a plurality of video images corresponding to the audio sub-file; Determine, according to the timestamp, an audio segment within a time period between the timestamp of the i-th video image and the timestamp of the (i+1)-th video image in the audio subfile; Determining text pronunciation information of the audio segment through an audio processing model, wherein the text pronunciation information is used to represent pronunciation features of the text corresponding to the audio segment; Determining mouth shape pronunciation information of the mouth shape information by using a mouth shape pronunciation prediction model, wherein the mouth shape pronunciation information is used to represent pronunciation features of a sound that can be produced based on the mouth shape information; Performing spectrum analysis on the audio clip to determine spectrum information of the audio clip; Determining pixel values of text corresponding to the audio subfile in the subtitles of the i-th video image based on the text pronunciation information, the lip pronunciation information, the spectrum information, and the expression information; The training steps of the mouth shape pronunciation prediction model include: Obtaining first sample videos of multiple test subjects performing multiple pronunciations; Acquire multiple first images of a first sample video, and obtain first sample lip shape information of a test person in the first image through an image information processing model; Segmenting the audio file corresponding to the first sample video according to the timestamp of the first image to obtain sample audio segments, and processing the sample audio segments through the audio processing model to obtain sample text pronunciation information; Processing the first sample mouth shape information according to the mouth shape pronunciation prediction model to obtain sample mouth shape pronunciation information; Performing spectrum analysis on the sample audio segment to obtain first sample spectrum information; Obtaining reference pronunciation information according to the first sample spectrum information; Determining a first comprehensive loss function of a lip pronunciation prediction model, an audio processing model, and an image information processing model based on the sample text pronunciation information, the sample lip pronunciation information, and the reference pronunciation information; According to the first comprehensive loss function, the lip pronunciation prediction model, the audio processing model and the image information processing model are trained to obtain the trained lip pronunciation prediction model, the trained audio processing model and the trained image information processing model.
2. The subtitle matching display method based on intelligent image processing according to claim 1, characterized in that: Determining a first comprehensive loss function of a lip pronunciation prediction model, an audio processing model, and an image information processing model based on the sample text pronunciation information, the sample lip pronunciation information, and the reference pronunciation information, including: According to the formula Determine the first comprehensive loss function LOSS1 of the mouth pronunciation prediction model, audio processing model and image information processing model, where P j,k,text is the sample text pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test person, P j,k,ms The sample mouth shape pronunciation information corresponding to the first sample mouth shape information of the kth first image of the jth test person, P j,k,R is the reference pronunciation information of the sample audio segment between the timestamp of the kth first image and the timestamp of the k+1th first image of the jth test person, (P j,k,text ) T P j,k,text , w1 and w2 are preset weights, if is a conditional function, m is the number of first images in the first sample video, n is the number of test personnel, k≤m, j≤n, and k, j, m and n are all positive integers.
3. The subtitle matching display method based on intelligent image processing according to claim 1, characterized in that: Determining pixel values of text corresponding to the audio subfile in the subtitles of the i-th video image according to the text pronunciation information, the lip pronunciation information, the spectrum information, and the expression information includes: According to the spectrum information, obtain the maximum pronunciation frequency of the audio clip; According to the formula Determine the pixel adjustment coefficient ε of the text corresponding to the audio sub-file in the subtitle of the i-th video image i , where f i,max is the maximum pronunciation frequency of the audio segment in the time period between the i-th video image and the i+1-th video image, f p P is the preset pronunciation frequency. i,text is the text pronunciation information of the audio segment in the time period between the i-th video image and the i+1-th video image, P i,ms is the lip shape pronunciation information of the ith video image, (P i,text ) T P i,text The transposed vector, E i is the expression information of the i-th video image, E 1,p is the preset first expression information, (E i ) T For E i The transposed vector of The pixel value of the text corresponding to the audio sub-file in the subtitle of the i-th video image is determined according to the pixel adjustment coefficient and the basic pixel value.
4. The subtitle matching display method based on intelligent image processing according to claim 1, characterized in that: Determining a display size of text corresponding to the audio subfile in a subtitle according to the audio subfile and the expression information of a target object in a video image includes: Determine the maximum volume information of an audio clip; Determine the average volume information of an audio file; Setting the second expression information related to the pronunciation volume; A display size of text corresponding to the audio subfile in the subtitles is determined according to the maximum volume information, the average volume information, and the second expression information.
5. The subtitle matching display method based on intelligent image processing according to claim 4, characterized in that: Determining a display size of text corresponding to the audio subfile in a subtitle according to the maximum volume information, the average volume information, and the second expression information includes: According to the formula Determine the size adjustment coefficient σ of the text corresponding to the audio subfile in the subtitle of the i-th video image i , where V i,max is the maximum volume information of the audio clip in the time period between the i-th video image and the i+1-th video image, V ave is the average volume information, E i is the facial expression information of the i-th video image, (E i ) T For E i The transposed vector, E 2,p is the second expression information, if is a conditional function; The display size of the text corresponding to the audio sub-file in the subtitle of the i-th video image is determined according to the size adjustment coefficient and the basic size value.
6. A subtitle matching and display system based on intelligent image processing, used to execute the method according to any one of claims 1 to 5, characterized in that: include: A parsing module, used to parse the video to be processed to obtain multiple video images; An information module, configured to process the target object region in the video image using an image information processing model to obtain lip shape information and facial expression information of the target object; The text module is used to process the audio file corresponding to the video to be processed through a text recognition model to determine the text information corresponding to the audio file; The video image module is used to process the audio file corresponding to the video to be processed, obtain the audio sub-file corresponding to the pronunciation of each text in the text information, and determine the video image corresponding to each audio sub-file; a pixel value module for determining pixel values of text corresponding to the audio subfile in the subtitles based on the audio subfile and the lip shape information and facial expression information of the target object in the video image; A display size module, configured to determine the display size of text in the subtitles corresponding to the audio subfile based on the expression information of the target object in the audio subfile and the video image; a display information module, configured to obtain display information of the text corresponding to the audio subfile in the video image according to the pixel value of the text and the display size; The subtitle module is used to obtain the subtitles of the video to be processed based on the display information of each text.
7. A subtitle matching display device based on intelligent image processing, characterized in that: include: processor; A memory for storing processor-executable instructions; wherein the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that Computer program instructions are stored thereon, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Multi-modal sentiment analysis method and system fusing image subtitles and BERT
CN117115516A
Live video stream processing method and device, electronic equipment and storage medium
CN115086753A
Video display method and device, equipment and storage medium
CN116055792A
Intelligent person auxiliary communication glasses based on deep learning
CN118262598A