Method, device, storage medium and program product for intelligently repairing a slip of the tongue in a video

By separating the audio and video streams of video files, extracting features, and constructing a multi-dimensional cost function, the problem of visual and auditory inconsistencies caused by verbal slips in video editing is solved, achieving smooth transitions and natural coherence in the video.

CN121908089BActive Publication Date: 2026-07-24BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING QINGSONG YIKANG INFORMATION TECHNOLOGY CO LTD
Filing Date
2026-03-24
Publication Date
2026-07-24

Smart Images

  • Figure CN121908089B_ABST
    Figure CN121908089B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a method and device for intelligently repairing a slip of the tongue in a video, a storage medium and a program product. The method comprises: parsing a video file to be repaired into an audio stream file and a video picture stream file; extracting a slip of the tongue segment from the audio stream file and the video picture stream file to obtain an audio slip of the tongue segment and a video picture slip of the tongue segment; extracting features from the audio slip of the tongue segment and the video picture slip of the tongue segment to obtain audio features and video picture features; constructing a multi-dimensional cost comprehensive function based on the audio features and the video picture features, wherein the dimensions include at least one of video picture continuity, audio physical feature stability and prosodic consistency; solving the multi-dimensional cost comprehensive function to obtain an optimal cutting point; and repairing the video file based on the optimal cutting point to generate a new video file. The method can simultaneously realize smooth transition of video character action, natural coherence of audio prosody, and avoid audio-visual discontinuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method, apparatus, storage medium, and program product for intelligently correcting verbal slips in videos. Background Technology

[0002] In content production scenarios such as audio-visual videos and podcasts, editing out slips of the tongue is a core step in improving content quality, and the low efficiency and high cost of manual editing urgently need to be addressed. Currently, the mainstream solution adopts a "cascaded" automated processing: first, audio is converted to text using ASR technology to locate slip-of-the-tongue keywords and redundant content; then, physical segmentation is performed using silent clips to complete the removal of slip-of-the-tongue segments and the splicing of effective content.

[0003] However, this traditional approach based on a single audio modality does not take into account the continuity of video footage. Editing often occurs when the speaker's head moves or their lip movements change, resulting in obvious visual abrupt changes in the synthesized video. Moreover, simple physical splicing cannot guarantee that the tone and breathing rhythm of the preceding and following sentences are consistent, causing rhythmic breaks and a harsh listening experience. Summary of the Invention

[0004] In view of this, the present disclosure provides a method, apparatus, storage medium and program product for intelligently correcting verbal slips in videos, which can simultaneously achieve smooth transitions of video characters' movements and natural and coherent audio rhythm, avoiding the audiovisual discontinuity problem caused by traditional hard cutting.

[0005] In a first aspect, embodiments of this disclosure provide a method for intelligently correcting verbal slips in videos, employing the following technical solution:

[0006] The video file to be repaired is parsed into an audio stream file and a video stream file;

[0007] Extract speech error segments from the audio stream file and the video stream file to obtain audio speech error segments and video speech error segments;

[0008] Feature extraction is performed on the audio and video clips containing verbal slips to obtain audio and video features.

[0009] Based on the audio features and the video features, a multi-dimensional cost synthesis function is constructed, wherein the dimension includes at least one of video continuity, audio physical feature stability and prosodic consistency.

[0010] Solve the multi-dimensional cost synthesis function to obtain the optimal shear point;

[0011] Based on the optimal cutting point, the video file is repaired to generate a new video file.

[0012] Optionally, the step of extracting verbal slips from the audio stream file and the video stream file to obtain audio and video verbal slips includes:

[0013] Perform speech recognition on the audio stream file to obtain the text.

[0014] Perform semantic analysis on the text to determine the start and end time range of the slip of the tongue segment;

[0015] According to the start and end time range, the audio stream file and the video stream file are used to extract verbal slips to obtain audio verbal slips and video verbal slips.

[0016] Optionally, the step of constructing a multi-dimensional cost synthesis function based on the audio features and the video frame features includes:

[0017] Spatiotemporal alignment and semantic complementarity are performed on the audio features and video features to obtain audio-video alignment features, and the audio-video alignment features are then decomposed into environmental features and subject modality features.

[0018] Based on the subject modal features, a multi-dimensional cost function is constructed, wherein the cost function includes at least one of the following: video frame continuity cost function, audio physical feature stability cost function, and prosodic consistency cost function;

[0019] Based on the aforementioned environmental characteristics, the weight coefficients of each cost function are obtained;

[0020] Based on the cost function and the corresponding weight coefficients, a multi-dimensional cost synthesis function is constructed.

[0021] Optionally, the subject modal features include normalized two-dimensional coordinate vectors of facial key points in each video frame and three-dimensional Euler angle vectors of the head;

[0022] Based on the normalized two-dimensional coordinate vector of the facial key points and the three-dimensional Euler angle vector of the head, a video frame continuity cost function is constructed.

[0023] The expression for the video frame continuity cost function is:

[0024]

[0025] In the formula, The cost function representing the continuity of video frames; The time index representing the candidate shear point; Normalized weighting coefficients representing micro-level characteristics; This represents the total number of facial landmarks. An index representing facial landmarks; Denotes the Euclidean norm; Indicates the time index of the frame preceding the candidate cut point; Indicates the time index of the frame following the candidate cut point; Indicates the first The first video frame at time [time] Normalized two-dimensional coordinate vectors of key facial features; Indicates the first The first video frame at time [time] Normalized two-dimensional coordinate vectors of key facial features; Normalized weighting coefficients representing macroscopic posture; Represents the Manhattan norm; Indicates the first The three-dimensional Euler angle vector of the head in the video frame at a given moment; Indicates the first The three-dimensional Euler angle vector of the head in the video frame at a given time.

[0026] Optionally, the main modal features also include the amplitude values ​​of audio sampling points;

[0027] Based on the amplitude values ​​of the audio sampling points, an audio physical feature stability cost function is constructed;

[0028] The expression for the audio physical feature stability cost function is:

[0029]

[0030] In the formula, This represents the cost function for the stability of audio physical characteristics; This represents the weighting coefficient that adjusts the absolute value of the amplitude; Indicates the first The amplitude value of the audio sampling point at a given moment; Indicates the first The absolute value of the amplitude at the audio sampling point at any given time; Weighting coefficients representing the importance of energy in the regulation region; This indicates the window size for short-time energy calculations; Indicates in window Index variable for internal audio sampling points; Indicates the first The amplitude value of the audio sampling point at a given time.

[0031] Optionally, the main modal features also include the fundamental frequency of the audio and the short-time energy envelope of the audio;

[0032] Based on the fundamental frequency and short-time energy envelope of the audio, a prosodic consistency cost function is constructed;

[0033] The expression for the prosodic consistency cost function is:

[0034]

[0035] In the formula, Represents the prosodic consistency cost function; Indicates the first The fundamental frequency of the audio at any given moment; Indicates the first The fundamental frequency of the audio at any given moment; Indicates the first The short-time energy envelope of the moment audio; Indicates the first The short-time energy envelope of the moment audio.

[0036] Optionally, obtaining the weight coefficients of each cost function based on the environmental features includes:

[0037] The environmental features are then categorized and spliced ​​together to generate multidimensional environmental splicing features.

[0038] The multidimensional environment splicing features are input into a lightweight fully connected network to obtain the weight coefficients of each cost function output by the lightweight fully connected network.

[0039] Secondly, this disclosure also provides a system for intelligently correcting verbal slips in videos, employing the following technical solution:

[0040] The file parsing module is used to parse the video file to be repaired into audio stream files and video stream files;

[0041] The segment extraction module is used to extract speech error segments from the audio stream file and the video stream file to obtain audio speech error segments and video speech error segments.

[0042] The feature extraction module is used to extract features from the audio slip-up segments and the video slip-up segments to obtain audio features and video features;

[0043] The function construction module is used to construct a multi-dimensional cost synthesis function based on the audio features and the video frame features, wherein the dimensions include at least one of video frame continuity, audio physical feature stability and prosodic consistency;

[0044] The function solving module is used to solve the multi-dimensional cost synthesis function and obtain the optimal shearing point;

[0045] The video repair module is used to repair the video file based on the optimal cut point and generate a new video file.

[0046] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution:

[0047] The computer device includes:

[0048] At least one processor; and,

[0049] A memory communicatively connected to the at least one processor; wherein,

[0050] The memory stores instructions that can be executed by the at least one processor, which enables the at least one processor to perform any of the above-described methods for intelligently correcting verbal slips in videos.

[0051] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to perform any of the above-described methods for intelligently correcting verbal slips in videos.

[0052] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0053] The method for intelligently correcting verbal slips in videos provided in this disclosure separates the audio and video data by parsing the video to be corrected into independent audio and video streams, laying the foundation for subsequent bimodal synchronous analysis. Then, the method performs joint localization of verbal slip segments on the two data streams, simultaneously extracting the corresponding audio and video verbal slip segments to ensure spatiotemporal consistency of the audio and video verbal slip areas. Based on this, deep feature extraction is performed on the audio and video verbal slip segments respectively, obtaining multidimensional features that characterize the audio's physical properties, prosodic structure, and the video's content, lip movements, and posture. Subsequently, using video continuity, audio physical feature stability, and prosodic consistency as key constraints, a multidimensional cost synthesis function is constructed, unifying and quantifying video smoothness, audio stability, and sentence coherence into a solvable optimization objective. By optimizing the multi-dimensional cost function, the optimal cut point can be automatically located, simultaneously satisfying the requirements of no abrupt changes in visuals, no audio distortion, and no rhythmic breaks. Based on this cut point, the original video is then precisely cut and seamlessly synthesized. This achieves intelligent repair of speech errors without relying on manual frame-by-frame proofreading. The repaired video maintains a high degree of synchronization between the speaker's lip movements, posture, and speech, while preserving natural and coherent intonation, breathing rhythm, and semantic expression. This significantly improves the video's smoothness, integrity, and realism from both visual and auditory perspectives, effectively reducing the damage to the overall viewing experience caused by speech error repair. Compared to existing repair schemes that rely solely on a single audio modality, this solution overcomes the limitation of focusing only on audio while neglecting visual continuity. It fundamentally avoids visual abrupt changes caused by improper cutting during head movements and lip movements. Furthermore, by using a multi-dimensional cost function to intelligently solve for the optimal cut point, rather than simple physical splicing, it effectively solves the shortcomings of traditional solutions, such as rhythmic breaks and abrupt sounds, achieving significant improvements in both repair accuracy and audiovisual experience.

[0054] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0055] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 A flowchart illustrating the method for intelligently correcting verbal slips in videos provided in this embodiment of the disclosure;

[0057] Figure 2 A flowchart illustrating the method for extracting verbal slips provided in this embodiment of the disclosure;

[0058] Figure 3 A flowchart illustrating the method for constructing a multi-dimensional cost synthesis function provided in this embodiment of the disclosure;

[0059] Figure 4 A flowchart illustrating the method for obtaining the weight coefficients of the cost function provided in this embodiment of the disclosure;

[0060] Figure 5 A schematic diagram of the system for intelligently correcting verbal slips in videos provided in this embodiment of the disclosure;

[0061] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0062] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0063] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0064] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0065] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0066] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0067] Reference Figure 1 This disclosure provides a method for intelligently correcting verbal slips in videos, including the following steps:

[0068] S1: Parse the video file to be repaired into an audio stream file and a video stream file;

[0069] S2: Extract verbal slips from the audio stream file and the video stream file to obtain audio verbal slips and video verbal slips;

[0070] S3: Extract features from the audio and video clips containing verbal slips to obtain audio and video features;

[0071] S4: Based on the audio features and the video features, construct a multi-dimensional cost synthesis function, wherein the dimension includes at least one of video continuity, audio physical feature stability and prosodic consistency;

[0072] S5: Solve the multi-dimensional cost synthesis function to obtain the optimal shearing point;

[0073] S6: Based on the optimal cutting point, repair the video file and generate a new video file.

[0074] This disclosure provides a method for intelligently correcting verbal slips in videos. By parsing the video to be corrected into independent audio and video streams, it achieves separate processing of audio and video data, laying the foundation for subsequent bimodal synchronous analysis. Then, it performs joint localization of verbal slip segments from the two data streams, simultaneously extracting the corresponding audio and video verbal slip segments to ensure spatiotemporal consistency of the audio and video slip areas. Based on this, it performs deep feature extraction on the audio and video verbal slip segments respectively, obtaining multi-dimensional features that characterize the audio's physical properties, prosodic structure, and the video's content, lip movements, and posture. Subsequently, using video continuity, audio physical feature stability, and prosodic consistency as key constraints, a multi-dimensional cost synthesis function is constructed, unifying and quantifying video smoothness, audio stability, and sentence coherence into a solvable optimization objective.

[0075] By optimizing the multi-dimensional cost function, the optimal cutting point can be automatically located, simultaneously satisfying no abrupt changes in the image, no audio distortion, and no rhythmic breaks. Based on this cutting point, the original video is then precisely cut and seamlessly synthesized. This enables intelligent repair of verbal slips without relying on manual frame-by-frame proofreading. The repaired video maintains a high degree of synchronization between the person's lip movements, posture, and speech, while also preserving the natural coherence of intonation, breathing rhythm, and semantic expression. This significantly improves the smoothness, integrity, and realism of the video from both visual and auditory dimensions, effectively reducing the damage to the overall viewing experience caused by verbal slip repair.

[0076] Compared to existing restoration solutions that rely solely on a single audio modality, this solution overcomes the limitation of focusing only on audio while neglecting the continuity of the visuals. It fundamentally avoids visual abrupt changes caused by improper cutting during the speaker's head movements and lip movements. At the same time, it achieves intelligent optimal cutting point solving through a multi-dimensional cost function, rather than simple physical splicing, effectively solving the defects of sentence rhythm discontinuity and abrupt sound in traditional solutions. It achieves significant improvements in both restoration accuracy and audiovisual experience.

[0077] In S1, the video file to be repaired is received and parsed into independent audio stream files (such as converting to WAV format) and video frame stream files (such as extracting frame sequences in MP4 format).

[0078] In S2, refer to Figure 2 The flowchart illustrating the method for extracting verbal slips demonstrates that "extracting verbal slips from audio stream files and video stream files to obtain audio and video slip segments" includes the following steps:

[0079] S21: Perform speech recognition on the audio stream file to obtain the text.

[0080] S22: Perform semantic analysis on the text to determine the start and end time range of the slip of the tongue;

[0081] S23: Extract speech error segments from the audio stream file and video stream file according to the start and end time range to obtain audio speech error segments and video speech error segments.

[0082] In S21, the end-to-end automatic speech recognition (ASR) model is invoked to transcribe the audio stream file segment by segment, and the audio timestamp corresponding to each text segment is recorded synchronously. At the same time, the text filtering module removes meaningless interjections and repetitive expressions from the recognition results, and finally generates accurate text bound to the audio stream timeline, providing basic data for subsequent error localization.

[0083] In S22, when standard subtitle text exists, the start and end time range of the slip of the tongue segment is determined by comparing the acquired text with the standard subtitle text. When there is no standard subtitle text, the Large Language Model (LLM) is invoked to perform deep semantic analysis on the text bound to the audio stream timeline. First, based on contextual logical consistency verification, suspected slips of the tongue with semantic contradictions, repetitive expressions, and incorrect word usage are identified in the text. Then, combined with the audio timestamp information corresponding to the text segment, the corresponding interval in the audio and video stream is matched in reverse. Finally, the semantic rationality of the suspected slip of the tongue is judged a second time by the LLM to accurately define the start and end time range of the slip of the tongue segment, providing a clear target interval for the subsequent selection of the cutting window.

[0084] In step S23, preset intervals before and after the start and end time ranges are identified. Within each interval, the frame with the best image quality is determined. The time points corresponding to these two frames are used as the start and end ranges for the final segment extraction. For example, if the start and end time range of a slip of the tongue segment is 5s to 15s, the frame with the best image quality is selected within the preset interval of 4s–5s before and after the start point (e.g., 4.6s), and the frame with the best image quality is selected within the preset interval of 15s–16s before and after the end point (e.g., 15.3s). Finally, 4.6s to 15.3s is used as the extraction range for the slip of the tongue segment in the audio stream file and the video stream file. The definition of "best image quality" is: selecting the best from audiovisual and visual perspectives, defining multiple evaluation dimensions, and combining these evaluation results to obtain the optimal frame.

[0085] The above solution uses speech recognition and semantic analysis to accurately locate the time interval of the slip of the tongue. Then, combined with the picture quality optimization strategy, it adaptively selects the best picture frame within a local preset interval of start and end time points. This achieves accurate location of the slip of the tongue segment and optimal extraction of picture quality, ensuring that the extracted segment not only corresponds to the slip of the tongue content but also has the best picture effect, thereby improving the accuracy of audio and video editing and content processing and the quality of the finished product.

[0086] In S3, a dual-thread approach is used to extract features from both audio and video clips of verbal slips. The audio processing module extracts audio features from the audio slip segments, including but not limited to spectrograms, zero-crossing rates, zero-crossing rate variations, short-time energy, energy transitions, and prosodic features. The specific spectrogram can be a Mel spectrogram; the zero-crossing rate refers to the number of times an audio signal crosses zero level per unit time, and is a basic indicator describing the time-domain characteristics of a speech signal; the zero-crossing rate variation refers to the amount or rate of change of the zero-crossing rate within adjacent time periods, reflecting the degree of fluctuation of the zero-crossing rate over time; short-time energy refers to the energy value of each frame of the audio signal after it is divided into frames, that is, the sum of the squares of the signal amplitude within the frame, and is a basic time-domain characteristic describing the energy distribution of the speech signal; energy jump refers to the short-time energy difference between two adjacent frames of audio signal, or the amplitude of the short-time energy change, reflecting the degree of abrupt change in signal energy; prosodic features include the fundamental frequency (FO) and the energy envelope. The fundamental frequency is the basic frequency of the speech signal, corresponding to the frequency of vocal cord vibration, and is one of the core parameters characterizing the prosodic features of speech, directly reflecting the pitch of the speech; the energy envelope is a smooth curve describing the change of speech signal energy over time, and is one of the core time-domain parameters characterizing the prosodic features of speech.

[0087] The video processing module extracts video features from the verbal slip segments in the video footage. These features include, but are not limited to, facial features, lip features, head pose features, and visual features. Facial features include the coordinates of key facial feature points (such as the pixel positions of the corners of the eyes and the tip of the nose), facial expression categories (such as neutral, smiling, and serious basic expressions), and the position and size of the face bounding box (representing the area of ​​the face in the frame). Lip features include key points of the lip contour (such as the upper lip peak, lower lip valley, and the closed state of the corners of the mouth), and the degree of lip opening (such as the quantified values ​​of closed, half-open, and fully open). Lip movement trajectory (displacement changes of lip feature points between adjacent frames), etc.; head pose features include head Euler angles (pitch angle, yaw angle, roll angle, representing the angles of head up and down, left and right, and rotation), head movement speed (rate of change of head pose angle between adjacent frames), and the relative position of the head in the picture (such as whether it is centered and the magnitude of offset), etc.; visual features include the overall light flow field of the picture (representing the pixel movement trend of adjacent frames), the texture features of the background area (such as the pixel distribution pattern of walls and curtains), and the brightness and contrast of the picture (reflecting the lighting stability of the shooting environment), etc.

[0088] In S4, refer to Figure 3 The flowchart illustrating the method for constructing a multi-dimensional cost synthesis function, "Constructing a multi-dimensional cost synthesis function based on audio features and video image features," includes the following steps:

[0089] S41: Perform spatiotemporal alignment and semantic complementation on audio features and video features to obtain audio-video alignment features, and then decompose the audio-video alignment features into environmental features and subject modality features;

[0090] S42: Based on the subject modal features, construct a multi-dimensional cost function, wherein the cost function includes at least one of the following: video frame continuity cost function, audio physical feature stability cost function, and prosodic consistency cost function;

[0091] S43: Based on environmental characteristics, obtain the weight coefficients of each cost function;

[0092] S44: Construct a multi-dimensional cost synthesis function based on the cost function and the corresponding weight coefficients.

[0093] In S41, feature embedding and positional encoding preprocessing are performed. The sequence V, composed of extracted visual features from the video frame, and the sequence A, composed of audio features, are encoded and mapped through a 1D convolutional layer with kernel=3 and stride=1 and a global average pooling layer, respectively. This mapping is uniformly mapped to a latent feature space of dimension D (D=64 / 128 is recommended), eliminating the difference in feature dimensionality between modalities to enable subsequent interaction. To address the issue of different audio and video sampling rates (e.g., 25fps video, 16kHz audio), relative positional encoding is introduced for the mapped audio and video temporal features. While preserving the temporal dependencies of each modality, a temporal reference benchmark is established for cross-modal spatiotemporal alignment, avoiding temporal disorder caused by sampling frequency differences.

[0094] After feature preprocessing, precise spatiotemporal alignment and preliminary semantic complementarity of audio and video are achieved through bidirectional cross-modal attention interaction. On one hand, audio-guided visual alignment is performed, using the audio feature sequence A as the query vector and the video image feature sequence V as the key and value vectors. An attention weight matrix W is calculated through scaling and dot product to generate the aligned video feature V. align =W⋅V accurately identifies the video visual frame most relevant to the pronunciation action of the current audio frame (e.g., locating the lip-closing action frame corresponding to a plosive sound), solving the physical delay problem of audio-visual asynchrony; on the other hand, it implements visually guided audio enhancement, using the video image feature sequence V as the query and the audio feature sequence A as the key and value, and weights the audio features through visual cues (such as frowning, wandering eyes) to generate aligned audio features A. align This enhances the sensitivity to audio features related to slips of the tongue, such as hesitation and pauses. Simultaneously, a monotonic alignment loss is introduced to construct temporal constraints, constraining A... align With V alignTo satisfy the requirement of spatiotemporal consistency with monotonically increasing time, the loss function L=L is used. cross_attn +λ⋅L monotonic (λ is recommended to be 0.1) Optimize the attention weights to ensure that audio frames and video frames correspond precisely on the timeline, where L... cross_attn For cross-attention reconstruction loss, L monotonic This represents the time-constrained loss.

[0095] Based on spatiotemporal alignment A align and V align Deep semantic complementarity enhancement is achieved through a multimodal gating unit to generate audio-video alignment features. The multimodal gating unit performs real-time assessment of the confidence level of single-modal features. When any single-modal feature becomes ambiguous (e.g., environmental noise blurs audio features, or video occlusion causes visual feature loss), a high-confidence feature from another modality is used to complement and correct the ambiguous feature, achieving cross-modal semantic mutual verification. The corrected A is then... align With V align By adding elements together, we obtain audio and video alignment features that simultaneously contain accurate physical synchronization information and semantic information after mutual verification, providing a high-dimensional, high-information-density feature foundation for subsequent feature decomposition.

[0096] Finally, modal decomposition is performed on the audio-video alignment features to obtain environmental features and subject modal features. Subject modal features extract the core effective information of the audio-video interaction, filtering out feature dimensions related to the subject's pronunciation actions, lip movements, and speech rhythm from the alignment features, and integrating them into subject modal features, while retaining the core semantic and spatiotemporal information required for error localization and cut point selection. Environmental features extract background information unrelated to the subject's behavior from the alignment features. Environmental features include audio environmental features and video image environmental features. Audio environmental features focus on the degree of environmental noise interference in the audio signal, including signal-to-noise ratio (SNR) and spectral flatness. SNR characterizes audio purity and is the ratio of signal power to noise power; a smaller value indicates a noisier environment. Spectral flatness characterizes the audio spectrum distribution characteristics and is the ratio of the geometric mean to the arithmetic mean of the spectrum amplitude; a value closer to 1 indicates a flatter spectrum (noise feature), and a value closer to 0 indicates a steeper spectrum (speech feature). The environmental features of video footage focus on the overall stability of the image and the degree of environmental interference, including image clutter (e.g., represented by image texture entropy) and motion intensity (e.g., represented by facial keypoint displacement variance). Image texture entropy characterizes image clutter and is an indicator of the uniformity of image texture distribution; a higher value indicates a more chaotic image texture (e.g., dynamic clutter in the background). Facial keypoint displacement variance characterizes motion intensity; a higher value indicates more significant displacement of facial keypoints in adjacent frames, meaning more intense movement of the person or image. Accurate separation of these two types of features is achieved, separating core features from redundant background features, thus improving the efficiency of downstream tasks in utilizing effective features.

[0097] In S42, the main modal features include the normalized 2D coordinate vectors of facial key points in each video frame and the 3D Euler angle vector of the head. The normalized 2D coordinate vectors reflect the micro-expression positions of the facial key points. These facial key points can form a 68-point or 468-point facial network. Based on the normalized 2D coordinate vectors of the facial key points and the 3D Euler angle vector of the head, a video frame continuity cost function is constructed. The expression of the video frame continuity cost function is:

[0098]

[0099] In the formula, The cost function representing the continuity of video frames; The time index of the candidate cut point is represented, and the search domain of t is the range of speech error fragment extraction. Normalized weighting coefficients representing micro-level characteristics; This represents the total number of facial landmarks. An index representing facial landmarks; This represents the Euclidean norm (L2 Norm). This represents the time index of the frame preceding the candidate cut point. ; This represents the time index of the frame following the candidate cut point. ; Indicates the first The first video frame at time [time] Normalized two-dimensional coordinate vectors of key facial features; Indicates the first The first video frame at time [time] Normalized two-dimensional coordinate vectors of key facial features; Normalized weighting coefficients representing macroscopic posture; This represents the Manhattan norm (L2 Norm). Indicates the first The three-dimensional Euler angle vector of the head in the video frame at a given moment; Indicates the first The three-dimensional Euler angle vector of the head in the video frame at time t. The normalized two-dimensional coordinate vector can be uniformly represented as: , This feature primarily captures subtle shifts in micro-expressions such as mouth closure and eyelid opening / closing, ensuring no sudden changes in eye closure or mouth opening before or after cropping; the three-dimensional Euler angle vector can be uniformly represented as... , Indicates pitch angle. Indicates the yaw angle. Indicates the roll angle. Indicates transpose. This is used to punish violent head rotations, ensuring a smooth transition in the speaker's facial orientation before and after the cut.

[0100] The video continuity cost function is used to quantify the difference in facial topology between two frames before and after a candidate cut point. To avoid the limitations of a single indicator, this scheme constructs a hybrid metric model that includes microscopic facial feature deformation and macroscopic head posture. The system calculates the weighted geometric distance between the frame before and after the cut point as the cost of the video continuity dimension.

[0101] In the above scheme, by introducing a visual spatiotemporal feature evaluation mechanism based on lip closure, head posture Euler angles, and movement amplitude, it can actively avoid areas of abrupt motion change during the cut point decision process. This effectively maintains the subject's posture, lip shape, and spatiotemporal continuity before and after the cut frame, significantly reducing visual abrupt changes and scene jumps caused by cutting, and improving the visual consistency of the editing result. For example, in traditional editing, the cut point falls on the instant the speaker looks up. In the previous frame, the speaker's gaze is still downward, but in the next frame, the gaze has turned to the camera, causing a significant abrupt shift in position. However, in this scheme, the cut point automatically moves to the frame where the speaker's head movement ends and the head is stable. The speaker's head position is basically consistent in the previous and next frames, significantly reducing the visual jump. As another example, in traditional editing, the cut point falls on the speaker turning their head. In the previous frame, the speaker is still facing the left side of the screen, but in the next frame, due to the cut jump, the speaker has turned to face forward, resulting in an abrupt change in head posture and disrupting the continuity of movement. During editing, this solution uses spatiotemporal feature analysis to identify that the angular velocity during the head-turning action is relatively large. It automatically moves the cut point to the area where the head-turning action is completed and the head posture is stable, so that the speaker's face is facing the same direction in the previous and next shots, forming a continuous visual transition without abrupt changes, thus improving the spatiotemporal consistency of the cut boundary.

[0102] The main modal features also include the amplitude values ​​of audio sampling points. Based on the amplitude values ​​of audio sampling points, an audio physical feature stability cost function is constructed. The expression of the audio physical feature stability cost function is as follows:

[0103]

[0104] In the formula, This represents the cost function for the stability of audio physical characteristics; This represents the weighting coefficient that adjusts the absolute value of the amplitude; Indicates the first The amplitude value of the audio sampling point at a given moment; Indicates the first The absolute value of the amplitude at the audio sampling point at any given time; Weighting coefficients representing the importance of energy in the regulation region; The window size represents the short-time energy calculation, which is the total number of audio sampling points involved in the calculation (e.g., the number of audio sampling points included in a 20ms duration). It is a fixed window length parameter that determines the time range covered when calculating short-time energy. Indicates in window The index variable for the internal audio sampling points, with a value range from 0 to K-1, is used to traverse each audio sampling point within the window and calculate the sum of squares of the amplitude values ​​of the first K sampling points in turn. Indicates the first The amplitude value of the audio sampling point at a given moment; Indicated by Let K be the short-time energy with a window length of K at the endpoint. The closer this value is to 0, the closer the current point is to the zero-crossing point. Cutting at this point can minimize the spectral leakage caused by signal truncation. The lower the value, the more likely the time period is in a silent or low-energy area of ​​speech (such as pauses between sentences), making it suitable for seamless splicing.

[0105] The audio physical characteristic stability cost function combines instantaneous amplitude and short-term energy, and can find the moment when the waveform amplitude is the smallest and the energy is the weakest as the shear point to eliminate the popping sound or phase discontinuity caused by forcibly cutting off at the peak.

[0106] The main modal features also include the fundamental frequency and short-time energy envelope of the audio. Based on the fundamental frequency and short-time energy envelope of the audio, a prosodic consistency cost function is constructed. The expression of the prosodic consistency cost function is as follows:

[0107]

[0108] In the formula, Represents the prosodic consistency cost function; Indicates the first The fundamental frequency of the audio at any given moment, measured in Hz; Indicates the first The fundamental frequency of the audio at any given moment; Indicates the first The short-time energy envelope of the moment audio; Indicates the first The short-time energy envelope of the temporal audio. Used to measure the degree of pitch change before and after cutting; Used to measure the degree of sudden change in volume before and after cutting.

[0109] The prosodic consistency cost function measures prosodic differences by comparing changes in fundamental frequency (F0) and energy envelope. It is used to reflect the continuity of speech prosody before and after a candidate cut point. When the cost is large, it indicates that the cut point may cause intonation breaks or rhythmic abrupt changes, which is not conducive to the natural connection of synthesized audio.

[0110] The above scheme utilizes prosodic features such as fundamental frequency variation trends, energy envelope characteristics, and breathing rhythm patterns to automatically identify natural pauses and prosodic stability areas in audio signals. Based on this multimodal prosodic consistency analysis, the system can avoid placing the cut point on long vowel stretching segments, abrupt intonation changes, or inhalation phases, ensuring continuity in rhythm, strength, and intonation before and after the cut, thereby improving the natural listening experience and speech smoothness of the edited audio. For example, in traditional editing, the cut point might fall on a rapid vocal segment within the speaker's sentence, where the previous sentence is still in an upward intonation and the energy peak hasn't fallen, while the next audio segment starts with a flatter intonation due to the different context, creating a momentary break in pitch and energy. In this scheme, the system detects a significant upward trend in the fundamental frequency variation curve within fast-paced segments, making it unsuitable for cutting. Therefore, it automatically moves the cut point to the next natural pause where the speech slows down and the energy envelope stabilizes, maintaining continuity in intonation and energy levels between the two segments, enhancing the prosodic continuity of the audio at the cut point, and resulting in a more natural listening experience.

[0111] In S43, refer to Figure 4 The flowchart illustrating the method for obtaining the weight coefficients of the cost function shows that "obtaining the weight coefficients of each cost function based on environmental characteristics" includes the following steps:

[0112] S431: Combine environmental features by category to generate multi-dimensional environmental splicing features;

[0113] S432: Input the multidimensional environment splicing features into the lightweight fully connected network and obtain the weight coefficients of each cost function output by the lightweight fully connected network.

[0114] In S431, environmental features include image texture entropy, facial keypoint displacement variance, signal-to-noise ratio, and spectral flatness. These four types of features are stitched together to obtain 4-dimensional environmental stitching features.

[0115] In S432, the 4D environment stitching features are input into a lightweight fully connected network. This network, after training, learns the mapping relationship between environment features and optimal weights, outputting the weight coefficients for each cost. The lightweight fully connected network is a neural network with an extremely simple structure, consisting of a few layers of data processing units (fully connected layers) linearly stacked. It does not contain complex image recognition or temporal processing structures (such as convolutional layers or recurrent layers). In this scheme, the lightweight fully connected network acts as an intelligent parameter decision-maker, specifically including an input layer, hidden layers, and an output layer. The input layer receives normalized 4D environment stitching features; the hidden layer has M (e.g., 16) dimensional neurons, uses the ReLU activation function to capture the nonlinear correlation between stitching features and achieves a slight dimensionality increase from 4D to MD, outputting a high-dimensional nonlinear feature vector; the output layer receives the M-dimensional high-dimensional feature vector output from the hidden layer, maps it to 3D space through a linear transformation, and then uses the Softmax activation function to normalize the 3D vector, finally outputting the weight coefficients for the three costs (the sum of the weights is 1).

[0116] The training process of the lightweight fully connected network adopts a two-stage training strategy of self-supervised pre-training combined with fine-tuning with a small number of real data samples. First, the basic capabilities of the model are pre-trained using a large amount of synthetic mis-speech data. Then, a small amount of real mis-speech data is used for fine-tuning and optimization to make the model adapt to the environmental features and cost function weight coefficient mapping relationship in real scenarios. The specific training process is as follows.

[0117] First, self-supervised pre-training is conducted. A large number of readily available, fluent spoken video sequences are used as positive samples. A synthetic dataset containing simulated speech errors is constructed through randomization operations generated by reverse speech error generation, providing sufficient training samples for the model. The original fluent video sequence is denoted as X={(F1,A1),(F2,A2),…,(F… j A j ),…,(F n A n )}, where n is the total number of frames in the video sequence, j is the time step index of the sequence, and F j For the video image frame corresponding to the j-th time step, A j The audio signal segment corresponding to the j-th time step and related to F j Display time alignment. Perform multi-dimensional randomization on the original sequence to generate synthetic data: with probability p r Perform repeated operations on video frames ∈ [0.05, 0.15] with probability p. d ∈[0.05,0.15] perform a discard operation to simulate stuttering and frame skipping errors; perform non-linear randomized splicing of video frames within a random window of length U∈[3,7]: [F c ,…,F c+U→Shuffle(⋅), simulates a slip of the tongue due to disordered word order or logic, where c is the index of the starting frame of the out-of-order window, [F c ,…,F c+U [] represents a continuous video frame sequence of length U starting from frame c, → indicates performing a Shuffle operation on the left-hand frame sequence, and Shuffle(⋅) indicates randomly rearranging the input continuous U video frames; for audio signal A j Injection satisfies η∼N(0,σ) 2 Additive white Gaussian noise is used to obtain the audio segment A after noise injection. j =A j +η, where the noise standard deviation σ∈[0.001,0.01], is used to simulate audio-related slips of the tongue such as breathing sounds and environmental noise. Based on this synthetic dataset, the model is self-supervised pre-trained to learn the basic feature representation capabilities in slip-of-the-tongue scenarios.

[0118] After pre-training, fine-tuning is performed using a few samples based on real slip-up data to compensate for the feature differences between synthetic data and real-world scenarios. A dataset of thousands of video clips containing real slip-ups is collected as the fine-tuning dataset. This dataset contains real slip-up features that are difficult to simulate with synthetic data, such as micro-expression cues like shifty eyes and frowning before and after a slip-up, and prosodic hesitation features like prolonged sounds, incomplete pronunciation, and filler words. The pre-trained model weights are loaded into a lightweight fully connected network, and transfer learning techniques such as low-rank adaptation (LoRA) fine-tuning or full-parameter fine-tuning are used to fine-tune the model using the real slip-up dataset, enabling the model to learn the correlation between slip-up features and environmental features in real-world scenarios.

[0119] Finally, supervised training was conducted for the core task of the model to achieve accurate mapping from multi-dimensional environment splicing features to cost function weight coefficients. A labeled environment feature vector sample set was constructed, containing multi-dimensional environment splicing features for various speech error scenarios, as well as corresponding cost function weight coefficient labels. Using this sample set as training data, supervised training was performed on the pre-trained and fine-tuned lightweight fully connected network. During training, mean squared error loss (MSE) was used as the loss function, and the Adam optimizer was used to iteratively optimize the network parameters until the model converged, enabling the network to accurately output the corresponding weight coefficients of each cost function based on the input multi-dimensional environment splicing features.

[0120] Traditional editing methods rely on fixed parameter decisions, making it difficult to adapt to environmental changes such as background noise, microphone type, and the speaker's movements. This results in poor robustness and generalization, hindering the consistent output of high-quality edits. This new approach introduces a lightweight, fully connected network to dynamically and adaptively adjust the weights of the multimodal cost function. By real-time sensing of the environmental noise level (signal-to-noise ratio) and the intensity of the subject's motion (keypoint displacement) in the video clip, the system intelligently determines which modal feature is more reliable in the current environment and adjusts its weight accordingly. This fundamentally solves the problem of the drastic decrease in robustness of traditional fixed-weight models in varying shooting environments (such as noisy environments or large movements), ensuring that the algorithm can stably select the optimal cut point in various scenarios and significantly improving the generalization and quality stability of the editing results.

[0121] For example, in traditional editing, fixed high weights are used to evaluate audio features. In high-noise environments such as outdoors or in cafes, background noise (non-human voice) is misidentified as speech pauses or slips of the tongue, leading to incorrect calculation of audio feature costs. The system may therefore incorrectly decide on the cut point in a region with high background noise but continuous speech, causing abrupt interruptions in the speech flow. This solution, however, can handle high-noise environments. A lightweight, fully connected network detects low audio signal-to-noise ratios and determines that audio features are unreliable, automatically reducing the weight of audio physical feature costs (e.g., β). The decision module then assigns higher weights to reliable visual and prosodic features, thus ignoring the interference of background noise on audio physical features and accurately locating the cut point in a true natural pause or low visual abruptness region, ensuring editing accuracy. For another example, in traditional editing, when the speaker turns their head significantly, due to insufficient weighting of visual features, the system may prioritize cutting from the point with the smoothest audio prosody. If this point happens to be located on a frame with drastic changes in head angular velocity, the edited image will appear to have a noticeable momentary shift. This solution can handle environments with high motion intensity. The lightweight fully connected network detects that the displacement variance of facial key points is high (i.e., high motion intensity), and determines that the visual continuity risk is high during this period. It automatically and significantly increases the weight of the visual continuity cost (such as alpha). At this time, the decision module will avoid the frame with the most intense motion and select the cut point after the motion transition or after the action is completed. This effectively eliminates the sense of screen jump caused by large motion amplitude.

[0122] In S44, all cost functions are normalized so that their values ​​are in the interval [0,1]. Based on the weight coefficient of each cost function, all normalized cost functions are weighted and summed to obtain a multi-dimensional cost synthesis function. The expression of the multi-dimensional cost synthesis function is as follows:

[0123]

[0124] In the formula, This represents the multi-dimensional cost synthesis function, i.e., the first... The overall cost score of the candidate cut point at time step; the lower the score, the more suitable the candidate cut point is for editing. These represent the weighting coefficients of the video frame continuity cost function; The weighting coefficients represent the cost function for the stability of audio physical characteristics; The weight coefficients represent the prosodic consistency cost function. .

[0125] In S5 and S6, when the multi-dimensional cost synthesis function is obtained, it is solved directly, and the one with the lowest comprehensive cost score is selected as the optimal shearing point from multiple candidate shearing points.

[0126] The video frames before and after the optimal cut point are compared to obtain the visual difference value. It is then determined whether the visual difference value is greater than a preset threshold. If it is greater than the preset threshold (e.g., the head must turn), the optical flow frame interpolation module is triggered. This module uses a deep neural network to calculate the dense optical flow field of the reference frames before and after the cut, generating 1-3 non-linear deformed frames (transition frames). The pixel features of the previous frame are smoothly distorted to the position of the next frame to achieve seamless healing. If the difference is not greater than the preset threshold, the frames are directly and seamlessly stitched together. After stitching is completed, the repaired smooth video file is output.

[0127] In summary, this solution solves the visual abruptness problem caused by traditional "hard cuts," minimizing the difference in lip shape and head pose before and after the cut point. It also solves the problem of discontinuous rhythm (intonation / rhythm) at the cut point, ensuring the natural fluency of the synthesized audio. Furthermore, it addresses the lack of robustness and generalization of fixed-weight decision models in varying environments, achieving adaptive optimization for different environmental conditions (such as background noise levels and subject movement intensity), ensuring that the most reliable modal features are automatically selected for cut point decisions in various shooting scenarios.

[0128] Reference Figure 5 This disclosure provides a system for intelligently correcting verbal slips in videos, comprising:

[0129] The file parsing module 101 is used to parse the video file to be repaired into an audio stream file and a video stream file.

[0130] The segment extraction module 102 is used to extract speech error segments from audio stream files and video frame stream files to obtain audio speech error segments and video frame speech error segments.

[0131] The feature extraction module 103 is used to extract features from audio and video clips containing verbal slips, and to obtain audio and video features.

[0132] The function construction module 104 is used to construct a multi-dimensional cost synthesis function based on audio features and video image features, wherein the dimensions include at least one of video image continuity, audio physical feature stability and prosodic consistency;

[0133] Function solving module 105 is used to solve the multi-dimensional cost synthesis function and obtain the optimal shear point;

[0134] The video repair module 106 is used to repair video files based on the optimal cut points and generate new video files.

[0135] The various variations and specific examples of the intelligent video correction method provided above are also applicable to the intelligent video correction system provided in this disclosure. Through the foregoing detailed description of the intelligent video correction method, those skilled in the art can clearly understand the implementation method of the intelligent video correction system. For the sake of brevity, it will not be described in detail here.

[0136] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0137] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the methods for intelligently correcting verbal slips in videos described in the foregoing embodiments of this disclosure.

[0138] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0139] like Figure 6 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 6The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0140] like Figure 6 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0141] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 6 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.

[0142] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the method for intelligently correcting slips of the tongue in a video according to embodiments of this disclosure are performed.

[0143] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0144] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the methods for intelligently correcting verbal slips in videos described in the foregoing embodiments of the present disclosure are performed.

[0145] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0146] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0147] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0148] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.

[0149] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0150] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0151] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0152] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0153] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for intelligently correcting verbal slips in videos, characterized in that, include: The video file to be repaired is parsed into an audio stream file and a video stream file; Extract speech error segments from the audio stream file and the video stream file to obtain audio speech error segments and video speech error segments; Feature extraction is performed on the audio and video clips containing verbal slips to obtain audio and video features. Based on the audio features and the video features, a multi-dimensional cost synthesis function is constructed, wherein the dimensions include video continuity, audio physical feature stability, and prosodic consistency. The step of constructing a multi-dimensional cost synthesis function based on the audio features and the video image features includes: Spatiotemporal alignment and semantic complementarity are performed on the audio features and video features to obtain audio-video alignment features, and the audio-video alignment features are then decomposed into environmental features and subject modality features. Based on the aforementioned subject modal features, a multi-dimensional cost function is constructed, wherein the cost function includes a video frame continuity cost function, an audio physical feature stability cost function, and a prosodic consistency cost function; Based on the aforementioned environmental characteristics, the weight coefficients of each cost function are obtained; Based on the cost function and the corresponding weight coefficients, a multi-dimensional cost synthesis function is constructed; Solve the multi-dimensional cost synthesis function to obtain the optimal shear point; Based on the optimal cutting point, the video file is repaired to generate a new video file.

2. The method for intelligently correcting verbal slips in videos according to claim 1, characterized in that, The step of extracting speech error segments from the audio stream file and the video stream file to obtain audio speech error segments and video speech error segments includes: Perform speech recognition on the audio stream file to obtain the text. Perform semantic analysis on the text to determine the start and end time range of the slip of the tongue segment; According to the start and end time range, the audio stream file and the video stream file are used to extract verbal slips to obtain audio verbal slips and video verbal slips.

3. The method for intelligently correcting verbal slips in videos according to claim 1, characterized in that, The main modal features include the normalized two-dimensional coordinate vector of facial key points in each video frame and the three-dimensional Euler angle vector of the head; Based on the normalized two-dimensional coordinate vector of the facial key points and the three-dimensional Euler angle vector of the head, a video frame continuity cost function is constructed. The expression for the video frame continuity cost function is: In the formula, The cost function representing the continuity of video frames; The time index representing the candidate shear point; Normalized weighting coefficients representing micro-level characteristics; This represents the total number of facial landmarks. An index representing facial landmarks; Denotes the Euclidean norm; Indicates the time index of the frame preceding the candidate cut point; Indicates the time index of the frame following the candidate cut point; Indicates the first The first video frame at time [time] Normalized two-dimensional coordinate vectors of key facial features; Indicates the first The first video frame at time [time] Normalized two-dimensional coordinate vectors of key facial features; Normalized weighting coefficients representing macroscopic posture; Represents the Manhattan norm; Indicates the first The three-dimensional Euler angle vector of the head in the video frame at a given moment; Indicates the first The three-dimensional Euler angle vector of the head in the video frame at a given time.

4. The method for intelligently correcting verbal slips in videos according to claim 3, characterized in that, The main modal features also include the amplitude values ​​of audio sampling points; Based on the amplitude values ​​of the audio sampling points, an audio physical feature stability cost function is constructed; The expression for the audio physical feature stability cost function is: In the formula, This represents the cost function for the stability of audio physical characteristics; This represents the weighting coefficient that adjusts the absolute value of the amplitude; Indicates the first The amplitude value of the audio sampling point at a given time; Indicates the first The absolute value of the amplitude at the audio sampling point at any given time; Weighting coefficients representing the importance of energy in the regulation region; This indicates the window size for short-time energy calculations; Indicates in window Index variable for internal audio sampling points; Indicates the first The amplitude value of the audio sampling point at a given time.

5. The method for intelligently correcting verbal slips in videos according to claim 4, characterized in that, The main modal features also include the fundamental frequency of the audio and the short-time energy envelope of the audio; Based on the fundamental frequency and short-time energy envelope of the audio, a prosodic consistency cost function is constructed; The expression for the prosodic consistency cost function is: In the formula, This represents the prosodic consistency cost function; Indicates the first The fundamental frequency of the moment audio; Indicates the first The fundamental frequency of the moment audio; Indicates the first The short-time energy envelope of the moment audio; Indicates the first The short-time energy envelope of the moment audio.

6. The method for intelligently correcting verbal slips in videos according to claim 1, characterized in that, The step of obtaining the weight coefficients of each cost function based on the environmental characteristics includes: The environmental features are then categorized and spliced ​​together to generate multidimensional environmental splicing features. The multidimensional environment splicing features are input into a lightweight fully connected network to obtain the weight coefficients of each cost function output by the lightweight fully connected network.

7. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method for intelligently correcting verbal slips in videos as described in any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method for intelligently correcting verbal slips in a video as described in any one of claims 1-6.

9. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method for intelligently correcting verbal slips in videos as described in any one of claims 1-6.