A method, apparatus, device and medium for repairing multi-lens video
By performing multimodal information analysis and generating state constraint rules on multi-shot videos, the problem of state discontinuity across shots is identified and repaired, thereby improving the narrative continuity and quality of multi-shot videos.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHUOZHUO TECH
- Filing Date
- 2026-07-02
- Publication Date
- 2026-07-31
AI Technical Summary
AI video generation models often produce multi-camera videos with discontinuous states across shots.
By performing multimodal information analysis on each shot in a multi-camera video, the narrative state is determined. Based on the narrative state of the preceding shots and the action events, state constraint rules are generated to identify defect types and to repair them using corresponding methods, including state consistency, event supplementation, narrative explanation, and editing and structural adjustment repair.
It effectively detects and repairs the discontinuity of cross-shot states in multi-shot videos, improving the narrative continuity and quality of the videos.
Smart Images

Figure CN122492702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for repairing multi-camera videos. Background Technology
[0002] With the development of artificial intelligence (AI) video generation technology, users can input text prompts, character reference images, scene reference images, etc. into AI video generation models to obtain multi-camera videos such as short dramas and advertisements output by AI video generation models.
[0003] Currently, single-shot videos output by AI video generation models are generally of high quality. However, multi-shot videos output by AI video generation models are prone to issues with discontinuity across shots. See also Figure 1 The image is a schematic diagram of a multi-camera video. Figure 1 The discontinuous state shown across shots refers to the following: the man has already opened the door and left the room in shot 1, but then reappears in the room in shot 2 without explanation.
[0004] Therefore, how to detect the cross-shot state of multi-shot videos and repair multi-shot videos when the cross-shot state is discontinuous has become an urgent technical problem to be solved. Summary of the Invention
[0005] To address the aforementioned issues, this application provides a method, apparatus, device, and medium for repairing multi-camera videos, capable of detecting the cross-camera state of multi-camera videos and repairing multi-camera videos when the cross-camera state is discontinuous.
[0006] The embodiments of this application disclose the following technical solutions: In a first aspect, this application discloses a method for restoring multi-camera videos, the method comprising: Acquire multi-camera video; Multimodal information analysis is performed on each shot in the multi-shot video to determine the narrative state of each shot; the narrative state is used to characterize the character state, environmental state, and temporal state in the shot; each shot includes a preceding shot and a target shot; the preceding shot is scheduled before the target shot. Based on the narrative state of the preceding shot and the action or plot events that occur in the preceding shot, determine the state constraint rules for the target shot. When the narrative state of the target shot does not satisfy the state constraint rule, and there is no explanatory event to release the state constraint rule, the defect type of the target shot is determined. Based on the defect type, a target repair method is determined from a set of preset repair methods; The target lens is repaired using the aforementioned target repair method.
[0007] Optionally, the step of performing multimodal information analysis on each shot in the multi-shot video to determine the narrative state of each shot includes: Based on the video production auxiliary information of the multi-camera video, a key entity table is constructed; the video production auxiliary information includes at least one of storyboard script, prompts, character settings, prop settings, and scene settings. For each shot in the multi-camera video, the target object is determined based on the visual feature similarity between each object in the shot and each key entity in the key entity table; Multimodal information analysis is performed on each target object to determine the narrative state of each target object.
[0008] Optionally, the explanatory events include at least one of the following: time jump events, memory events, dream events, fantasy events, montage events, explicit action events, plot explanation events, occlusion explanation events, camera direction change events, scene switching events, and automatic state change events.
[0009] Optionally, the preset repair method set includes state consistency repair, event supplementation repair, narrative explanation repair, and editing and structural adjustment repair; The method of repairing the target lens using the target repair method includes: When the target repair method is state consistency repair, the character state, environment state or time state in the target shot is adjusted by any one of the following methods: local redraw, local completion, region replacement and regeneration. When the target repair method is event-supplementary repair, a transition shot is inserted between the preceding shot and the target shot, or an action segment is added to the target shot; the transition shot or the action segment is used to explain the changes in character state, environmental state, or time state in the target shot; When the target restoration method is narrative interpretation restoration, any one of the following narrative interpretations is added to the target shot: subtitles, dialogue text, time markers, transition markers, and script descriptions; the narrative interpretation is used to explain the changes in character state, environmental state, or time state in the target shot; When the target repair method is editing and structural adjustment repair, the order of the target shot and its adjacent shots is adjusted, or the transition method between the target shot and its adjacent shots is changed.
[0010] Optionally, the step of repairing the target lens using the target repair method includes: Determine the continuity defect score of the target shot; the continuity defect score is related to the degree of matching between the narrative state of the target shot and the state constraint rule; The defect severity level of the target lens is determined based on the continuity defect score. When the number of lenses to be repaired exceeds a preset threshold, the lenses to be repaired are repaired sequentially according to the degree of defect.
[0011] Secondly, this application discloses a multi-lens video restoration device, the device comprising: a lens acquisition module, a state determination module, a rule determination module, a type determination module, a method determination module, and a lens restoration module; The lens acquisition module is used to acquire multi-lens video; The state determination module is used to perform multimodal information analysis on each shot in the multi-shot video to determine the narrative state of each shot; the narrative state is used to characterize the character state, environmental state, and temporal state in the shot; each shot includes a preceding shot and a target shot; the preceding shot is scheduled before the target shot. The rule determination module is used to determine the state constraint rules for the target shot based on the narrative state of the preceding shot and the action events or plot events that occur in the preceding shot. The type determination module is used to determine the defect type of the target shot when the narrative state of the target shot does not satisfy the state constraint rule and there is no explanatory event for releasing the state constraint rule. The method determination module is used to determine the target repair method from a preset repair method set according to the defect type; The lens repair module is used to repair the target lens using the target repair method.
[0012] Optionally, the state determination module is specifically used for: constructing a key entity table based on the video production auxiliary information of the multi-shot video; the video production auxiliary information includes at least one of storyboard script, prompts, character settings, prop settings, and scene settings; for each shot in the multi-shot video, determining a target object based on the visual feature similarity between each object in the shot and each key entity in the key entity table; and performing multimodal information analysis on each target object to determine the narrative state of each target object.
[0013] Optionally, the explanatory events include at least one of the following: time jump events, memory events, dream events, fantasy events, montage events, explicit action events, plot explanation events, occlusion explanation events, camera direction change events, scene switching events, and automatic state change events.
[0014] Optionally, the preset repair method set includes state consistency repair, event supplementation repair, narrative explanation repair, and editing and structural adjustment repair; The shot repair module is specifically used for: when the target repair method is state consistency repair, adjusting the character state, environment state, or time state in the target shot through any one of the following methods: local redraw, local completion, region replacement, and regeneration; when the target repair method is event supplement repair, inserting a transition shot between the preceding shot and the target shot, or supplementing the target shot with action segments; the transition shot or the action segments are used to explain the changes in character state, environment state, or time state in the target shot; when the target repair method is narrative explanation repair, supplementing the target shot with any one of the following narrative explanations: subtitles, dialogue text, time markers, transition markers, and script descriptions; the narrative explanations are used to explain the changes in character state, environment state, or time state in the target shot; when the target repair method is editing and structural adjustment repair, adjusting the sequence between the target shot and its adjacent shots, or changing the transition method between the target shot and its adjacent shots.
[0015] Optionally, the lens repair module is specifically used to: determine the continuity defect score of the target lens; the continuity defect score is related to the degree of matching between the narrative state of the target lens and the state constraint rule; determine the defect severity level of the target lens based on the continuity defect score; when the number of lenses to be repaired is greater than a preset number threshold, repair the lenses to be repaired sequentially according to the defect severity level.
[0016] Thirdly, this application discloses a multi-lens video restoration device, the device comprising: a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement the various steps of the multi-lens video restoration method as described in the first aspect.
[0017] Fourthly, this application discloses a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of the multi-lens video restoration method as described in the first aspect.
[0018] Compared with the prior art, this application has the following beneficial effects: This application provides a method, apparatus, device, and medium for repairing multi-shot videos. By determining the narrative state of each shot in a multi-shot video and generating state constraint rules for a target shot based on the narrative state of the preceding shot and the action or plot events occurring in the preceding shot, this application can detect cross-shot states in multi-shot videos. Furthermore, when the narrative state of a target shot does not satisfy the state constraint rules for that target shot, and there is no explanatory event to release the state constraint rules, this application determines the defect type of the target shot and repairs the target shot according to the target repair method corresponding to the defect type, thereby enabling the repair of multi-shot videos when cross-shot states are discontinuous. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of a multi-camera video. Figure 2 A flowchart illustrating a method for repairing multi-camera video provided in this application embodiment; Figure 3 A schematic diagram of a multi-lens video restoration device provided in an embodiment of this application; Figure 4 This is a schematic diagram of a computer-readable medium provided in an embodiment of this application. Detailed Implementation
[0021] As described earlier, single-shot videos output by AI video generation models are generally of high quality. However, multi-shot videos output by AI video generation models are prone to issues with discontinuity across shots. For example... Figure 1 The discontinuous state shown across shots refers to the following: the man has already opened the door and left the room in shot 1, but then reappears in the room in shot 2 without explanation.
[0022] Through research, the inventors have proposed a method, apparatus, device, and medium for repairing multi-shot videos. This application determines the narrative state of each shot in a multi-shot video and generates state constraint rules for the target shot based on the narrative state of the preceding shot and the action or plot events that occurred in the preceding shot. This enables the detection of cross-shot states in multi-shot videos. Furthermore, when the narrative state of the target shot does not satisfy the state constraint rules for the target shot, and there is no explanatory event to release the state constraint rules, this application determines the defect type of the target shot and repairs the target shot according to the target repair method corresponding to the defect type. This allows for the repair of multi-shot videos when the cross-shot states are discontinuous.
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0024] See Figure 2 The figure is a flowchart of a multi-camera video restoration method provided in an embodiment of this application. The method includes: S201: Acquire multi-camera video.
[0025] Multi-shot videos refer to videos with at least one shot transition (i.e., videos consisting of two or more shots). Due to shot transitions, changes may occur between shots in terms of character states (e.g., emotions, postures, appearances), environmental states (e.g., opening and closing of doors and windows, prop positions), and temporal states (e.g., time jumps, flashbacks, etc.). Therefore, it is necessary to detect and restore the narrative continuity between shots. It should be noted that "multi-shot" can be formed by editing, splicing, or other editing methods; this application does not limit this.
[0026] In one specific implementation, the multi-camera video can be the video output by the AI video generation model after inputting video production auxiliary information into it. This video production auxiliary information can be at least one of storyboards, text prompts, character settings, prop settings, and scene settings. This application does not limit the scope of this information.
[0027] S202: Perform multimodal information analysis on each shot in the multi-cam video to determine the narrative state of each shot; the narrative state is used to characterize the state of the characters, environment, and time in the shot; each shot includes a preceding shot and a target shot; the preceding shot is scheduled before the target shot.
[0028] In this embodiment, it is not necessary to perform multimodal information analysis on all objects in a multi-shot video. This is because many objects exist only as background objects or non-critical objects, such as passersby in the background, decorative paintings on the wall, and distant vehicles. Performing multimodal information analysis on all objects in a multi-shot video would not only lead to an exponential increase in computational overhead and waste of computational resources, but the random fluctuations of background or non-critical objects could easily introduce false matches, thus affecting the accuracy of subsequent narrative continuity detection. Therefore, this application performs the following steps A1-A3: A1: Construct a key entity table based on the video production auxiliary information of multi-camera videos; the video production auxiliary information includes at least one of storyboard scripts, prompts, character settings, prop settings, and scene settings.
[0029] First, key entities are extracted from the video production support information of multi-camera videos using Natural Language Processing (NLP) technology. Key entities include at least one of key characters, key objects, key scenes, and key scene anchors. Then, each extracted key entity is assigned a unique global identifier (i.e., key entity ID) to construct a key entity table. For example, if the video production support information is a storyboard: "Character A picks up his phone in the living room and leaves the room into the hallway," then the corresponding key entity table can be shown in Table 1 below: Table 1
[0030] A2: For each shot in a multi-camera video, determine the target object based on the visual feature similarity between each object in the shot and each key entity in the key entity table.
[0031] First, for each shot in the multi-shot video, image features of the objects in the shot are extracted using an object detection network, and text features of key entities in the key entity table are extracted using a text feature extraction network. Second, to achieve cross-modal comparison, a pre-trained visual-language large model can map the aforementioned image and text features to a feature space of the same dimension, obtaining object feature vectors and key entity feature vectors respectively. Subsequently, the visual feature similarity between the object feature vector and the key entity feature vector is calculated (e.g., cosine similarity or Euclidean distance). If the similarity is greater than or equal to a preset similarity threshold (e.g., 0.85), the object in that shot is identified as the target object.
[0032] In one example, if the object is a character, the target object can be determined based on the similarity between each object in the shot and each key entity in the key entity table in terms of appearance features (e.g., skin color, body shape), facial features, clothing features, and posture features.
[0033] In another example, if the object is an item, the target object can be determined based on the similarity between each object in the shot and the appearance features (such as shape, color, and texture) of each key entity in the key entity table. Furthermore, for smaller items, a pre-trained human pose estimation model can be invoked, combining the interaction features between the item and the character's hand (such as determining whether the item's two-dimensional coordinates are within the reachable space of the character's hand key points) as an auxiliary similarity criterion to determine the target object.
[0034] It should be noted that if the visual feature similarity between each object in the shot and each key entity in the key entity table is less than the preset similarity threshold, then the object in the shot is identified as a non-target object and marked as identity_uncertain.
[0035] A3: Perform multimodal information analysis on each target object to determine the narrative state of each target object.
[0036] First, acquire multimodal information for each target object. This multimodal information may include at least one of the following: image information, video timing information, audio information, subtitle information, dialogue text, script text, on-screen text, and plot context information.
[0037] Subsequently, multimodal information analysis is performed on each target object based on its multimodal information to determine its narrative state. This multimodal information analysis may include performing person detection, object detection, instance segmentation, and keypoint detection on the target object using a visual detection network; action recognition and interaction relationship recognition using a temporal action network; and subtitle parsing, speech recognition, optical character recognition, and script text alignment using a speech and text model.
[0038] Narrative state is used to represent the state of the characters, the environment, and the time in a shot.
[0039] The character state is used to comprehensively describe the character's behavior and physiological and psychological characteristics in the current shot. Character state includes the character's action state, emotional state, appearance state, and relationship state, as shown in Table 2 below: Table 2
[0040] In one example, if multimodal information analysis of the target object detects that the target object's hand is touching the doorknob and the target object's body is moving towards the exit, then the target object's role action state is determined to be opening the door.
[0041] In another example, if multimodal information analysis of the target object detects that the target object is exhibiting behaviors such as backing away and rapid breathing, and the dialogue includes phrases like "run away" and "don't come any closer," then the target object's emotional state is determined to be either tension or fear.
[0042] The environmental state describes the scene's topological logic, the ownership of key items, and the physical open / closed attributes of scene anchor points. The environmental state includes the environmental location state, the environmental object state, and the current environmental state, as shown in Table 3 below: Table 3
[0043] Among them, time state is used to characterize the narrative time relationship between the current target shot and the previous shot, such as the same time, time jump, memory, dream, montage, etc. Specifically, time state can be determined by comprehensively analyzing time prompts in subtitles, dialogue, and scripts, recognizing transition methods (such as black screen, white screen), extracting visual style changes (such as color tone changes, blurring effects), and recognizing text from optical character recognition on the screen.
[0044] In one example, if a black screen transition appears in the shot along with the text "A few days later," then the time state is determined to be a time jump.
[0045] In another example, if the scene is clearly blurred and accompanied by the dialogue "I remembered that time", then the time state is determined to be a memory.
[0046] S203: Based on the narrative state of the preceding shot and the action or plot events that occurred in the preceding shot, determine the state constraint rules for the target shot.
[0047] Multi-shot videos include prequel shots and target shots. The prequel shots are ordered before the target shots on the narrative timeline or playback timeline (they do not have to be physically adjacent, as long as they are in the same narrative segment).
[0048] For any target shot to be verified, a state transition graph is generated based on the narrative state of its preceding shots and the action or plot events that occurred in those preceding shots. Based on this state transition graph, the state constraint rules for the target shot can be determined.
[0049] At the computer data structure level, this state transition diagram can be mathematically represented as a triple: G=(V,E,C). Where V is the set of state nodes, E is the set of state transition edges, and C is the set of state constraint rules.
[0050] State nodes are used to structurally represent a specific state value of a key entity within a particular shot or time period, in a specific dimension. For example, state nodes can be as follows: V1 = {entity: person_A, shot:shot_001, dimension: spatial_state, value: inside_room}, which indicates that person A's spatial state in shot 001 is "inside the room"; V2 = {entity: phone_1, shot: shot_002, dimension:object_state, value: held_by_person_A}, which indicates that phone 1's object state in shot 002 is "held by person A".
[0051] State transition edges are used to represent the transition or driving relationship between two state nodes. When a specific event is detected by an action recognition network or dialogue semantic analysis, a corresponding transition edge is created in the graph. For example, state transition edges can be as follows: E1 = exit_room(person_A, room_1), indicating that person A leaves the room; E2 = enter_room(person_A, room_1), indicating that person A enters the room; E3 = pick_up(person_A, phone_1), indicating that person A picks up their phone.
[0052] The set of state constraint rules is obtained through topological reasoning of state nodes and state transition edges using a pre-defined logic compilation engine. The logic compilation engine has a pre-built base of common-sense rules about the physical world (e.g., "You cannot remain in a room after leaving it," "An item should be held in a holding state from the time it is picked up until it is put down"). When a state transition edge from a previous shot is detected, the logic compilation engine queries the corresponding state node in the target shot based on the action semantics of that state transition edge. If the state value of that state node violates the expected state value derived from that state transition edge, then the corresponding constraint rule is generated.
[0053] In one example, if a state transition edge E = exit_room(person_A, room_1) is detected in the preceding shot, meaning person A has left room 1, the corresponding generated state constraint rule is: Constraint_Spatial: Person_A NOT IN room_1. This state constraint rule indicates that in subsequent target shots, person_A's legal spatial state should not be within room_1 again until enter_room (re-entering the room), time_jump (time jump to the next day, etc.) is detected, or a specific script interpretation is provided through dialogue. The underlying narrative causal reasoning logic for generating this rule is that the spatial transfer of physical entities is continuous; if person A leaves a specific scene, there must be a legitimate and reasonable return action or time span explanation for them to reappear in the scene in subsequent shots.
[0054] In another example, if a state transition edge E = pick_up(person_A, phone_1) is detected in the preceding shot, meaning person A picks up phone 1, the corresponding generated state constraint rule is: Constraint_Object: phone_1 == held_by_person_A. This state constraint rule indicates that, unless an action event such as put_down, give_to_others, drop / lost occurs, or occlusion occurs (e.g., a hand in a pocket, the item is not visible but logically still held), phone_1 should remain held by person_A, or maintain an interpretable holding continuity with person_A. The generation logic of this rule is that after a character picks up an item, that item should not disappear inexplicably in a subsequent shot without any plot development or action connection.
[0055] S204: When the narrative state of the target shot does not satisfy the state constraint rules and there is no explanatory event to release the state constraint rules, determine the defect type of the target shot.
[0056] When the narrative state of a target shot does not meet the state constraint rules for that shot, it is not immediately considered a continuity defect. Instead, it is further determined whether there is an explanatory event to remove the state constraint rules. An explanatory event refers to an event that can provide reasonable support for the change in narrative state from the perspectives of narrative logic, time progression, shot expression, action process, or scene change. Specifically, the existence of an explanatory event can be determined by identifying the transition effects between video frames, extracting the semantic features of dialogue or subtitles, or detecting specific action features. For example, explanatory events include at least one of the following: time jump events, flashback events, dream events, fantasy events, montage events, explicit action events, plot explanation events, occlusion explanation events, camera direction change events, scene switching events, and automatic state change events. If an explanatory event exists, it indicates that the change in narrative state has narrative rationality and there is no continuity defect; if an explanatory event does not exist, it indicates that the change in narrative state lacks reasonable support and there is a continuity defect.
[0057] When the narrative state of a target shot does not satisfy the state constraint rules for that target shot, and there is no explanatory event to release the state constraint rules, the defect type of the target shot is determined. Defect types include character state continuity defects, environmental state continuity defects, temporal state continuity defects, and object continuity defects.
[0058] A character state continuity defect refers to an unreasonable abrupt change in a character's action state, emotional state, appearance state, or relationship state, which cannot be supported by an explanatory event. For example: Character A is limping with an injured left leg in a previous shot, but is running normally in the target shot without explanation; Character A shows fear and tension in a previous shot, but becomes calm and composed without transition in the target shot; Character A's face is covered in blood in a previous shot, but is clean and tidy without any cleaning process in the target shot; Character A faces Character B in a previous shot, but in the target shot, they are facing completely opposite directions without any turning action.
[0059] Environmental continuity defects refer to unreasonable abrupt changes in a scene or scene anchor point that cannot be supported by an explanatory event. For example: the door is open in the preceding shot, but suddenly closes in the target shot; the light is on in the preceding shot, but goes out inexplicably in the target shot; the window is intact in the preceding shot, but suddenly breaks in the target shot; character A has left room_1 in the preceding shot, but is inexplicably back in room_1 in the target shot; character A is still at the doorway in the preceding shot, but appears directly on the street in the distance in the target shot without any scene transition or time jump indication.
[0060] A temporal continuity defect refers to an unreasonable abrupt change in the narrative time expressed by the target shot and the narrative time expressed by the preceding shot, and this abrupt change cannot be supported by an explanatory event. For example: in the preceding shot, character A is still performing an action, but in the target shot, the long-term result of that action has already appeared, without any time progression indication; in the preceding shot, the environment is still daytime, but in the target shot, it directly switches to nighttime, without any time jump or transition explanation; in the preceding shot, the order of events is clear, but the target shot presents a state that logically occurred earlier, without any flashback explanation.
[0061] Object continuity defects refer to unreasonable abrupt changes in a character's possession, placement, state of existence (appearing or disappearing out of nowhere), or ownership of items, which cannot be supported by an explanatory event. For example: in a previous shot, the phone is held by character A, but in the target shot, the phone appears on the table without explanation; in a previous shot, the action of character A handing a letter to character B has not yet occurred, but in the target shot, the letter is already held by character B.
[0062] S205: Determine the target repair method from the preset repair method set according to the defect type.
[0063] The preset repair methods include state consistency repair, event supplementation repair, narrative explanation repair, and editing and structural adjustment repair.
[0064] Among them, state consistency repair is used to directly adjust the visual content in the target shot to make it consistent with the expected state derived from the previous shot (i.e., the state constraint rules of the target shot).
[0065] Event-based supplementary repair is used to insert transitional shots between preceding shots and target shots to explain changes in character state, environmental state, or time state in the target shot, or to add action segments to the target shot to make the originally discontinuous state changes reasonable.
[0066] Narrative interpretation restoration is used to provide a narrative explanation for changes in character state, environmental state, or time state in a target shot by supplementing any one of the following narrative interpretations: subtitles, dialogue text, time markers, transition markers, and script descriptions, so as to make the originally discontinuous state changes reasonable.
[0067] Editing and structural adjustment repairs are used to eliminate or mitigate continuity defects by adjusting the sequence of the target shot and its adjacent shots, or by changing the transitions between the target shot and its adjacent shots.
[0068] S206: Repair the target lens using a target repair method.
[0069] In one specific implementation, when the target repair method is state consistency repair, the character state, environment state, or temporal state in the target shot is adjusted through any one of the following methods: local redrawing, local completion, region replacement, and regeneration. Specifically, firstly, regions in the target shot that do not conform to the state constraint rules are identified as mask regions to be repaired. Secondly, the correct state image of the corresponding entity in the preceding shot is used as a reference condition, and the state constraint rules are converted into text prompts (e.g., "Character A is in the room"). Finally, the target shot, mask regions, reference conditions, and text prompts are input into a pre-trained image or video generation model for joint inference to generate the adjusted target shot, thereby achieving any one of the following: local redrawing, local completion, region replacement, and regeneration.
[0070] In one specific implementation, when the target repair method is event-supplementary repair, a transition shot is inserted between the preceding shot and the target shot, or an action segment is added to the target shot; the transition shot or action segment is used to explain the changes in the character's state, the environment's state, or the time state in the target shot.
[0071] In one specific implementation, when the target repair method is narrative interpretation repair, any one of the following narrative interpretations is added to the target shot: subtitles, dialogue text, time markers, transition markers, and script descriptions; the narrative interpretation is used to explain changes in character state, environmental state, or time state in the target shot.
[0072] In one specific implementation, when the target repair method is editing and structural adjustment repair, the order of the target shot and its adjacent shots is adjusted, or the transition method between the target shot and its adjacent shots is changed.
[0073] In this embodiment, after determining the target repair method, the continuity defect score of the target shot can be determined first. In one specific implementation, the continuity defect score can be related to the degree of matching between the narrative state of the target shot and the state constraint rules for the target shot. In another specific implementation, the continuity defect score can also be calculated as shown in the following formula (1): Score=Wd×D_state+Wc×C_conflict+Wt×T_proximity+Wi×I_importance-We×E_explanation-Wo×O_occlusion (1) Wherein, D_state represents the degree of difference between the narrative state of the target shot and the expected state of the target shot; C_conflict represents the degree of conflict between the narrative state of the target shot and the state constraint rules for the target shot; T_proximity represents the proximity of the preceding shot and the target shot in terms of narrative time; I_importance represents the importance of the character, environment, or time in the plot; E_explanation represents the strength of the explanatory event; and O_occlusion represents the deduction caused by occlusion, uncertain detection, or low confidence. The weight coefficients Wd, Wc, Wt, Wi, We, and Wo are determined as follows: a training set including multiple sets of video samples with labeled continuous defect scores is pre-constructed; the sample videos in the training set are used as input, and the manually labeled defect scores are used as supervision labels; iterative training is performed using linear regression or gradient-based parameter optimization algorithms until the loss function converges, thereby learning the above weight coefficients; or, those skilled in the art can pre-adjust them according to the specific requirements of the actual application scenario.
[0074] Secondly, the defect severity level of the target lens is determined based on the continuity defect score. For example, if the continuity defect score is less than or equal to a first score threshold, the defect severity level of the target lens is determined to be no defect; if the continuity defect score is greater than the first score threshold and less than or equal to a second score threshold, the defect severity level of the target lens is determined to be a suspected continuity defect; if the continuity defect score is greater than the second score threshold and less than or equal to a third score threshold, the defect severity level of the target lens is determined to be a definite continuity defect; if the continuity defect score is greater than the third score threshold, the defect severity level of the target lens is determined to be a severe continuity defect. Furthermore, the defect severity level of a severe continuity defect is higher than that of a definite continuity defect, higher than that of a suspected continuity defect, and higher than that of a defect-free lens.
[0075] Finally, when the number of shots to be repaired exceeds a preset threshold, the shots are repaired sequentially according to their defect severity level (e.g., shots with higher defect severity levels are repaired first). Furthermore, after repairing a target shot, multimodal information analysis can be performed again to determine its narrative state, and state constraint verification is re-executed to verify whether the repaired target shot meets continuity requirements. If the continuity defect score of the repaired target shot drops below a first score threshold, the repair is considered successful; if it remains above the first score threshold, further iterative repair or a switch in repair methods can be implemented. Thus, this re-inspection mechanism avoids introducing new continuity issues after a single repair, improving the overall narrative consistency of the final output video.
[0076] In summary, this application provides a method for repairing multi-shot videos. By determining the narrative state of each shot in a multi-shot video and generating state constraint rules for a target shot based on the narrative state of the preceding shot and the action or plot events occurring in the preceding shot, this application can detect cross-shot states in multi-shot videos. Furthermore, when the narrative state of a target shot does not satisfy the state constraint rules for that target shot, and there is no explanatory event to release the state constraint rules, this application determines the defect type of the target shot and repairs the target shot according to the target repair method corresponding to the defect type, thereby enabling the repair of multi-shot videos when cross-shot states are discontinuous.
[0077] See Figure 3 This figure is a schematic diagram of a multi-camera video restoration device provided in an embodiment of this application. The multi-camera video restoration device 300 includes: The lens acquisition module 301 is used to acquire multi-lens video; The state determination module 302 is used to perform multimodal information analysis on each shot in the multi-shot video to determine the narrative state of each shot; the narrative state is used to characterize the character state, environmental state and temporal state in the shot; each shot includes a preceding shot and a target shot; the preceding shot is scheduled before the target shot in time sequence; The rule determination module 303 is used to determine the state constraint rules for the target shot based on the narrative state of the preceding shot and the action events or plot events that occur in the preceding shot. The type determination module 304 is used to determine the defect type of the target shot when the narrative state of the target shot does not meet the state constraint rules and there is no explanatory event to release the state constraint rules. The method determination module 305 is used to determine the target repair method from a set of preset repair methods based on the defect type; Lens repair module 306 is used to repair the target lens through target repair method.
[0078] In one specific implementation, the state determination module 302 is specifically used to: construct a key entity table based on the video production auxiliary information of the multi-shot video; the video production auxiliary information includes at least one of storyboard, prompts, character settings, prop settings, and scene settings; for each shot in the multi-shot video, determine the target object based on the visual feature similarity between each object in the shot and each key entity in the key entity table; and perform multimodal information analysis on each target object to determine the narrative state of each target object.
[0079] In one specific implementation, the explanatory events include at least one of the following: time jump events, memory events, dream events, fantasy events, montage events, explicit action events, plot explanation events, occlusion explanation events, camera direction change events, scene switching events, and automatic state change events.
[0080] In one specific implementation, the set of preset repair methods includes state consistency repair, event supplementation repair, narrative explanation repair, and editing and structural adjustment repair. The shot repair module 306 is specifically used for: when the target repair method is state consistency repair, adjusting the character state, environment state, or time state in the target shot through any one of the following methods: local redraw, local completion, region replacement, and regeneration; when the target repair method is event supplement repair, inserting a transition shot between the preceding shot and the target shot, or supplementing the target shot with action clips; the transition shot or action clips are used to explain the changes in character state, environment state, or time state in the target shot; when the target repair method is narrative explanation repair, supplementing the target shot with any one of the following narrative explanations: subtitles, dialogue text, time markers, transition markers, and script descriptions; the narrative explanations are used to explain the changes in character state, environment state, or time state in the target shot; when the target repair method is editing and structural adjustment repair, adjusting the sequence between the target shot and its adjacent shots, or changing the transition method between the target shot and its adjacent shots.
[0081] In one specific implementation, the lens repair module 306 is specifically used to: determine the continuity defect score of the target lens; the continuity defect score is related to the degree of matching between the narrative state and the state constraint rules of the target lens; determine the defect level of the target lens based on the continuity defect score; when the number of lenses to be repaired is greater than a preset number threshold, repair the lenses to be repaired in sequence according to the defect level.
[0082] In summary, this application provides a device for repairing multi-shot videos. By determining the narrative state of each shot in a multi-shot video and generating state constraint rules for a target shot based on the narrative state of the preceding shot and the action or plot events occurring in the preceding shot, this application can detect the cross-shot state of multi-shot videos. Furthermore, when the narrative state of a target shot does not satisfy the state constraint rules for that target shot, and there is no explanatory event to release the state constraint rules, this application determines the defect type of the target shot and repairs the target shot according to the target repair method corresponding to the defect type, thereby enabling the repair of multi-shot videos when the cross-shot state is discontinuous.
[0083] This application also provides a corresponding multi-lens video restoration device and a computer-readable medium for implementing the multi-lens video restoration method provided in this application.
[0084] The multi-lens video restoration device includes a memory and a processor. The memory is used to store instructions or code, and the processor is used to execute the instructions or code to enable the device to perform a multi-lens video restoration method according to any embodiment of this application.
[0085] See Figure 4 This figure is a schematic diagram of a computer-readable medium provided in an embodiment of this application. The computer-readable medium 400 stores a computer program 411, which, when executed by a processor, implements the above-described... Figure 1 The steps for repairing multi-camera videos.
[0086] It should be noted that, in the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0087] It should be noted that the machine-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0088] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0089] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
[0090] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0091] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method of repairing a multi-camera video, characterized by, The method includes: Acquire multi-camera video; Multimodal information analysis is performed on each shot in the multi-shot video to determine the narrative state of each shot; the narrative state is used to characterize the character state, environmental state, and temporal state in the shot; each shot includes a preceding shot and a target shot; the preceding shot is scheduled before the target shot. Based on the narrative state of the preceding shot and the action or plot events that occur in the preceding shot, determine the state constraint rules for the target shot. When the narrative state of the target shot does not satisfy the state constraint rule, and there is no explanatory event to release the state constraint rule, the defect type of the target shot is determined. Based on the defect type, a target repair method is determined from a set of preset repair methods; The target lens is repaired using the aforementioned target repair method.
2. The method of claim 1, wherein, The step of performing multimodal information analysis on each shot in the multi-shot video to determine the narrative state of each shot includes: Based on the video production auxiliary information of the multi-camera video, a key entity table is constructed; the video production auxiliary information includes at least one of storyboard script, prompts, character settings, prop settings, and scene settings. For each shot in the multi-camera video, the target object is determined based on the visual feature similarity between each object in the shot and each key entity in the key entity table; Multimodal information analysis is performed on each target object to determine the narrative state of each target object.
3. The method of claim 1, wherein, The explanatory events include at least one of the following: time jump events, memory events, dream events, fantasy events, montage events, explicit action events, plot explanation events, occlusion explanation events, camera direction change events, scene switching events, and automatic state change events.
4. The method of claim 1, wherein, The set of preset repair methods includes state consistency repair, event supplementation repair, narrative explanation repair, and editing and structural adjustment repair; The method of repairing the target lens using the target repair method includes: When the target repair method is state consistency repair, the character state, environment state or time state in the target shot is adjusted by any one of the following methods: local redraw, local completion, region replacement and regeneration. When the target repair method is event-supplementary repair, a transition shot is inserted between the preceding shot and the target shot, or an action segment is added to the target shot; the transition shot or the action segment is used to explain the changes in character state, environmental state, or time state in the target shot; When the target restoration method is narrative interpretation restoration, any one of the following narrative interpretations is added to the target shot: subtitles, dialogue text, time markers, transition markers, and script descriptions; The narrative interpretation is used to explain changes in the character's state, the environment's state, or the time state in the target shot; When the target repair method is editing and structural adjustment repair, the order of the target shot and its adjacent shots is adjusted, or the transition method between the target shot and its adjacent shots is changed.
5. The method of claim 1, wherein, The method of repairing the target lens using the target repair method includes: Determine the continuity defect score of the target shot; the continuity defect score is related to the degree of matching between the narrative state of the target shot and the state constraint rule; The defect severity level of the target lens is determined based on the continuity defect score. When the number of lenses to be repaired exceeds a preset threshold, the lenses to be repaired are repaired sequentially according to the degree of defect.
6. A multi-camera video repair apparatus, characterized by comprising: The device includes: a lens acquisition module, a status determination module, a rule determination module, a type determination module, a method determination module, and a lens repair module; The lens acquisition module is used to acquire multi-lens video; The state determination module is used to perform multimodal information analysis on each shot in the multi-shot video to determine the narrative state of each shot; the narrative state is used to characterize the character state, environmental state, and temporal state in the shot; each shot includes a preceding shot and a target shot; the preceding shot is scheduled before the target shot. The rule determination module is used to determine the state constraint rules for the target shot based on the narrative state of the preceding shot and the action events or plot events that occur in the preceding shot. The type determination module is used to determine the defect type of the target shot when the narrative state of the target shot does not satisfy the state constraint rule and there is no explanatory event for releasing the state constraint rule. The method determination module is used to determine the target repair method from a preset repair method set according to the defect type; The lens repair module is used to repair the target lens using the target repair method.
7. The apparatus of claim 6, wherein, The state determination module is specifically used to: construct a key entity table based on the video production auxiliary information of the multi-camera video; the video production auxiliary information includes at least one of storyboard script, prompts, character settings, prop settings, and scene settings; for each shot in the multi-camera video, determine the target object based on the visual feature similarity between each object in the shot and each key entity in the key entity table; Multimodal information analysis is performed on each target object to determine the narrative state of each target object.
8. The apparatus according to claim 6, characterized in that, The explanatory events include at least one of the following: time jump events, memory events, dream events, fantasy events, montage events, explicit action events, plot explanation events, occlusion explanation events, camera direction change events, scene switching events, and automatic state change events.
9. A device for restoring multi-lens videos, characterized in that, The device includes: a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the multi-lens video restoration method as described in any one of claims 1 to 5.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the multi-lens video restoration method as described in any one of claims 1 to 5.