Multi-device image editing method and device and storage medium

By obtaining the camera metadata of the image acquisition device for intelligent screening and automated editing, the low efficiency of multi-device collaborative creation in the vehicle-mounted image acquisition system is solved, and high-quality, personalized video generation is achieved.

CN120658924APending Publication Date: 2025-09-16SZ ZHUOYU TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511030142.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing in-vehicle image acquisition systems are inefficient and lack logic in multi-device collaboration and content creation. It is difficult to uniformly dispatch materials from multiple devices, resulting in users requiring a large amount of manual screening and editing. Existing automatic editing tools have difficulty understanding lens motion characteristics and scene changes.

Method used

By obtaining device camera metadata from multiple image acquisition devices, including lens motion type and device motion trajectory, and combining it with user preference information for intelligent screening and automated editing, high-quality videos that meet user needs are generated.

Benefits of technology

It significantly improves the efficiency of multi-device collaborative processing, can intelligently dispatch image materials from different devices, generate artistic and logical videos, and enhance the user's creative experience and video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658924A_ABST
    Figure CN120658924A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-device image editing method and device and a storage medium, and relates to the technical field of vehicle-mounted intelligent cabin systems, and the method comprises the steps: obtaining an original image material set of a plurality of image collection devices; determining equipment operation mirror metadata respectively corresponding to each original image material in the original image material set; the equipment mirror operation metadata comprises at least one of the following: an equipment lens movement type, an equipment movement track and equipment mirror operation duration; screening at least one target image material matched with the user material preference information from the original image material set according to the operation mirror metadata of each device; and performing video editing according to each target image material to generate a target edited slice. Therefore, intelligent screening and automatic editing of the multi-view-angle materials are realized by introducing the equipment operation mirror metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicle-mounted intelligent cockpit systems, and in particular to a multi-device image editing method, device, and storage medium. Background Art

[0002] With the rapid development of smart cars and in-vehicle electronic systems, in-vehicle cockpits are no longer limited to traditional driving and basic entertainment functions, but are gradually expanding into multimedia and intelligent interaction. As a result, more and more vehicles are equipped with image acquisition devices, such as in-cabin cameras and external action cameras, for various scenarios such as driving monitoring, journey recording, and entertainment sharing.

[0003] However, existing vehicle-mounted and mobile image acquisition systems still have significant deficiencies in multi-device collaboration and content creation. First, current systems mostly rely on a single device for material management and editing, and lack the ability to uniformly schedule and collaboratively process materials across multiple devices. This means that when users are faced with multi-perspective, large-volume materials, they often need to spend a lot of manpower on manual screening and editing, resulting in overall low efficiency. In addition, existing automatic editing tools mainly rely on simple timing or device properties for material splicing, and have difficulty understanding and utilizing high-level content information such as lens motion characteristics and scene changes, resulting in a lack of logic and artistic expression in the resulting films.

[0004] To address the above issues, the industry has not yet proposed a better technical solution. Summary of the Invention

[0005] The embodiments of the present application provide a multi-device collaborative image control method, device, storage medium and program product for solving at least one of the above-mentioned technical problems.

[0006] In a first aspect, an embodiment of the present application provides a multi-device image editing method, comprising: obtaining a set of original image materials from multiple image acquisition devices; determining the device camera movement metadata corresponding to each original image material in the original image material set; the device camera movement metadata comprising at least one of the following: device lens movement type, device movement trajectory, and device camera movement duration; screening at least one target image material that matches user material preference information from the original image material set based on each of the device camera movement metadata; performing video editing based on each of the target image materials to generate a target editing film.

[0007] In a second aspect, an embodiment of the present application provides a storage medium, in which one or more programs including execution instructions are stored. The execution instructions can be read and executed by electronic devices (including but not limited to computers, servers, or network devices, etc.) to execute any of the above-mentioned multi-device image editing methods of the present application.

[0008] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the multi-device image editing methods described above in the present application.

[0009] In a fourth aspect, an embodiment of the present application further provides a computer program product, the computer program product comprising a computer program stored on a storage medium, the computer program comprising program instructions, and when the program instructions are executed by a computer, the computer executes any of the above-mentioned multi-device image editing methods. The beneficial effects of the embodiments of the present application are: By introducing device camera metadata, combined with intelligent material screening and automated editing technology, the efficiency of multi-device collaborative processing is significantly improved. It can intelligently dispatch video footage from different devices, accurately screening based on high-level information such as lens movement and trajectory, and generating high-quality videos that meet user creative preferences. As a result, through automated editing processes and matching personalized needs, users can more easily create artistic and logical works, significantly improving the efficiency and creative experience of the in-vehicle image creation system. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A flowchart of an example of a multi-device image editing method according to an embodiment of the present application is shown; Figure 2 An operational flowchart of another example of a multi-device image editing method according to an embodiment of the present application is shown; Figure 3 A business flow diagram of an example of a multi-device image editing method according to an embodiment of the present application is shown; Figure 4 This is a schematic structural diagram of an embodiment of an electronic device of the present application. DETAILED DESCRIPTION

[0012] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0013] It should also be noted that, in this document, the terms "include" and "comprising" include not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, the elements defined by the phrase "include..." do not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the elements.

[0014] In the technical solutions of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved shall comply with the provisions of relevant laws and regulations and shall not violate public order and good morals.

[0015] It should be noted that in the multi-device collaborative shooting solutions currently provided by related technologies, users usually need to use different software to control various types of shooting equipment. The lack of a unified control platform makes the operation cumbersome and increases the complexity of the operation. Especially when shooting efficiently, users have to switch software frequently, resulting in a waste of time and energy. In addition, there is a lack of centralized management of materials shot by different devices. The unified storage, search and retrieval of materials all rely on multiple platforms, which seriously affects work efficiency.

[0016] Secondly, current technologies often require multi-party collaboration to achieve synchronized shooting and intelligent post-production editing across multiple devices. Due to inconsistencies in footage, angles, preset trajectories, and other factors, a single operation is difficult to achieve, including coordinated shooting, footage alignment, and final editing. This not only increases labor costs but also reduces workflow flexibility and responsiveness.

[0017] Furthermore, current technologies lack the ability to uniformly capture and align metadata from different devices. The integration of metadata such as shooting time, shot content, and flight trajectory often relies on manual comparison and organization, which not only increases the likelihood of errors but also prolongs editing and post-production cycles. Consequently, in scenarios where multiple devices collaborate to shoot, the lack of an effective metadata management mechanism makes the integration and post-processing of footage more complex and inefficient.

[0018] It should be understood that the purpose of the above description of the current related art is only to facilitate the public to better understand the inventive spirit and motivation of this application, and is not to be construed as limiting this application. In addition, the technical solutions described in the above-mentioned current related art are not prior art and may also be undisclosed technical solutions, such as solutions under research or in the laboratory stage.

[0019] Figure 1 A flowchart of an example of a multi-device image editing method according to an embodiment of the present application is shown.

[0020] Regarding the execution subject of the method of the embodiment of the present application, it can be any controller or processor with computing or processing capabilities, such as a vehicle-mounted terminal or a mobile phone terminal, etc. In the following embodiments, the car cockpit computer will be used as an example for description. By introducing the collaborative processing of multi-device image materials and the analysis of device lens movement metadata, the shortcomings of existing vehicle-mounted and mobile image acquisition systems in multi-device collaboration and content creation are solved. It can more accurately screen out target materials that meet user preferences, avoid over-reliance on simple timing or device attribute splicing in traditional editing methods, ensure that the material splicing can better reflect the lens motion characteristics and scene changes, and greatly improve the user experience and video creation efficiency.

[0021] In some examples, the method of the embodiments of the present application can be integrated into an electronic device or terminal through software, hardware, or a combination of software and hardware, and the type of terminal or electronic device can be diverse, such as a mobile phone, tablet computer, desktop computer, or car terminal, etc.

[0022] like Figure 1 As shown, in step S110, a set of original image materials from multiple image acquisition devices is obtained.

[0023] It's important to note that with the widespread adoption of intelligent in-vehicle systems, more and more image capture devices are being installed in vehicles. These devices include in-cabin cameras, drones (equipped with image capture devices), handheld cameras, and mobile phones, capable of recording various perspectives around and within the vehicle. In some cases, in-cabin cameras can coordinate and manage multiple camera devices, enabling automated capture, unified footage management, and automated filming.

[0024] Here, for image acquisition devices of different types or interfaces, corresponding data transmission protocols can be used to transmit the raw image materials they capture. For the received raw image material set, the cockpit vehicle computer can classify and organize it according to device type, shooting time and shooting location.

[0025] It should be noted that different image acquisition devices have corresponding shooting perspectives, such as the sky perspective of a drone, the cabin perspective of a handheld camera, etc., which can support the unified collection of multi-perspective materials and provide a basis for the collaborative management of materials from multiple devices.

[0026] In step S120 , the device camera movement metadata corresponding to each original image material in the original image material set is determined.

[0027] Here, the device camera movement metadata may include the lens movement characteristics of the device during the shooting process. Specifically, the device camera movement metadata may include at least one of the following: the device lens movement type, the device movement trajectory, and the device camera movement duration.

[0028] Device camera metadata can be determined in a variety of ways, such as through sensor detection and data analysis. In some examples, each image capture device is equipped with a motion sensor and location signal marker module, which can record the device's motion information in real time, including the type of lens movement (such as translation, rotation, and zoom), the motion trajectory (such as the device's displacement path in space), and the duration of the device's motion (the length of time the lens continues to move). The camera metadata of all image capture devices is synchronized with the image material in real time, ensuring that each image material can be accurately associated with the device's camera movement characteristics at the time of capture.

[0029] In step S130 , video editing is performed based on the device camera metadata and the original image material set to generate a target editing film.

[0030] Here, after completing the collection and association of image materials and equipment camera movement metadata, using this information for precise editing can enable more intelligent and detailed editing of the materials, ensuring that the edited film meets user needs and exhibits high-quality image effects.

[0031] In some embodiments, the screening of editing materials based on device camera movement metadata can be supported. For example, if the user requires the editing to include materials with smooth shots, materials with a lens movement type of "pan" or "push-pull" can be selected for editing; if the user requires a highly dynamic picture, materials with a device movement trajectory of "fast horizontal movement" or "rapid rotation" can be selected. At this time, device camera movement metadata, such as lens movement type and device movement trajectory, will serve as an important basis for screening materials. In addition, the duration of specific materials can be adjusted according to user requirements. For example, if the material of a certain shot is too long, the duration of the material of the corresponding shot can be adjusted.

[0032] It should be noted that video editing based on device camera metadata can support transitional editing of video footage from different perspectives, such as shot switching and fade-in and fade-out, making the transitions between footage more natural and smooth. The editing mode for these transitional editing combinations can be defined using an intelligent editing algorithm based on device camera metadata or by integrating user editing requirements data, and this should not be limited here. Therefore, combining camera metadata optimizes the multi-perspective video editing process, further enhancing the visual and expressive quality of the video, and thus generating high-quality video footage.

[0033] In some examples of the embodiments of the present application, the multi-view video editing process can be performed by integrating user editing requirements with device metadata, thereby supporting accurate editing of multi-view video materials.

[0034] More specifically, user editing requirement data is obtained, where the user editing requirement data includes a device camera mode required for editing. Exemplarily, the user editing requirement data is collected through user interaction, such as through a user operation interface.

[0035] The device camera mode may include at least one device camera metadata and may also include other auxiliary information to form a more comprehensive and personalized camera mode. For example, the device camera mode may include device camera metadata, such as whether one or more specific camera motion modes (such as push-pull, pan, and pan) are required; the device camera mode may also include camera switching styles, such as whether smooth camera transitions or abrupt camera switches are required; and the device camera mode may also include shot duration and shot transition effects, such as fade-in, fade-out, and quick cuts.

[0036] In some implementations, natural language processing can also be used to determine the corresponding device camera movement mode. For example, a user can input "I want a 30-second swooping camera movement followed by a 10-second cabin panning camera movement" into the cabin computer via voice or text, and the corresponding device camera movement mode can be identified using NLP technology.

[0037] Device lens motion types include any of the following: swooping shots, dolly shots, panning shots, panning shots, and follow shots. A swooping shot refers to a device (e.g., a drone) shooting vertically downwards, typically used to show a panoramic view of the scenery or people below, giving a bird's-eye view. A dolly shot (e.g., a cockpit camera) refers to a device lens that zooms in or out by changing its focal length, creating a change in focus, often used to highlight or distance the subject. A panning shot refers to a device (e.g., a handheld camera) rotating horizontally or vertically around a central point. A panning shot refers to a device (e.g., a drone or handheld camera) moving the lens horizontally or vertically smoothly, often used to show the breadth of a scene or to follow the movement of a target object. A follow shot refers to a device (e.g., a drone) shooting closely behind a specific target object or person, often used to capture the movement of an object or person in a dynamic scene.

[0038] Then, the original video material set is edited according to the user's editing requirement data and the device's camera movement metadata to generate the target editing film.

[0039] Here, by integrating the user's editing demand data with the device's camera metadata, the most suitable lenses can be selected for splicing and transition, ensuring that the final edited film is of high quality and meets personalized needs.

[0040] For example, the device's camera metadata can be matched to the user's editing requirements. For example, if the user specifies that they want to see a "pull shot," all clips that meet this requirement will be filtered out. Furthermore, the duration of the selected clips can be adjusted based on the user's needs. If a clip is too long, it can be edited and shortened.

[0041] Furthermore, based on user requirements, transition effects can be intelligently added between consecutive shots. For example, if a user requests a "smooth transition," a "fade in / out" effect will be inserted between the edit points. Finally, the edited footage is synthesized on the timeline in the order specified by the user, ultimately generating the desired video.

[0042] Therefore, by integrating camera metadata with the user's editing needs, it is possible to accurately select and edit video materials that meet the needs, achieve precise editing of multi-perspective video materials, and generate video films that meet personalized requirements.

[0043] In some examples of the present application, various image acquisition devices can be remotely controlled through the cockpit computer. For example, in an in-vehicle smart cockpit scenario, the cockpit domain serves as the center, connecting the vehicle-mounted drone, handheld imaging devices (such as action cameras, pocket cameras, and micro single-lens cameras), and in-cabin imaging devices (such as DMS sensors) via wired (USB, etc.) or wireless (Wi-Fi, etc.) connections.

[0044] Specifically, when an imaging device manipulation request is detected, target image acquisition device information and corresponding manipulation content in the imaging device manipulation request are parsed, where the manipulation content includes shooting control information and / or motion control information.

[0045] In some embodiments, control requests can be triggered via a touchscreen, voice recognition module, or other input device (e.g., onboard buttons or physical switches) in the cockpit. When a user requests to adjust the imaging device (e.g., adjust the drone's angle or activate an action camera), the cockpit system detects this request.

[0046] The in-cabin computer's intelligent control system parses control requests, which include target imaging device information (such as device type and ID) and control content (such as camera control and device motion control), to identify the specific operation required. Camera control information may include starting or stopping camera, adjusting focus, exposure settings, and shooting mode. Motion control information may include the device's motion mode and trajectory.

[0047] Then, the device connection protocol that matches the target image acquisition device information is called to encapsulate the control content to generate corresponding device control instructions. The device control instructions are then sent to the corresponding target image acquisition device to control the target image acquisition device.

[0048] Each image acquisition device (such as drones, sports cameras, and in-car imaging devices) has a unique identifier (such as device ID, type, and model). Therefore, it is necessary to find the connection information and protocol requirements related to the device.

[0049] In some implementations, depending on the device connection method (wired USB connection or wireless Wi-Fi connection), different communication protocols may be used for data transmission and instruction execution.

[0050] For example, when an image acquisition device is connected to the vehicle via USB, the system will select the appropriate device protocol to encapsulate the control content based on the communication requirements of the USB protocol and the image transmission module. At this time, data can be transmitted through the standard USB communication protocol and the relevant API interface can be called to control the imaging device.

[0051] When the image acquisition device is connected to the vehicle computer via Wi-Fi, the system needs to exchange data with the image transmission module using the Wi-Fi protocol. At this point, the vehicle computer, operating in STA (Station) mode, connects to the device's AP (Access Point) mode, encapsulating and transmitting control commands to the device via the wireless network. Based on the target device's information and control content, the commands are encapsulated into standardized data packets to meet the target device's protocol requirements. Therefore, by precisely selecting the connection protocol and encapsulating the control commands, whether using a wired USB connection or a wireless Wi-Fi connection, commands can be intelligently encapsulated and transmitted based on different communication methods, enabling efficient and reliable remote device control.

[0052] Through the above-mentioned communication connection with each image acquisition device, users can also simultaneously view or switch the images captured in real time by different image acquisition devices. For example, the cockpit display screen can display the preview of the vehicle-mounted drone shooting, the preview of the handheld image shooting and the preview of the image device in the cockpit. Users can view the real-time images captured by different devices through the cockpit display screen. These images can be displayed in the same interface or in different interfaces (users can switch the interfaces by themselves).

[0053] Figure 2 An operational flowchart of another example of a multi-device image editing method according to an embodiment of the present application is shown.

[0054] like Figure 2 As shown, in step S210, a multi-device collaborative shooting script is obtained, and parsed to generate device camera movement instructions corresponding to each image acquisition device.

[0055] Here, the device camera movement instruction is used to indicate at least one of the following camera movement control information: device lens movement type, device movement trajectory, and device camera movement duration.

[0056] In some implementations, the multi-device collaborative shooting script predefines device camera metadata for different image capture devices, enabling each device to capture raw image material based on the device camera metadata. This eliminates the need for image capture devices to be equipped with motion sensors or location markers; the pre-specified device camera metadata can effectively correlate the raw image material captured by each device.

[0057] In step S220, each device mirror movement instruction is sent to the corresponding image acquisition device respectively, so as to drive each image acquisition device to acquire original image materials according to the corresponding device mirror movement instruction.

[0058] Here, the appropriate transmission method is selected based on the connection type of each device (e.g., wired USB connection, wireless Wi-Fi connection, etc.). Furthermore, the details of command packaging and transmission for generating corresponding device camera movement commands based on the script content can be partially referred to in the description of other embodiments above and will not be repeated here.

[0059] In the embodiments of this application, a multi-device collaborative shooting script generates camera movement instructions for each imaging device, thereby controlling multiple image acquisition devices to shoot according to a predetermined plan, ensuring efficient and accurate capture of multi-angle and multi-view footage. Thus, by combining predefined device camera movement metadata to generate device camera movement instructions, multiple image acquisition devices are controlled to coordinate shooting operations, improving the degree of automation in the shooting process and the quality of the footage.

[0060] In step S230 , the device camera movement metadata corresponding to each original image material in the original image material set is determined.

[0061] Specifically, during multi-device collaborative filming, multiple image capture devices will shoot according to preset camera movement instructions and device camera movement metadata. The image material captured by each device not only includes the image content, but also carries the corresponding device camera movement metadata. This ensures the accurate association between the original image material and the corresponding device camera movement metadata. Through effective metadata extraction and material matching, detailed motion information for each shot can be obtained.

[0062] In step S240 , at least one target image material that matches the user's material preference information is screened from the original image material set.

[0063] Here, content selection is performed based on user content preference information. This user content preference information can be multi-dimensional, such as various device camera metadata (e.g., user-preferred camera motion types) or scene content preferences (e.g., landscape types that match the user's preferences). For example, user content preference information can be updated based on the user's historical editing behavior and / or user profile information. For example, by analyzing the user's past editing history, the user can identify which types of shots or motions best suit the user's style, thereby inferring the user's preferences for new content.

[0064] Regarding the update of user material preference information, in some examples of the embodiments of the present application, historical editing operation records and / or user preference setting information can be obtained. The historical editing operation records contain multiple historical editing videos and corresponding user operations on the videos, which provide important information about the user's preferred content during the editing process. For example, users often choose certain types of scenes, lens movement methods or specific transition effects, which can all serve as a basis for judging their future preferences. In addition, the user preference setting information contains at least one user preferred scene and / or user preferred camera movement, which directly reflects the user's personalized needs for scene content, lens type, etc. during the material screening and editing process, such as a preference for "sunny natural scenery" or "smooth lens push-pull effect."

[0065] Furthermore, user material preference information is determined based on historical editing operation records and / or user preference settings. For example, machine learning techniques can be incorporated to learn from user behavior and infer potential user needs, thereby continuously optimizing user material preference information. Thus, by combining historical editing operation records and preference settings, user material preference information can be dynamically adjusted, allowing material selection to more dynamically tailor to the user's individual needs.

[0066] In some examples of the embodiments of the present application, the analysis of scene semantic information can also be introduced to automatically screen target image materials. Specifically, the scene semantic information corresponding to each original image material in the original image material set is extracted. The scene semantic information may include scene content classification (such as sunset scenery), object and character detection, etc., to help the system understand the content of the image material more comprehensively. Furthermore, the scene semantic information of each original image material is matched with the user material preference information to screen at least one target image material from the original image material set. Thus, through the multi-dimensional matching of semantic information and user preferences, a deeper contextual understanding is provided for the material screening of multi-view videos, and materials that meet user preferences can be selected more accurately.

[0067] In step S250, each target image material is video-edited according to the device camera metadata to generate a target edited film.

[0068] Here, based on the previously screened target image material and its corresponding device camera metadata, the video editing task is performed in combination with the specific motion information of each shot, and a final edited film that meets the user's preferences can be generated.

[0069] In some examples of the embodiments of the present application, the original image material set and / or at least one target image material are classified and displayed according to the device camera movement metadata.

[0070] For example, all materials using panning lenses are classified into one category, materials using rotating lenses are classified into another category, and materials can even be further divided according to the shooting trajectory of the device. In this way, the classification management of multi-viewing image materials of different devices is achieved, which helps to quickly locate the corresponding materials and improve the efficiency of material screening and use. Specifically, independent display areas or filter labels can be provided for each type of lens motion or material with a specific trajectory, so that users can choose to view different types of materials according to their needs, and realize unified management of different imaging devices. Through the graphical interface, editors can intuitively browse and select materials that meet the current editing needs. For example, if a panning lens scene is required, the user only needs to click on the "panning lens" category to automatically display all image materials that meet this category. The images generated by multiple devices can be managed in a unified manner, such as previewing or deleting materials, which greatly simplifies the user management and screening process of materials.

[0071] In some other examples of the embodiments of the present application, the original image material set and / or at least one target image material may also be displayed by classifying the materials according to the shooting devices of the materials.

[0072] For example, footage captured by all vehicle-mounted cameras can be grouped into one category, footage captured by drones into another category, and footage captured by handheld cameras into another category. This allows for categorized management of multi-angle footage captured by different devices. Specifically, separate display areas or filter tabs can be provided for footage captured by each device, allowing users to view different types of footage based on their needs, achieving unified management of different imaging devices. Through a graphical interface, editors can intuitively browse and select footage that meets their current editing needs.

[0073] Furthermore, in practice, a new device can be connected to the cockpit vehicle computer, and the shooting materials stored in the new device can be uploaded to the material library corresponding to the cockpit vehicle computer, and these shooting materials can be classified, stored and managed.

[0074] Figure 3 A business flow diagram of an example of a multi-device image editing method according to an embodiment of the present application is shown.

[0075] like Figure 3As shown, multiple image acquisition devices (such as drones and handheld devices) execute tasks through automated camera movement scripts to provide rich, multi-dimensional imagery, enabling precise editing during post-editing. Specifically, drones can execute a variety of complex camera movements targeting moving targets, such as hedging shots, zooming out shots, and spiral shots with continuously changing radiuses. These multi-angle shooting methods provide a rich source of material for editing. Handheld devices can perform in-cabin camera movements, such as surround shots and specific shots of the driver, co-pilot, and rear seats. Furthermore, users can freely hold the device to shoot, increasing the diversity and flexibility of their shots.

[0076] During the entire shooting process, all the shooting materials and camera movement information (including lens movement type, movement trajectory, duration, shooting date, etc.) of all devices are recorded in real time and transmitted to the editor. The editor system not only receives the original image material, but also automatically identifies the highlight clips in the video based on the camera movement metadata and combines AI scene recognition technology to perform intelligent editing. The system automatically determines the timing of lens switching based on the requirements of different camera movement types to ensure a natural and smooth transition between different lenses. For example, the system can select specific lens types for connection according to user needs. If the user wants a swooping camera movement followed by a panoramic camera movement in the cabin, the editor will automatically match the appropriate material based on the camera movement metadata and perform a smooth lens transition.

[0077] In addition, the editor can also generate a time-related travel record film based on the time span of the shooting date. The system automatically organizes the footage into a chronological travel record based on the date information, showing the complete process from start to finish. For example, if the footage shot by the user spans a day or a few hours, it can support automatic generation of a film within a time span, showing the complete process of travel or activity. By supporting the generation of film based on date spans, not only the efficiency of material integration is improved, but also the final edited content is more coherent and narrative.

[0078] Therefore, by combining automated camera scripts, intelligent editing and AI scene recognition technology, efficient processing and precise editing of multi-device image materials can be achieved, allowing users to obtain rich, smooth and demand-compliant editing films, improving the efficiency of the editing process and video quality.

[0079] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of combined actions, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application. In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0080] In some embodiments, the embodiments of the present application also provide a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes any one of the above-mentioned multi-device image editing methods.

[0081] In some embodiments, an embodiment of the present application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a multi-device image editing method.

[0082] The apparatus of the embodiment of the present application described above can be used to execute the multi-device image editing method of the embodiment of the present application, and accordingly achieve the technical effects achieved by the multi-device image editing method of the embodiment of the present application, which will not be described in detail here. In the embodiment of the present application, the relevant functional modules can be implemented by a hardware processor.

[0083] Figure 4 This is a hardware structure diagram of an electronic device for executing a multi-device image editing method provided by another embodiment of the present application. Figure 4 As shown, the device includes: One or more processors 410 and memory 420, Figure 4 A processor 410 is taken as an example.

[0084] The device for executing the multi-device video editing method may further include: an input device 430 and an output device 440 .

[0085] The processor 410, the memory 420, the input device 430 and the output device 440 may be connected via a bus or other means. Figure 4 The bus connection is taken as an example.

[0086] Memory 420, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the multi-device image editing method in the embodiments of the present application. Processor 410 executes the non-volatile software programs, instructions, and modules stored in memory 420 to execute various server functional applications and data processing, thereby implementing the multi-device image editing method in the above-mentioned method embodiment.

[0087] The memory 420 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the device, etc. In addition, the memory 420 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 420 may optionally include a memory remotely located relative to the processor 410, and these remote memories may be connected to the device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0088] The input device 430 may receive input digital or character information and generate signals related to user settings and function control of the device. The output device 440 may include a display device such as a display screen.

[0089] The one or more modules are stored in the memory 420 and, when executed by the one or more processors 410 , perform the multi-device image editing method in any of the above method embodiments.

[0090] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0091] The electronic devices of the embodiments of the present application exist in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and are primarily designed to provide voice and data communications. These terminals include smartphones (e.g., iPhones), multimedia phones, feature phones, and low-end phones.

[0092] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access. These terminals include PDAs, MIDs, and UMPCs, such as the iPad.

[0093] (3) Portable entertainment devices: These devices can display and play multimedia content. These devices include audio and video players (such as iPods), handheld game consoles, e-books, smart toys, and portable car navigation devices.

[0094] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but because it needs to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.

[0095] (5) Other electronic devices with data interaction functions.

[0096] In some embodiments, this application further provides a mobile platform equipped with the computer device described in any embodiment of this application. Mobile platforms include, but are not limited to, vehicles, tracked robots, bipedal robots, quadrupedal robots, etc., where the vehicles may be passenger cars, pickup trucks, and vans. It should be noted that the above are merely examples, and this application does not limit the specific form of the mobile platform.

[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0098] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a general hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-device image editing method, comprising: Obtaining a collection of original image materials from multiple image acquisition devices; Determining device camera movement metadata corresponding to each original image material in the original image material set; the device camera movement metadata includes at least one of the following: device lens movement type, device movement trajectory, and device camera movement duration; Video editing is performed based on the device camera metadata and the original image material set to generate a target editing film.

2. The method according to claim 1, further comprising: When an imaging device manipulation request is detected, parsing target image acquisition device information and corresponding manipulation content in the imaging device manipulation request; The control content includes shooting control information and / or motion control information; The device connection protocol matching the target image acquisition device information is called to encapsulate the control content to generate corresponding device control instructions, where the device control instructions are used to control the target image acquisition device.

3. The method according to claim 1, wherein The obtaining of the original image material set of multiple image acquisition devices includes: Obtaining a multi-device collaborative shooting script and parsing it to generate a device camera movement instruction corresponding to each image acquisition device; the device camera movement instruction is used to indicate at least one of the following camera movement control information: device lens movement type, device movement trajectory, and device camera movement duration; Each of the device mirror movement instructions is sent to the corresponding image acquisition device respectively to drive each image acquisition device to acquire original image materials according to the corresponding device mirror movement instruction.

4. The method according to claim 1, wherein The video editing is performed based on the device camera metadata and the original image material set to generate a target editing film, including: Acquiring user editing requirement data, wherein the user editing requirement data includes a device camera mode required for editing; The original image material set is video edited according to the user editing requirement data and the device camera movement metadata to generate a target editing film.

5. The method according to claim 1, wherein The video editing is performed based on the device camera metadata and the original image material set to generate a target editing film, including: Screening at least one target image material that matches the user's material preference information from the original image material set; Video editing is performed on each of the target image materials according to the device camera metadata to generate a target editing film.

6. The method according to claim 5, wherein: The step of screening at least one target image material that matches the user's material preference information from the original image material set includes: Extracting scene semantic information corresponding to each original image material in the original image material set; The scene semantic information of each of the original image materials is matched with the user material preference information to filter at least one target image material from the original image material set.

7. The method according to claim 6, wherein: Before matching the scene semantic information of each of the original image materials with the user material preference information to filter at least one target image material from the original image material set, the method further includes: Obtaining historical editing operation records and / or user preference setting information; the historical editing operation records include multiple historical editing results and corresponding user operations for the results, and the user preference setting information includes at least one user preferred scene and / or user preferred camera movement; User material preference information is determined according to the historical editing operation record and / or the user preference setting information.

8. The method according to claim 5, further comprising: The original image material set and / or the at least one target image material are classified and displayed according to the device lens movement metadata and / or the shooting device.

9. The method according to any one of claims 1 to 8, wherein The device lens motion type includes any one or more of the following: a swooping shot, a zooming shot, a panning shot, a panning shot, and a following shot; The image acquisition device includes any one or more of the following: a vehicle-mounted camera, a drone, and a handheld camera.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image editing apparatus, image editing method and program

    CN101989173A

  • Video editing method, device and equipment and storage medium

    CN111757149A

  • Video editing method and device, equipment and medium

    CN118018812A

  • Shooting control method, mechanical arm and storage medium

    CN118433542A

  • Broadcast television news video auxiliary editing method and system

    CN118828054A