VLOG video acquisition method and device, equipment and storage medium

Video clips are captured through drones and cameras, combined with video image feature extraction and VLOG templates, and the VLOG video of the target user is automatically obtained, solving the problem of cumbersome and low efficiency in the existing technology, and achieving efficient and reliable video acquisition.

CN119922352AInactive Publication Date: 2025-05-02XIANGJIANG LAB
View PDF 17 Cites 0 Cited by

Patent Information

Application Number
CN202510405345.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the VLOG video acquisition process of the target user is cumbersome, resulting in low acquisition efficiency and susceptible to manual intervention.

Method used

Through drones and cameras, video clips are taken in the scenic area, video image features are extracted and matched, and VLOG templates and translation generation models are combined to automatically obtain the target user's VLOG video.

Benefits of technology

It reduces the time for VLOG video acquisition, improves the acquisition efficiency, and improves the reliability of video acquisition, avoiding the impact of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922352A_ABST
    Figure CN119922352A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of videos and information, and discloses a VLOG video acquisition method, device and equipment and a storage medium, and the method comprises the steps: obtaining a plurality of video clips obtained by photographing a current user in a scenic spot through an unmanned plane and a camera; performing feature extraction on the video image in each video clip to obtain a video face feature corresponding to each video clip; acquiring a face image of a target user, performing feature extraction on the face image of the target user to obtain face features of the target user, matching the face features of the target user with the video face features corresponding to each video clip, and selecting the successfully matched video clip as a target video clip; determining a target VLOG template corresponding to the target video clip; and processing the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain a VLOG video of the target user. According to the invention, the acquisition efficiency of the VLOG video of the target user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of video technology and information technology, and in particular to a VLOG video acquisition method, device, equipment and storage medium. Background Art

[0002] VLOG video, full name Video Blog, is a multimedia log format that combines images, sounds and text narration. VLOG video not only vividly displays the creator's personal experience and what he sees and hears, but also incorporates his unique thoughts, feelings and emotional expressions. This form of video content has quickly become popular on social media platforms due to its authenticity, intimacy and infectiousness, and has become a popular way of self-expression and sharing.

[0003] However, the existing process of acquiring the VLOG video of the target user is cumbersome, which is not conducive to improving the efficiency of acquiring the VLOG video. The reason is that the existing technology mainly adopts the manual acquisition method to acquire the VLOG video of the target user, and the manual acquisition method consumes a lot of human resources and time resources, increases the acquisition time of the VLOG video of the target user, and is easily affected by manual intervention, so it is not conducive to improving the efficiency of acquiring the VLOG video. Summary of the invention

[0004] The present invention provides a VLOG video acquisition method, device, computer equipment and storage medium to solve the technical problem that the existing VLOG video acquisition process of target users is complicated and is not conducive to improving the VLOG video acquisition efficiency.

[0005] In a first aspect, a VLOG video acquisition method is provided, comprising: Get multiple video clips of the current user captured by drones and cameras in the scenic area; Obtaining video images in each video segment; Extracting features from the video images in each video clip to obtain video face features corresponding to each video clip; Acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; Determine a target VLOG template corresponding to the target video clip based on the VLOG template set and a predefined method; The first translation amount and the second translation amount are obtained through a preset translation amount generation model, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user.

[0006] Further, the obtaining of the video image in each video clip includes: Get the video format of each video clip; According to the video format of each video clip, each video clip is parsed and processed to obtain the video image in each video clip.

[0007] Furthermore, the extracting features of the video images in each video clip to obtain the video face features corresponding to each video clip includes: Reading a facial feature extraction model in a preset file, and inputting the video image in each video clip into the facial feature extraction model; The facial feature extraction model is used to extract features from the video images in each video clip to obtain the video facial features corresponding to each video clip.

[0008] Further, the acquiring of the face image of the target user, performing feature extraction on the face image of the target user to obtain face features of the target user, matching the face features of the target user with face features of the video corresponding to each video clip, and selecting the successfully matched video clip as the target video clip includes: Obtaining a selfie image uploaded by a user terminal device, capturing a face image of a target user in the selfie image, performing feature extraction on the face image of the target user, and obtaining face features of the target user; Obtain a matching instruction, execute the matching instruction, match the target user's facial features with the video facial features corresponding to each video clip, and select the successfully matched video clip as the target video clip.

[0009] Further, the determining the target VLOG template corresponding to the target video clip based on the VLOG template set and the predefined method includes: Reading a VLOG template set, wherein the VLOG template set includes a plurality of VLOG templates; The score corresponding to each VLOG template is obtained, and the VLOG template with the highest score is selected as the target VLOG template corresponding to the target video clip.

[0010] Further, the method of generating a model by a preset translation amount, obtaining a first translation amount and a second translation amount, processing the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user includes: A first translation amount and a second translation amount are obtained by using a preset translation amount generation model, and a face rectangular frame in a target video segment is moved to a center point of a picture by using the first translation amount and the second translation amount, and a target video segment in which the face rectangular frame is moved to a center point of a picture is selected as an adjusted video segment; An import instruction is obtained, the import instruction is executed, the adjusted video clip is imported into a target VLOG template, and the target VLOG template into which the adjusted video clip is imported is selected as the VLOG video of the target user.

[0011] Further, after obtaining the first translation amount and the second translation amount by using the preset translation amount generation model, processing the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user, the VLOG video acquisition method includes: The VLOG video is saved in a private cloud, a preview video is captured in the VLOG video, a video link of the preview video is created, and the video link of the preview video is pushed to the user terminal device; When receiving a confirmation instruction returned by the user terminal device according to the video link of the preview video, a video link of the VLOG video is created, and the video link of the VLOG video is pushed to the user terminal device.

[0012] In a second aspect, a VLOG video acquisition device is provided, comprising: The first acquisition module is used to acquire multiple video clips of the current user captured by the drone and the camera in the scenic area; A second acquisition module is used to acquire the video image in each video clip; An extraction module is used to extract features from the video images in each video clip to obtain video face features corresponding to each video clip; The third acquisition module is used to acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; A determination module, used for determining a target VLOG template corresponding to a target video clip based on a VLOG template set and a predefined method; The fourth acquisition module is used to obtain the first translation amount and the second translation amount through a preset translation amount generation model, and process the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user.

[0013] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned VLOG video acquisition method when executing the computer program.

[0014] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned VLOG video acquisition method are implemented.

[0015] The present application provides a VLOG video acquisition method, device, computer equipment and storage medium, which acquire multiple video clips of the current user captured by a drone and a camera in a scenic area; acquire a video image in each video clip; perform feature extraction on the video image in each video clip to obtain the video face feature corresponding to each video clip; acquire the face image of a target user, perform feature extraction on the face image of the target user to obtain the face feature of the target user, match the face feature of the target user with the video face feature corresponding to each video clip, and select the video clip with successful matching as the target video clip; determine the target VLOG template corresponding to the target video clip based on a VLOG template set and a predefined method; obtain a first translation amount and a second translation amount through a preset translation amount generation model, and The first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user. The beneficial effects are in two aspects. On the one hand, the first translation amount and the second translation amount are obtained through a preset translation amount generation model, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user. Since manual acquisition is not required, the acquisition time of the VLOG video of the target user is reduced, which is beneficial to improving the acquisition efficiency of the VLOG video of the target user. On the other hand, since the VLOG video of the target user is automatically acquired and will not be affected by manual intervention, it is beneficial to improve the reliability of the acquired VLOG video of the target user. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0017] Figure 1 Schematic diagram of the application environment of the VLOG video acquisition method in one embodiment of the present invention; Figure 2A schematic diagram of a flow chart of a VLOG video acquisition method provided by an embodiment of the present invention; Figure 3 yes Figure 2 Schematic diagram of the process of step S23; Figure 4 yes Figure 2 Schematic diagram of the process of step S25; Figure 5 yes Figure 2 Schematic diagram of the process of step S26; Figure 6 is a schematic diagram of the structure of a VLOG video acquisition device in one embodiment of the present invention; Figure 7 It is a schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] See also Figure 1 , Figure 1 Schematic diagram of the application environment of the VLOG video acquisition method in one embodiment of the present invention. The VLOG video acquisition method provided by the embodiment of the present invention can be applied in the following embodiments: Figure 1 In a server device of an application environment, the server device communicates with the user device through a network.

[0020] The server device obtains multiple video clips of the current user captured by the drone and camera in the scenic area; Obtaining video images in each video segment; Extracting features from the video images in each video clip to obtain video face features corresponding to each video clip; Acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; Determine a target VLOG template corresponding to the target video clip based on the VLOG template set and a predefined method; The first translation amount and the second translation amount are obtained through a preset translation amount generation model, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user.

[0021] The VLOG video is saved in a private cloud, a preview video is captured in the VLOG video, a video link of the preview video is created, and the video link of the preview video is pushed to the user-end device; when a confirmation instruction is received from the user-end device based on the video link of the preview video, a video link of the VLOG video is created, and the video link of the VLOG video is pushed to the user-end device.

[0022] In the scheme implemented by the above-mentioned VLOG video acquisition method, device, equipment and medium, the beneficial effects lie in two aspects. On the one hand, through a preset translation amount generation model, a first translation amount and a second translation amount are obtained, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user. Since manual acquisition is not required, the acquisition time of the VLOG video of the target user is reduced, which is beneficial to improving the acquisition efficiency of the VLOG video of the target user. On the other hand, since the VLOG video of the target user is automatically acquired and will not be affected by manual intervention, it is beneficial to improve the reliability of the acquired VLOG video of the target user.

[0023] Among them, user-end devices include but are not limited to personal computers, Internet of Vehicles terminals, tablet computers and portable wearable devices.

[0024] The server device may be implemented as an independent server or a server cluster consisting of multiple servers.

[0025] The present invention is described in detail below through specific embodiments. Figure 2 , Figure 2 A flow chart of a VLOG video acquisition method provided by an embodiment of the present invention includes the following steps: S21, obtaining multiple video clips of the current user captured by the drone and the camera in the scenic area; Exemplarily, multiple video clips of the current user captured by a drone and a camera in a scenic area are obtained, including: Get the coordinates of the check-in point of the scenic spot, select the coordinates of the check-in point of the scenic spot as the shooting coordinates, get the opening time of the check-in point of the scenic spot, and select the opening time of the check-in point of the scenic spot as the shooting time; The shooting coordinates and shooting time are sent to the drone and the camera, and multiple video clips of the current user shot by the drone and the camera in the scenic area are obtained according to the shooting coordinates and the shooting time.

[0026] Among them, the coordinates of the check-in point of the scenic spot are selected as the shooting coordinates, and the shooting coordinates are sent to the drone and camera. After obtaining the shooting coordinates, the drone can automatically plan an optimal flight path according to the terrain and obstacles. This can not only reduce flight time and energy consumption, but also avoid getting lost or deviating from the planned path during the flight, ensuring that the drone can safely and stably arrive at the check-in point of the scenic spot.

[0027] Among them, obtaining multiple video clips of the current user shot by drones and cameras in the scenic area according to the shooting coordinates and shooting time avoids the privacy infringement problems that may be caused by all-weather and all-round uninterrupted monitoring. Since the shooting is only carried out at the shooting coordinates and shooting time, it balances the security needs and personal privacy rights. In addition, by setting the shooting time, the data storage of drones and cameras can be effectively reduced, the storage cost of drones and cameras can be reduced, and the efficiency of video retrieval can be improved, because this can quickly locate the required video clips according to the shooting time.

[0028] Among them, the check-in points of scenic spots refer to the iconic places that users tend to visit and leave mementos when visiting scenic spots. These places often have unique attractions, which may be magnificent natural landscapes, historic buildings, or creative art installations. The existence of check-in points in scenic spots not only satisfies users' pursuit of fresh experiences, but also provides them with materials for sharing on social media. By taking photos at these places, users can record the beautiful moments of their travels and show their travel experiences on social platforms.

[0029] S22, obtaining a video image in each video segment; Wherein, obtaining the video image in each video clip includes: Get the video format of each video clip; According to the video format of each video clip, each video clip is parsed and processed to obtain the video image in each video clip.

[0030] Exemplarily, obtaining the video format of each video clip includes: The header information of each video segment is obtained, and the video format of each video segment is obtained from the header information of each video segment.

[0031] Among them, the video format is the basis for parsing video content. The video format determines the encoding method and organizational structure of the video data. By accurately identifying the video format, the parser can select the correct decoding algorithm and process to ensure the accurate restoration of the video clips.

[0032] Among them, by parsing each video clip, the video image in each video clip is obtained, which can significantly reduce the demand for data storage and transmission, because compared with video clips, video images are easier to store efficiently and transmit quickly under limited resource conditions.

[0033] S23, extracting features from the video images in each video clip to obtain video face features corresponding to each video clip; Among them, video face features refer to the facial information extracted from video images that can distinguish individual identities. Video face features mainly include facial contour features, facial texture features, and dynamic expression features.

[0034] S24, obtaining a face image of a target user, performing feature extraction on the face image of the target user to obtain face features of the target user, matching the face features of the target user with face features of the video corresponding to each video clip, and selecting a successfully matched video clip as a target video clip; The step of obtaining a face image of a target user, extracting features of the face image of the target user to obtain face features of the target user, matching the face features of the target user with face features of the video corresponding to each video clip, and selecting a successfully matched video clip as a target video clip includes: Obtaining a selfie image uploaded by a user terminal device, capturing a face image of a target user in the selfie image, performing feature extraction on the face image of the target user, and obtaining face features of the target user; Obtain a matching instruction, execute the matching instruction, match the target user's facial features with the video facial features corresponding to each video clip, and select the successfully matched video clip as the target video clip.

[0035] Among them, selfie images refer to personal images taken and recorded by user-end devices.

[0036] For the sake of illustration, an example is given below: For example, there are 3 current users, namely user A, user B, user C, and user D.

[0037] The drone and camera shoot video clip 1 of user A in the scenic area; The drone and camera shoot video clip 2 of user B in the scenic area; The drone and camera shoot video clip 3 of user C in the scenic area; The drone and the camera shoot the video clip 4 obtained by user D in the scenic area; When the user-end device is the device of user A, a selfie image uploaded by the device of user A is obtained, a facial image of user A is captured in the selfie image, features are extracted from the facial image of user A to obtain facial features of user A, a matching instruction is obtained, the matching instruction is executed, and the facial features of user A are matched with the video facial features corresponding to video clips 1, 2, 3, and 4, respectively, and video clip 1 that successfully matches is selected as the target video clip.

[0038] When the user-end device is the device of user B, a selfie image uploaded by the device of user B is obtained, a facial image of user B is captured in the selfie image, features are extracted from the facial image of user B to obtain facial features of user B, a matching instruction is obtained, the matching instruction is executed, and the facial features of user B are matched with the video facial features corresponding to video clips 1, 2, 3, and 4, respectively, and video clip 2 that successfully matches is selected as the target video clip.

[0039] When the user-end device is the device of user C, a selfie image uploaded by the device of user C is obtained, a facial image of user C is captured in the selfie image, features of the facial image of user C are extracted, facial features of user C are obtained, a matching instruction is obtained, the matching instruction is executed, the facial features of user C are matched with the video facial features corresponding to video clips 1, 2, 3, and 4, respectively, and video clip 3 that is successfully matched is selected as the target video clip.

[0040] When the user-end device is the device of user D, a selfie image uploaded by the device of user D is obtained, a facial image of user D is captured in the selfie image, features are extracted from the facial image of user D to obtain facial features of user D, a matching instruction is obtained, the matching instruction is executed, the facial features of user D are matched with video facial features corresponding to video clips 1, 2, 3, and 4, respectively, and video clip 4 that has been successfully matched is selected as the target video clip.

[0041] S25, determining a target VLOG template corresponding to the target video clip based on the VLOG template set and a predefined method; The target VLOG template is the preferred VLOG template, which generally includes the structure layout, style design, transition effect, subtitle style and music of the video.

[0042] S26, obtaining a first translation and a second translation through a preset translation generation model, processing the first translation, the second translation, the target video clip and the target VLOG template to obtain a VLOG video of the target user.

[0043] Exemplarily, the translation amount generation model is: ; .

[0044] P1 is the first translation, which is used to describe the translation of the face rectangle in the horizontal direction. P2 is the second translation, which is used to describe the translation of the face rectangle in the vertical direction. Z is the magnification factor, w is the width of the target video clip, and h is the width of the target video clip; x is the horizontal coordinate of the upper left corner of the face rectangle; y is the vertical coordinate of the upper left corner of the face rectangle; m is the width of the face rectangle, and n is the height of the face rectangle. The coordinates of the center point of the picture are (w / 2, h / 2).

[0045] Among them, the coordinates of the center point of the picture are (w / 2, h / 2).

[0046] Among them, the horizontal coordinate of the upper left corner of the face rectangular frame, the vertical coordinate of the upper left corner of the face rectangular frame, the width of the face rectangular frame, the height of the face rectangular frame, and the coordinates of the center point of the picture are obtained, and the first translation and the second translation are generated according to the horizontal coordinate of the upper left corner of the face rectangular frame, the vertical coordinate of the upper left corner of the face rectangular frame, the width of the face rectangular frame, the height of the face rectangular frame, the coordinates of the center point of the picture, and the translation generation model. Since the first translation is used to describe the translation of the face rectangular frame in the horizontal direction, and the second translation is used to describe the translation of the face rectangular frame in the vertical direction, the face rectangular frame in the target video clip can be moved to the center point of the picture through the first translation and the second translation. This automated processing method greatly improves the efficiency and convenience of video processing. Users can obtain professional-level visual effects without tedious adjustments, saving time and energy.

[0047] Wherein, after obtaining the first translation and the second translation by using the preset translation amount generation model, processing the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user, the VLOG video acquisition method includes: The VLOG video is saved in a private cloud, a preview video is captured in the VLOG video, a video link of the preview video is created, and the video link of the preview video is pushed to the user terminal device; When receiving a confirmation instruction returned by the user terminal device according to the video link of the preview video, a video link of the VLOG video is created, and the video link of the VLOG video is pushed to the user terminal device.

[0048] Capture preview videos in VLOG videos, including: Get the preset duration, and based on the preset duration, capture a preview video of the preset duration in the VLOG video.

[0049] For the sake of illustration, an example is given below: For example, the preset duration is 5 seconds, and a 5-second preview video is captured in the VLOG video.

[0050] For example, the preset duration is 10 seconds, and a 10-second preview video is captured in the VLOG video.

[0051] For example, the preset duration is 20 seconds, and a 20-second preview video is captured in the VLOG video.

[0052] For example, the preset duration is 30 seconds, and a 30-second preview video is captured in the VLOG video.

[0053] Among them, the VLOG video is stored in the private cloud. Since the private cloud provides a highly secure storage environment, strict access control and data encryption measures are implemented to ensure that the VLOG video is not illegally accessed or leaked, which directly protects the privacy rights and interests of users and the copyright security of the VLOG video. In addition, the private cloud has strong storage scalability and flexibility, and can dynamically adjust the storage space as the VLOG video grows, avoiding the capacity limit of local storage, which not only reduces the storage cost of VLOG videos, but also improves the management efficiency of VLOG videos.

[0054] Among them, when a confirmation instruction is received from the user-end device based on the video link of the preview video, it means that the preview video has been confirmed by the target user, a video link of the VLOG video is created, and the video link of the VLOG video is pushed to the user-end device. The target user clicks on the video link of the VLOG video through the user-end device to download the VLOG video.

[0055] In the embodiment of the present invention, the beneficial effects lie in two aspects. On the one hand, a first translation and a second translation are obtained through a preset translation generation model, and the first translation, the second translation, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user. Since manual acquisition is not required, the acquisition time of the VLOG video of the target user is reduced, which is beneficial to improving the acquisition efficiency of the VLOG video of the target user. On the other hand, since the VLOG video of the target user is automatically acquired and will not be affected by manual intervention, it is beneficial to improve the reliability of the acquired VLOG video of the target user.

[0056] See also Figure 3 , Figure 3 yes Figure 2 The flow chart of step S23 is described in detail as follows: S31, reading a facial feature extraction model in a preset file, and inputting the video image in each video clip into the facial feature extraction model; Among them, the facial feature extraction model adopts a pre-trained model. The pre-trained model has been fully trained on a large data set and can capture rich facial features. This ensures that the facial feature extraction model has high accuracy and robustness in facial feature extraction tasks.

[0057] In addition, by using pre-trained models, developers do not need to train facial feature extraction models from scratch, saving a lot of time and computing resources.

[0058] Optionally, the facial feature extraction model adopts a YOLOV10 model or an Insightface model.

[0059] Among them, the full name of YOLOV10 is You Only Look Once version 10.

[0060] Among them, the YOLOV10 model is a real-time target detection model. The YOLOV10 model has been significantly optimized in terms of detection accuracy and computing resource balance. The YOLOV10 model has excellent multi-scale detection capabilities. The YOLOV10 model uses multi-scale feature fusion technology, which can simultaneously detect targets of different sizes and improve detection coverage. This feature makes the YOLOV10 model more capable in dealing with complex scenes and diverse targets.

[0061] Among them, the InsightFace model is a face recognition algorithm framework based on deep learning. The InsightFace model is highly efficient. The InsightFace model uses advanced deep learning algorithms and neural network structures to quickly extract facial features from input images and complete face recognition within milliseconds. This high efficiency enables the InsightFace model to process large amounts of face data in real time and meet the needs of various real-time application scenarios.

[0062] In addition, the InsightFace model uses a large-scale face dataset during training and adopts distributed training technology to improve training speed and efficiency. This enables the InsightFace model to learn richer and more diverse facial features, further improving recognition accuracy and robustness.

[0063] S32, performing feature extraction on the video image in each video clip by using the facial feature extraction model to obtain video facial features corresponding to each video clip.

[0064] Secondly, the facial feature extraction model is highly robust and can ensure stable feature extraction in various environments to obtain the video facial features corresponding to each video clip.

[0065] In an embodiment of the present invention, the facial feature extraction model is used to extract features from the video images in each video clip to obtain the video facial features corresponding to each video clip. This method can automatically process a large number of video images, quickly capture video facial features, and greatly improve the extraction speed and accuracy of video facial features.

[0066] See also Figure 4 , Figure 4 yes Figure 2 The flow chart of step S25 is described in detail as follows: S41, reading a VLOG template set, where the VLOG template set includes a plurality of VLOG templates; Among them, the VLOG template collection is a collection of VLOG templates. The VLOG template collection contains a variety of VLOG templates of different styles and themes to meet different creative needs.

[0067] S42, obtaining the score corresponding to each VLOG template, and selecting the VLOG template with the highest score as the target VLOG template corresponding to the target video clip.

[0068] In the embodiment of the present invention, since the VLOG template with the highest score has passed extensive testing and feedback, the VLOG template with the highest score is selected as the target VLOG template corresponding to the target video clip, which can ensure the practicality and applicability of the target VLOG template.

[0069] See also Figure 5 , Figure 5 yes Figure 2 The flow chart of step S26 is described in detail as follows: S51, obtaining a first translation amount and a second translation amount through a preset translation amount generation model, moving a face rectangular frame in a target video segment to a center point of a picture through the first translation amount and the second translation amount, and selecting a target video segment in which the face rectangular frame is moved to a center point of a picture as an adjusted video segment; S52, obtaining an import instruction, executing the import instruction, importing the adjusted video clip into a target VLOG template, and selecting the target VLOG template into which the adjusted video clip is imported as the VLOG video of the target user.

[0070] Among them, the target VLOG template usually includes the video's structural layout, style design, transition effects, subtitle style and music.

[0071] In an embodiment of the present invention, an import instruction is obtained, the import instruction is executed, the adjusted video clip is imported into the target VLOG template, and the target VLOG template into which the adjusted video clip is imported is selected as the VLOG video of the target user. Since manual acquisition is not required, the acquisition time of the VLOG video of the target user is reduced, which is beneficial to improving the acquisition efficiency of the VLOG video of the target user.

[0072] See also Figure 6 , Figure 6 : is a schematic diagram of the structure of a VLOG video acquisition device in one embodiment of the present invention. Figure 6 As shown, the VLOG video acquisition device includes a first acquisition module 101, a second acquisition module 102, an extraction module 103, a third acquisition module 104, a determination module 105, and a fourth acquisition module 106. The detailed description of each functional module is as follows: The first acquisition module 101 is used to acquire multiple video clips of the current user captured by the drone and the camera in the scenic area; A second acquisition module 102, used to acquire a video image in each video clip; The extraction module 103 is used to extract features from the video images in each video clip to obtain video face features corresponding to each video clip; The third acquisition module 104 is used to acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; A determination module 105, configured to determine a target VLOG template corresponding to a target video clip based on a VLOG template set and a predefined method; The fourth acquisition module 106 is used to obtain the first translation amount and the second translation amount through a preset translation amount generation model, and process the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user.

[0073] In one embodiment, the second acquisition module 102 includes: An acquisition subunit, used for acquiring a video format of each video clip; The parsing subunit is used to parse each video segment according to the video format of each video segment to obtain the video image in each video segment.

[0074] In one embodiment, the extraction module 103 includes: A first extraction subunit is used to read a face feature extraction model in a preset file and input the video image in each video clip into the face feature extraction model; The second extraction subunit is used to perform feature extraction on the video image in each video clip through the facial feature extraction model to obtain the video facial features corresponding to each video clip.

[0075] In one embodiment, the third acquisition module 104 includes: The interception subunit is used to obtain the selfie image uploaded by the user terminal device, intercept the face image of the target user in the selfie image, perform feature extraction on the face image of the target user, and obtain the face features of the target user; The matching subunit is used to obtain the matching instruction, execute the matching instruction, match the target user's facial features with the video facial features corresponding to each video clip, and select the successfully matched video clip as the target video clip.

[0076] In one embodiment, the determining module 105 includes: A reading subunit, used for reading a VLOG template set, wherein the VLOG template set includes a plurality of VLOG templates; A sub-unit is selected to obtain a score corresponding to each VLOG template, and a VLOG template with the highest score is selected as a target VLOG template corresponding to the target video clip.

[0077] In one embodiment, the fourth acquisition module 106 includes: A generating subunit is used to generate a model by using a preset translation amount, obtain a first translation amount and a second translation amount, move a face rectangular frame in a target video segment to a center point of a picture by using the first translation amount and the second translation amount, and select a target video segment in which the face rectangular frame is moved to a center point of the picture as an adjusted video segment; The processing subunit is used to obtain an import instruction, execute the import instruction, import the adjusted video clip into a target VLOG template, and select the target VLOG template into which the adjusted video clip is imported as the VLOG video of the target user.

[0078] In one embodiment, the VLOG video acquisition device further includes: The first push module is used to save the VLOG video in the private cloud, intercept the preview video in the VLOG video, create a video link of the preview video, and push the video link of the preview video to the user terminal device; The second push module is used to create a video link of the VLOG video and push the video link of the VLOG video to the user terminal device when receiving a confirmation instruction returned by the user terminal device according to the video link of the preview video.

[0079] In the embodiment of the present invention, the beneficial effects lie in two aspects. On the one hand, a first translation and a second translation are obtained through a preset translation generation model, and the first translation, the second translation, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user. Since manual acquisition is not required, the acquisition time of the VLOG video of the target user is reduced, which is beneficial to improving the acquisition efficiency of the VLOG video of the target user. On the other hand, since the VLOG video of the target user is automatically acquired and will not be affected by manual intervention, it is beneficial to improve the reliability of the acquired VLOG video of the target user.

[0080] The specific definition of the VLOG video acquisition device can refer to the definition of the VLOG video acquisition method in the above text, which will not be repeated here.

[0081] Each module in the above-mentioned VLOG video acquisition device can be implemented in whole or in part by software, hardware and their combination. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each of the above modules.

[0082] See also Figure 7 , Figure 7 : is a schematic diagram of the structure of a computer device in an embodiment of the present invention. In one embodiment, a computer device is provided. The computer device is a server device, and its internal structure diagram can be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with external devices.

[0083] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the following steps may be implemented: Get multiple video clips of the current user captured by drones and cameras in the scenic area; Obtaining video images in each video segment; Extracting features from the video images in each video clip to obtain video face features corresponding to each video clip; Acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; Determine a target VLOG template corresponding to the target video clip based on the VLOG template set and a predefined method; The first translation amount and the second translation amount are obtained through a preset translation amount generation model, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user.

[0084] In some embodiments, the processor is configured to implement: Get the video format of each video clip; According to the video format of each video clip, each video clip is parsed and processed to obtain the video image in each video clip.

[0085] In some embodiments, the processor is configured to implement: Reading a facial feature extraction model in a preset file, and inputting the video image in each video clip into the facial feature extraction model; The facial feature extraction model is used to extract features from the video images in each video clip to obtain the video facial features corresponding to each video clip.

[0086] In some embodiments, the processor is configured to implement: Obtaining a selfie image uploaded by a user terminal device, capturing a face image of a target user in the selfie image, performing feature extraction on the face image of the target user, and obtaining face features of the target user; Obtain a matching instruction, execute the matching instruction, match the target user's facial features with the video facial features corresponding to each video clip, and select the successfully matched video clip as the target video clip.

[0087] In some embodiments, the processor is configured to implement: Reading a VLOG template set, wherein the VLOG template set includes a plurality of VLOG templates; The score corresponding to each VLOG template is obtained, and the VLOG template with the highest score is selected as the target VLOG template corresponding to the target video clip.

[0088] In some embodiments, the processor is configured to implement: A first translation amount and a second translation amount are obtained by using a preset translation amount generation model, and a face rectangular frame in a target video segment is moved to a center point of a picture by using the first translation amount and the second translation amount, and a target video segment in which the face rectangular frame is moved to a center point of a picture is selected as an adjusted video segment; An import instruction is obtained, the import instruction is executed, the adjusted video clip is imported into a target VLOG template, and the target VLOG template into which the adjusted video clip is imported is selected as the VLOG video of the target user.

[0089] In some embodiments, the processor is configured to implement: The VLOG video is saved in a private cloud, a preview video is captured in the VLOG video, a video link of the preview video is created, and the video link of the preview video is pushed to the user-end device; when a confirmation instruction is received from the user-end device based on the video link of the preview video, a video link of the VLOG video is created, and the video link of the VLOG video is pushed to the user-end device.

[0090] An embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0091] The computer-readable storage medium stores program codes, which can be called by a processor to execute the VLOG video acquisition method described in the above method embodiment.

[0092] The computer-readable storage medium has a storage space for program codes.

[0093] The program code includes the code of any step in the VLOG video acquisition method described in the above method embodiment.

[0094] For example, the program code is called by the processor and can execute the following steps: Get multiple video clips of the current user captured by drones and cameras in the scenic area; Obtaining video images in each video segment; Extracting features from the video images in each video clip to obtain video face features corresponding to each video clip; Acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; Determine a target VLOG template corresponding to the target video clip based on the VLOG template set and a predefined method; The first translation amount and the second translation amount are obtained through a preset translation amount generation model, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user.

[0095] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), and a network processor (NP); it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components.

[0096] The above description and the accompanying drawings sufficiently illustrate the embodiments of the present disclosure to enable those skilled in the art to practice them. In this article, each embodiment may focus on the differences from other embodiments, and the same or similar parts between the embodiments may refer to each other. For the methods and products disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0097] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. Technicians can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0098] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0099] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A VLOG video acquisition method, characterized in that: include: Get multiple video clips of the current user captured by drones and cameras in the scenic area; Obtaining video images in each video segment; Extracting features from the video images in each video clip to obtain video face features corresponding to each video clip; Acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; Determine a target VLOG template corresponding to the target video clip based on the VLOG template set and a predefined method; The first translation amount and the second translation amount are obtained through a preset translation amount generation model, and the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user.

2. The VLOG video acquisition method according to claim 1, characterized in that: The step of obtaining the video image in each video segment includes: Get the video format of each video clip; According to the video format of each video clip, each video clip is parsed and processed to obtain the video image in each video clip.

3. The VLOG video acquisition method according to claim 1, characterized in that: The step of extracting features from the video image in each video clip to obtain video face features corresponding to each video clip includes: Reading a facial feature extraction model in a preset file, and inputting the video image in each video clip into the facial feature extraction model; The facial feature extraction model is used to extract features from the video images in each video clip to obtain the video facial features corresponding to each video clip.

4. The VLOG video acquisition method according to claim 1, characterized in that: The step of obtaining a face image of a target user, extracting features of the face image of the target user to obtain face features of the target user, matching the face features of the target user with face features of the video corresponding to each video clip, and selecting a successfully matched video clip as a target video clip includes: Obtaining a selfie image uploaded by a user terminal device, capturing a face image of a target user in the selfie image, performing feature extraction on the face image of the target user, and obtaining face features of the target user; Obtain a matching instruction, execute the matching instruction, match the target user's facial features with the video facial features corresponding to each video clip, and select the successfully matched video clip as the target video clip.

5. The VLOG video acquisition method according to claim 1, characterized in that: The step of determining a target VLOG template corresponding to a target video clip based on a VLOG template set and a predefined method includes: Reading a VLOG template set, wherein the VLOG template set includes a plurality of VLOG templates; The score corresponding to each VLOG template is obtained, and the VLOG template with the highest score is selected as the target VLOG template corresponding to the target video clip.

6. The VLOG video acquisition method according to claim 1, characterized in that: The method of generating a model by using a preset translation amount, obtaining a first translation amount and a second translation amount, processing the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user includes: A first translation amount and a second translation amount are obtained by using a preset translation amount generation model, and a face rectangular frame in a target video segment is moved to a center point of a picture by using the first translation amount and the second translation amount, and a target video segment in which the face rectangular frame is moved to a center point of a picture is selected as an adjusted video segment; An import instruction is obtained, the import instruction is executed, the adjusted video clip is imported into a target VLOG template, and the target VLOG template into which the adjusted video clip is imported is selected as the VLOG video of the target user.

7. The VLOG video acquisition method according to claim 1, characterized in that: After the preset translation amount generation model is used to obtain the first translation amount and the second translation amount, the first translation amount, the second translation amount, the target video clip and the target VLOG template are processed to obtain the VLOG video of the target user, the VLOG video acquisition method includes: The VLOG video is saved in a private cloud, a preview video is captured in the VLOG video, a video link of the preview video is created, and the video link of the preview video is pushed to the user terminal device; When receiving a confirmation instruction returned by the user terminal device according to the video link of the preview video, a video link of the VLOG video is created, and the video link of the VLOG video is pushed to the user terminal device.

8. A VLOG video acquisition device, characterized in that: include: The first acquisition module is used to acquire multiple video clips of the current user captured by the drone and the camera in the scenic area; A second acquisition module is used to acquire the video image in each video clip; An extraction module is used to extract features from the video images in each video clip to obtain video face features corresponding to each video clip; The third acquisition module is used to acquire a face image of a target user, perform feature extraction on the face image of the target user to obtain face features of the target user, match the face features of the target user with the face features of the video corresponding to each video clip, and select the video clip with successful matching as the target video clip; A determination module, used for determining a target VLOG template corresponding to a target video clip based on a VLOG template set and a predefined method; The fourth acquisition module is used to obtain the first translation amount and the second translation amount through a preset translation amount generation model, and process the first translation amount, the second translation amount, the target video clip and the target VLOG template to obtain the VLOG video of the target user.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the VLOG video acquisition method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the VLOG video acquisition method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Video processing method and device and short video platform

    CN111460219A

  • Intelligent shooting method and device, server and storage medium

    CN111654619A

  • Video processing method and device, electronic equipment and computer readable storage medium

    CN112764845A

  • Face video clip collection method and system in scenic spot

    CN113837114A

  • Target processing method and device, equipment and storage medium

    CN114120121A