Video description generation method, device, terminal and computer-readable storage medium

By using a video description model on the server side to extract and describe the video cover images sent by smart home devices, the problem of users needing to further identify the video cover images to determine the content is solved, and the user's viewing efficiency is improved.

CN118675083BActive Publication Date: 2025-06-24SHENZHEN QIHOO INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410695591.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-06-24
Estimated Expiration
2044-05-30

AI Technical Summary

Technical Problem

When a user views a video cover image, he or she needs to further identify the video cover image to determine the image content corresponding to the video cover image, so that the user's viewing efficiency is low.

Method used

The server receives the video cover image corresponding to the target event video sent by the smart home device, uses the video description model to extract the video cover image, obtain the target image features, and generate a video description based on the target image features.

Benefits of technology

By generating video descriptions of video cover images, users can quickly understand the video content without further identification of video cover images, which improves users' viewing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118675083B_ABST
    Figure CN118675083B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method, device, terminal, and computer-readable storage medium for video description generation. The method is applied to a server and includes: receiving a video cover image corresponding to a target event video sent by a smart home device, where the video cover image is a cover image determined by the smart home device from the target event video corresponding to a basic event, and the target event video is an event video extracted by the smart home device from the security video collected in the environment where it is located; performing image feature extraction processing on the video cover image based on a video description large model to obtain target image features; and performing video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image. Thereby, it solves the technical problem that when a user views a video cover image, the user needs to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimodal technologies, and in particular, to a method, apparatus, terminal, and computer-readable storage medium for generating video descriptions. Background Art

[0002] In recent years, various intelligent monitoring terminals have rapidly emerged. After an intelligent monitoring terminal generates a corresponding video clip from a monitoring video, it will generate a video cover image corresponding to the video clip. However, when a user views the video cover image, the user needs to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency for the user. Summary of the Invention

[0003] Embodiments of this application provide a method, apparatus, terminal, and computer-readable storage medium for generating video descriptions, which can solve the technical problem that when a user views a video cover image, the user needs to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency for the user.

[0004] In a first aspect, an embodiment of this application provides a method for generating a video description, which is applied to a server. The method includes:

[0005] Receiving a video cover image corresponding to a target event video sent by a smart home device, where the video cover image is a cover image determined by the smart home device from a target event video corresponding to a basic event, and the target event video is an event video extracted by the smart home device from a security video collected from the environment where it is located;

[0006] Performing image feature extraction processing on the video cover image based on a video description large model to obtain target image features;

[0007] Performing video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0008] Optionally, the performing image feature extraction processing on the video cover image based on the video description large model to obtain target image features includes:

[0009] Performing image character feature extraction processing on the video cover image based on the video description large model to obtain character image features, performing image event feature extraction processing on the video cover image to obtain event image features, and performing image scene feature extraction processing on the video cover image to obtain scene image features;

[0010] Obtaining target image features based on the character image features, the event image features, and the scene image features.

[0011] Optionally, the image scene feature extraction process for the video cover image to obtain scene image features includes:

[0012] Obtain the basic scene set for the smart home device and the basic scene features corresponding to the basic scene;

[0013] Based on the video description large model, perform differential feature recognition processing on the cover scene of the video cover image and the basic scene to obtain target differential features;

[0014] Determine scene image features based on the basic scene features and the target differential features.

[0015] Optionally, the video description generation process for the target image features based on the video description large model to obtain a video description for the video cover image includes:

[0016] Through the video description large model, perform character description generation processing on the character image features to obtain a character description, perform event description generation processing on the event image features to obtain an event description, and perform scene description processing on the scene image features to obtain a scene description;

[0017] Through the video description large model, perform description association processing based on the character description, the event description, and the scene description to obtain a video description of the video cover image.

[0018] Optionally, after the video description generation process for the target image features based on the video description large model to obtain a video description for the video cover image, it further includes:

[0019] Obtain an associated video cover image corresponding to the video cover image, and determine an associated video description corresponding to the associated video cover image based on the video description large model; wherein, the associated video cover image is the cover image of an associated event video, and the associated event video is determined from all event videos according to time relevance for the target event video;

[0020] Perform video description association processing on the video description based on the associated video description to obtain a video association description for the video cover image.

[0021] Optionally, the method further includes:

[0022] Create an initial video description large model based on a basic large model;

[0023] Obtain sample video cover images and label sample video descriptions for the sample video cover images;

[0024] Input the sample video cover image into the initial video description large model for model training. Through the initial video description large model, perform image feature extraction processing on the sample video cover image to obtain sample image features, and perform video description generation processing on the sample image features to obtain a reference video description for the sample video cover image;

[0025] During the model training process, adjust the model parameters of the initial video description large model based on the sample video description and the reference video description to obtain a video description large model.

[0026] Optionally, the method further includes:

[0027] Determine the target event dynamic bar corresponding to the target event video in the smart home management interface, and display the video cover image and the video description in the target event dynamic bar;

[0028] In response to a video viewing request for the target event dynamic bar, display the target event video corresponding to the video cover image.

[0029] In a second aspect, an embodiment of the present application provides a video description generation method, which is applied to a smart home device. The method includes:

[0030] Collect the security video of the surrounding environment, and extract the target event video corresponding to the basic event from the security video;

[0031] Determine the video cover image corresponding to the target event video, and send the video cover image to the server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, and perform image feature extraction processing on the video cover image based on the video description large model to obtain target image features, and perform video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0032] In a third aspect, an embodiment of the present application provides a video description generation device, which is applied to a smart home device. The device includes:

[0033] A collection module, adapted to collect the security video of the surrounding environment, and extract the target event video corresponding to the basic event from the security video;

[0034] A sending module, adapted to determine a video cover image corresponding to the target event video, and send the video cover image to a server, so that the server receives the video cover image corresponding to the target event video sent by a smart home device, and performs image feature extraction processing on the video cover image based on a video description large model to obtain target image features, and performs video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0035] In a fourth aspect, an embodiment of the present application provides a video description generation device, which is applied to a smart home device. The device includes:

[0036] An acquisition module, adapted to acquire a security video of the surrounding environment, and extract a target event video corresponding to a basic event from the security video;

[0037] A sending module, adapted to determine a video cover image corresponding to the target event video, and send the video cover image to a server, so that the server receives the video cover image corresponding to the target event video sent by a smart home device, and performs image feature extraction processing on the video cover image based on a video description large model to obtain target image features, and performs video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0038] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes:

[0039] A processor; and

[0040] A memory arranged to store computer-executable instructions, and when the executable instructions are executed, the processor executes the method described in any one of the above.

[0041] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, and the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the method described in any one of the above is implemented.

[0042] The beneficial effects brought by the technical solutions provided in some embodiments of the present application at least include: receiving, by the server, a video cover image corresponding to a target event video sent by a smart home device, then using a video description large model to perform image feature extraction processing on the video cover image to obtain target image features, and then using the video description large model to perform video description generation processing on the target image features to obtain a video description for the video cover image. By generating the video description of the video cover image, the technical problem that when a user views the video cover image, the user needs to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in low viewing efficiency of the user, is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is an exemplary system architecture diagram of a video description generation method provided by an embodiment of the present application;

[0045] Figure 2 It is a flowchart of a video description generation method provided by an embodiment of the present application;

[0046] Figure 3 It is a flowchart of a method for determining target image features provided by an embodiment of the present application;

[0047] Figure 4 It is a flowchart of a method for determining scene image features provided by an embodiment of the present application;

[0048] Figure 5 It is a flowchart of a method for obtaining a video description for a video cover image provided by an embodiment of the present application;

[0049] Figure 6 It is a flowchart of a method for determining a video description of a video association description of a video cover image provided by an embodiment of the present application;

[0050] Figure 7 It is a flowchart of a method for determining a video description large model provided by an embodiment of the present application;

[0051] Figure 8 It is a flowchart of a method for displaying a video cover image and a video description provided by an embodiment of the present application;

[0052] Figure 9Schematic diagram of an interface of a target event dynamic bar provided by an embodiment of the present application;

[0053] Figure 10 Schematic flow diagram of a video description generation method applied to a smart home device provided by an embodiment of the present application;

[0054] Figure 11 Schematic structural diagram of a video description generation device applied to a server provided by an embodiment of the present application;

[0055] Figure 12 Schematic structural diagram of a video description generation device applied to a smart home device provided by an embodiment of the present application;

[0056] Figure 13 Schematic structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners

[0057] To make the features and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope protected by the embodiments of the present application.

[0058] In the related art, after intelligent monitoring terminals monitor their respective scenarios, corresponding monitoring videos will be generated for users to view. To facilitate users to understand the content of the monitoring video, a video cover image corresponding to the video segment will be generated. However, it is often difficult for users to directly understand the brief content of the video segment based on the video cover image, and users still need to carefully identify it, making it difficult for users to directly determine the video segment they need to view. Instead, they need to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency for users.

[0059] To solve the existing technical problems, the present application provides a method for generating video descriptions, which is applied to a server. The method includes: receiving a video cover image corresponding to a target event video sent by a smart home device, where the video cover image is a cover image determined by the smart home device from the target event video corresponding to a basic event, and the target event video is an event video extracted by the smart home device from the security video of the surrounding environment it collects; performing image feature extraction processing on the video cover image based on a video description large model to obtain target image features; and performing video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image. By generating the video description of the video cover image, the technical problem that when a user views the video cover image, they need to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency for the user, is solved.

[0060] Please refer to Figure 1 , Figure 1 which is an exemplary system architecture diagram of a method for generating video descriptions provided by an embodiment of the present application.

[0061] As Figure 1 shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired communication links or wireless communication links. For example, the wired communication links include optical fibers, twisted pairs, or coaxial cables, and the wireless communication links include Bluetooth communication links, Wireless-Fidelity (Wi-Fi) communication links, or microwave communication links, etc.

[0062] The terminal 101 can interact with the server 103 through the network 102 to receive messages from the server 103 or send messages to the server 103. Alternatively, the terminal 101 can interact with the server 103 through the network 102 to further receive messages or data sent by other users to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various terminals, including but not limited to smart watches, smart phones, tablets, laptop portable computers, and desktop computers, etc. When the terminal 101 is software, it can be installed in the above-listed terminals, which can be implemented as multiple software or software modules (for example, used to provide distributed services), or can be implemented as a single software or software module, and no specific limitation is made here.

[0063] Server 103 may be a business server that provides various services. It should be noted that server 103 can be hardware or software. When server 103 is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When server 103 is software, it can be implemented as multiple software or software modules (such as those used to provide distributed services) or as a single software or software module, and specific limitations are not made here.

[0064] In the embodiments of the present application, server 103 may receive a video cover image corresponding to a target event video sent by a smart home device. The video cover image is a cover image determined by the smart home device from the target event video corresponding to the basic event, and the target event video is an event video extracted by the smart home device from the security video of the surrounding environment it collects; perform image feature extraction processing on the video cover image based on a video description large model to obtain target image features; perform video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0065] Terminal 101 may be a smart home device. Terminal 101 may collect the security video of the surrounding environment, extract the target event video corresponding to the basic event from the security video; determine the video cover image corresponding to the target event video, and send the video cover image to the server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, performs image feature extraction processing on the video cover image based on a video description large model to obtain target image features, and performs video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0066] It should be understood that Figure 1 the numbers of terminals, networks, and servers in

[0067] Please refer to Figure 2 Figure 2 which is a schematic flowchart of a video description generation method provided by the embodiments of the present application. The execution subject of the embodiments of the present application may be a server that executes the video description generation method, or a processor in the server that executes the video description generation method, or a video description generation service in the server that executes the video description generation method. For the convenience of description, the following takes the execution subject as the processor in the server as an example to introduce the specific execution process of the video description generation method.

[0068] As Figure 2 shown, the video description generation method is applied to a server, and the method includes:

[0069] ​S202: Receive the video cover image corresponding to the target event video sent by the smart home device. The video cover image is the cover image determined by the smart home device from the target event video corresponding to the basic event, and the target event video is the event video extracted by the smart home device from the security video collected from the surrounding environment.

[0070] Among them, smart home devices include but are not limited to smart monitoring devices, smart door locks with monitoring functions, and smart doorbells with monitoring functions, etc. After installing the smart home device, collect the security video of the surrounding environment through the smart home device. After obtaining the security video, perform video partitioning processing on the security video based on the basic events in the security video, so as to obtain the target event videos corresponding to each basic event. Basic events include but are not limited to unlocking, opening the door, ringing the doorbell, someone passing by, someone staying, etc. The target event video records the complete process of the corresponding basic event from occurrence to end. After performing video partitioning processing on the security video based on the basic events in the security video, there may be blank event videos that are not partitioned into target event videos, and no further processing needs to be performed on this blank event video. The surrounding environment of the smart home device includes but is not limited to the living room at home, the bedroom at home, or the doorway, etc.

[0071] After obtaining the target event video based on the smart home device, the smart home device performs frame image extraction processing on the target event video to obtain video frame images, performs information degree scoring processing on the video frame images based on the image information degree scoring model, obtains the information degree score corresponding to each video frame image, determines the target information degree score from the information degree scores, and uses the target video frame image corresponding to the target information degree score as the target cover image corresponding to the target event video. The information degree of an image refers to the amount of information and the quality of information contained in the image. The higher the information degree score, the higher the amount of information and the quality of information contained in the image, and thus the target information degree score required is determined from the information degree scores.

[0072] After the smart home device obtains the video cover image corresponding to the target event video, the smart home device sends the video cover image corresponding to the target event video to the server, and the server receives the video cover image corresponding to the target event video sent by the smart home device.

[0073] S204: Perform image feature extraction processing on the video cover image based on the video description large model to obtain the target image features.

[0074] Among them, the target image features include human image features, event image features, and scene image features. The human image features include at least one of, but are not limited to, facial features, body features, clothing features, expression features, motion features, etc. The facial features include eye features, nose features, mouth features, face shape features, etc.; the body features include height features, body type features, skin color features, body posture features, etc.; the clothing features include clothing style features, color features, accessory features, etc.; the expression features include expression change features such as eye expression features, mouth corner features, facial muscle features, etc.; the motion features include limb motion features, gesture features, motion trajectory features, etc.

[0075] The event image features include at least one of, but are not limited to, the features of events or behaviors included in the image. The scene image features include at least one of, but are not limited to, the visual features of the scenes monitored by smart home devices, such as shape, color, texture, motion, direction, contrast, sharpness, etc., and may also include date features, weather features, etc.

[0076] After the video description large model performs image feature extraction processing on the video cover image, the target image features for the video cover image are obtained.

[0077] S206: Based on the video description large model, perform video description generation processing on the target image features to obtain a video description for the video cover image.

[0078] Among them, after obtaining the target image features of the video cover image, use the video description large model to perform video description generation processing on the target image features to obtain a video description for the video cover image. The video description large model can be a multi-modal large model. A multi-modal large model is a large deep learning model that can process multiple modalities of data (such as text, image, audio, etc.). It is usually composed of multiple sub-models. Each sub-model specializes in processing one modality of data and integrates information from different modalities through cross-modal learning to achieve more accurate and comprehensive information extraction and understanding.

[0079] The video description of the video cover image obtained by processing the target image features based on the video description large model is used to connect and describe the people, events, and scenes in the video cover image in text. For example, if the person is grandma, the event is going home, and the scene is home, the corresponding video description generated can be "Grandma is going home"; another example is that if the person is a courier, the event is ringing the doorbell, and the scene is at the door, the corresponding video description generated can be "The courier is ringing the doorbell". People include but are not limited to family members, the elderly, babies, food delivery workers, neighbors, and staff. In addition, in the actual scenario, there can also be other animals, such as pets. Events include but are not limited to going home, visiting, walking the dog, delivering food or express, passing by, and cleaning. Events can also include corresponding event types, and different event types can correspond to different event type icons. Event types include unlocking / opening the door, ringing the doorbell, someone passing by, someone staying, etc.

[0080] In the embodiment provided in the present application, the server receives the video cover image corresponding to the target event video sent by the smart home device, then uses the video description large model to perform image feature extraction processing on the video cover image to obtain the target image features, and then uses the video description large model to perform video description generation processing on the target image features to obtain the video description for the video cover image. By generating the video description of the video cover image, the technical problem that when the user views the video cover image, they need to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency of the user, is solved.

[0081] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a process for determining target image features provided by an embodiment of the present application. As Figure 3 shown, in S204, performing image feature extraction processing on the video cover image based on the video description large model to obtain the target image features includes:

[0082] S302: Performing image person feature extraction processing on the video cover image based on the video description large model to obtain person image features, performing image event feature extraction processing on the video cover image to obtain event image features, and performing image scene feature extraction processing on the video cover image to obtain scene image features.

[0083] Among them, performing image person feature extraction processing on the video cover image based on the video description large model to obtain person image features includes:

[0084] Perform face feature recognition on the person in the video cover image based on the video description large model to obtain face recognition features; and / or, perform industry feature recognition on the clothing of the person in the video cover image based on the video description large model to obtain industry category features; and / or, perform action feature recognition on the actions of the person in the video cover image based on the video description large model to obtain person action features; determine the person image features based on at least one of the face recognition features, industry category features, and person action features.

[0085] The face recognition features include facial features and expression features, the industry category features include clothing features, the person action features include body features and action features. The face recognition features are used to represent the identity of the person, and of course, can also represent the emotions of the person. The industry category features can also represent the occupation of the person, so as to determine the occupation information of the recognized person. The person action features can be used to represent the actions performed by the person.

[0086] In addition, perform image event feature extraction processing on the video cover image to obtain event image features, including:

[0087] Perform action feature recognition on the actions of the person in the video cover image based on the multi-modal large model to obtain person action features; determine the objects associated with the actions of the person, and extract the object features of the objects based on the multi-modal large model; perform feature fusion on the person action features and the object features to generate the event image features of the video cover image.

[0088] It is easy to understand that the person's actions combined with the objects associated with the person's actions can form corresponding events. Therefore, the event image features can be determined by feature fusion of the person action features and the object features. The person action features can be used to represent the actions performed by the person, and the object features are used to represent the object information, such as the object information can be the object name, object description information, etc.

[0089] At the same time, perform image scene feature extraction processing on the video cover image to obtain scene image features. The scene image features include, but are not limited to, the visual features of the scene monitored by smart home devices, such as shape, color, texture, motion, direction, contrast, clarity, etc., and can also include date features and weather features, etc.

[0090] S304: Obtain the target image features based on the person image features, event image features, and scene image features.

[0091] Among them, after obtaining the person image features, event image features, and scene image features, feature association can be performed on the person image features, event image features, and scene image features corresponding to the same video cover image, so as to obtain the target image features after feature association. The means of feature association include, but are not limited to, feature splicing, etc.

[0092] In the embodiments provided in this application, an image person feature extraction process, an image event feature extraction process, and an image scene feature extraction process are performed on the video cover image by using a video description large model, so as to obtain person image features, event image features, and scene image features. Then, target image features are obtained based on the person image features, event image features, and scene image features, so that the target image features have both person image features, event image features, and scene image features.

[0093] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a process for determining scene image features provided by an embodiment of this application. As Figure 4 shown, the image scene feature extraction process performed on the video cover image in S302 to obtain the scene image features includes:

[0094] S402: Obtain the basic scene set for the smart home device, and the basic scene features corresponding to the basic scene.

[0095] Among them, after the smart home device is installed, usually its corresponding environment remains fixed within a certain period of time. For example, the smart home device can be a smart door lock with a monitoring function. When the door is in the closed state, the scene monitored by the smart door lock is the doorway. Then, the background environment that has not changed for a long time corresponding to the doorway is the basic scene. The basic scene set for the smart home device can be the background environment that has not changed within a preset duration monitored by the smart home device.

[0096] After obtaining the basic scene set for the smart home device, the basic scene features corresponding to the basic scene can be extracted based on the video description large model. The basic scene features include the objects in the scene and their attributes, such as quantity, position, shape, color, etc.

[0097] S404: Perform a differential feature recognition process on the cover scene and the basic scene of the video cover image based on the video description large model to obtain target differential features.

[0098] Among them, since the video cover image is the cover image determined by the smart home device from the target event video corresponding to the basic event, there are usually differences between the cover scene and the basic scene of the video cover image. For example, there is an extra express box or a broom on the ground. Therefore, it is necessary to further accurately identify the differences between the cover scene and the basic scene. Thus, a differential feature recognition process is performed on the cover scene and the basic scene of the video cover image based on the video description large model to accurately identify the target differential features of the cover scene compared to the basic scene. Here, the cover scene usually does not include the people in the scene.

[0099] S406: Determine the scene image features based on the basic scene features and the target difference features.

[0100] Among them, the scene image features corresponding to the video cover image pay more attention to the features different from the ordinary basic scene. Therefore, the target difference features of the cover scene compared with the basic scene are extracted separately, and then the scene image features are determined through the basic scene features and the target difference features. Furthermore, the scene image features can pay more attention to the target difference features, which is convenient for subsequent targeted description of the scene.

[0101] In the embodiment provided by the present application, based on the video description large model, the difference feature recognition processing is performed on the cover scene and the basic scene of the video cover image, so as to accurately identify the target difference features of the cover scene compared with the basic scene. Then, the scene image features are determined through the basic scene features and the target difference features. Furthermore, the scene image features can pay more attention to the target difference features, which is convenient for subsequent targeted description of the scene.

[0102] Please refer to Figure 5 , Figure 5 , which is a schematic flow chart of a method for obtaining a video description of a video cover image provided by an embodiment of the present application. As Figure 5 shown, in S206, the video description generation processing is performed on the target image features based on the video description large model to obtain a video description of the video cover image, including:

[0103] S502: Perform person description generation processing on the person image features through the video description large model to obtain a person description, perform event description generation processing on the event image features to obtain an event description, and perform scene description processing on the scene image features to obtain a scene description.

[0104] Among them, the video description large model can use the preset person face database to perform person recognition on the person image features. For example, a preset person face database can be established based on the corresponding relationship between the face information of the pre-recorded person and the person's identity. After person recognition, the corresponding identity of the person can be determined, such as mom, dad, grandpa, grandma, or a specific person's name, etc.

[0105] When the video description large model uses the preset person face database to perform person recognition on the person image features and the recognition fails, the occupation of the person can be determined based on the person image features. For example, the occupation information of the recognized person can be determined based on the industry category features in the person image features, such as a courier, etc.

[0106] When the occupation information corresponding to the person cannot be determined, the person can be described based on the person action features in the person image features, such as a person holding paper and a pen, a person squatting, etc.

[0107] In addition, when generating an event description by processing event image features, the event image features can be determined by fusing the features of the person's actions and the object features. Thus, the event descriptions generated by processing the event image features can be "holding a box", "holding an express delivery", "sweeping the floor with a broom", etc.

[0108] Since the scene image features are determined by the basic scene features and the target difference features, when the scene image features do not change compared with the basic scene features, the scene description corresponding to the scene image features can be omitted; when the scene image features change compared with the basic scene features, the description corresponding to the changed features can be emphasized in the scene description corresponding to the scene image features. For example, the basic scene feature can be the doorway, and the scene description corresponding to the scene image features can be the doorway filled with express deliveries.

[0109] S504: Perform description association processing on the person description, event description, and scene description through the video description large model to obtain the video description of the video cover image.

[0110] Among them, after obtaining the person description, event description, and scene description, determine the associations among the person, event, and scene. For example, the person is grandma, the event associated with grandma is cleaning, and the scene associated with the event is the doorway. Then the video description of the video cover image can be "Grandma is cleaning at the doorway". Usually, since the scene is fixed, the scene description can be omitted, and the video description of the video cover image can be "Grandma is cleaning". Usually, the number of words in the video description of the video cover image is less than or equal to the preset number of words, and the preset number of words can be adjusted based on the interface display area corresponding to the video description. For example, the preset number of words can be 15 or 10.

[0111] In the embodiments provided in this application, the corresponding person description, event description, and scene description are determined based on the person image features, event image features, and scene image features, and then description association processing is performed on the person description, event description, and scene description, so as to obtain the video description of the video cover image.

[0112] Please refer to Figure 6 , Figure 6 which is a schematic flowchart of the process for determining the video description of the video association description of the video cover image provided by the embodiments of this application. As Figure 6 shown, after performing video description generation processing on the target image features based on the video description large model in S206 to obtain the video description of the video cover image, it further includes:

[0113] S602: Obtain the associated video cover image corresponding to the video cover image, and determine the associated video description corresponding to the associated video cover image based on the video description large model; wherein, the associated video cover image is the cover image corresponding to the associated event video, and the associated event video is determined from all event videos according to the time correlation of the target event video.

[0114] Among them, the video cover image is the cover image corresponding to the target event video, and the associated video cover image is the cover image corresponding to the associated event video. Obtain all event videos, perform chronological sorting on all event videos based on the time sequence, and then determine the associated event video corresponding to the target event video based on the chronological sorting of all event videos. For example, use the previous one or more event videos before the target event video as the associated event video corresponding to the target event video, and / or use the next one or more event videos after the target event video as the associated event video corresponding to the target event video. The time correlation can be understood as the chronological correlation of event occurrence, such as events occurring adjacent before or adjacent after.

[0115] After determining the associated video cover image corresponding to the video cover image, generate a similar method based on the video description of the video cover image, and generate the associated video description corresponding to the associated video cover image based on the video description large model.

[0116] S604: Perform video description association processing on the video description based on the associated video description to obtain the video association description for the video cover image.

[0117] Among them, after obtaining the associated video description corresponding to the associated video cover image, use the associated video description to perform video description association processing on the video description, so as to obtain the video association description for the video cover image.

[0118] Video description association processing is used to supplement and fuse the video description of the video cover image based on the associated video description, so as to provide users with more details. For example, the video description of the video cover image is "Mom stands at the door", but the associated video description corresponding to the associated video cover image of the associated event video at the next moment is "Mom opens the door", then the video association description generated after supplementing the video description of the video cover image based on the associated video description can be "Mom comes home".

[0119] In the embodiments provided in this application, supplement and fuse the video description of the video cover image based on the associated video description, so as to use the video cover images in different time periods to provide users with more details, and avoid the problem that the state at the moment corresponding to the video cover image is difficult to accurately reflect the content of the target event video.

[0120] Please refer toFigure 7 , Figure 7 is a schematic flowchart of a process for determining a video description large model provided by an embodiment of the present application. As Figure 7 shown, the method includes:

[0121] S702: Create an initial video description large model based on a basic large model.

[0122] Among them, the basic large model can be a multimodal large model. A multimodal large model is a large deep learning model that can process various modal data (such as text, images, audio, etc.). It is usually composed of multiple sub-models, each sub-model specializes in processing one type of modal data, and through cross-modal learning, the information of different modalities is integrated together to achieve more accurate and comprehensive information extraction and understanding. When using the basic large model to create an initial video description large model, the basic large model can be preliminarily trained to obtain the initial video description large model; or the model parameters of the basic large model can be adaptively adjusted to obtain the initial video description large model.

[0123] S704: Obtain a sample video cover image and annotate the sample video description for the sample video cover image.

[0124] Among them, after obtaining the sample video cover image, determine the sample video description corresponding to the sample video cover image, and then annotate the sample video description for the sample video cover image.

[0125] S706: Input the sample video cover image into the initial video description large model for model training. Through the initial video description large model, perform image feature extraction processing on the sample video cover image to obtain sample image features, and perform video description generation processing on the sample image features to obtain a reference video description for the sample video cover image.

[0126] Among them, input the sample video cover image into the initial video description large model, perform image feature extraction processing to obtain sample image features, and perform video description generation processing on the sample image features to obtain a reference video description for the sample video cover image. Thus, generate a reference video description for the sample video cover image corresponding to the initial video description large model.

[0127] S708: During the model training process, adjust the model parameters of the initial video description large model based on the sample video description and the reference video description to obtain a video description large model.

[0128] Among them, during the model training process, a model loss function is constructed based on the parameters corresponding to the sample video description and the reference video description respectively, so as to determine the model loss value of the initial video description large model by using the sample video description and the reference video description based on the model loss function. Then, the model parameters of the initial video description large model are adjusted by using the model loss value. After that, the model training is carried out again until the initial video description large model converges after the model parameters of the initial video description large model are adjusted, and a video description large model is obtained.

[0129] In the embodiment provided by the present application, the initial video description large model is trained by using the sample video cover image and the corresponding sample video description, so that the initial video description large model learns the features of video description based on the video cover image. During the model training process, the model parameters of the initial video description large model are adjusted based on the sample video description and the reference video description generated by the initial video description large model, and a video description large model is obtained.

[0130] Please refer to Figure 8 , Figure 8 which is a schematic flowchart of displaying a video cover image and a video description provided by an embodiment of the present application. Please refer to Figure 9 , Figure 9 which is a schematic interface diagram of a target event dynamic bar provided by an embodiment of the present application. As Figure 8 shown, the method includes:

[0131] S802: Determine the target event dynamic bar corresponding to the target event video in the smart home management interface, and display the video cover image and the video description in the target event dynamic bar.

[0132] Among them, when the smart home device is a door lock, determine the target event dynamic bar corresponding to the target event video in the door lock management interface, then display the target event dynamic bar, and display the video cover image and the video description in the target event dynamic bar.

[0133] As Figure 9 shown, the video descriptions include "Mom is holding the baby and coming home", "Grandma is coming home", "The courier is ringing the doorbell and holding a courier package", "The Meituan deliveryman is ringing the doorbell (not answered)", "Someone is cleaning", "Two female workers are standing at the door, holding a pen and paper", and the corresponding video cover images are on the right. Figure 9 The specific images are not shown in , and the specific images will be shown during the actual display process. In addition, the event type icon in the upper left corner of the video cover image can also indicate the event type, such as the doorbell icon corresponding to ringing the doorbell, or icons related to people passing by or staying. Among them, the intelligent brain is the server corresponding to the door lock.

[0134] The time can also be displayed in the target event dynamic bar, and at the same time, a selection area for time selection can be provided, such as Figure 9 "Today" in, and clicking on "All Messages" can filter the video description. A delete button is also displayed in the upper right corner of the target event dynamic bar, and can be deleted after selecting the corresponding video cover image and video description.

[0135] S804: In response to a video viewing request for the target event dynamic bar, display the target event video corresponding to the video cover image.

[0136] Among them, as Figure 9 shown, the red font shows the targeted display after clicking on "Meituan deliveryman ringing the doorbell (not answered)" based on the video viewing request. After "Meituan deliveryman ringing the doorbell (not answered)", it is to prompt the user of the target display area corresponding to the video viewing request, and at this time, play the corresponding target event video.

[0137] In the embodiment provided by the present application, by displaying the video cover image and the video description in the target event dynamic bar, the user can quickly understand the video content corresponding to the target event video, which is convenient for the user to view targeted and quickly.

[0138] Please refer to Figure 10 , Figure 10 which is a schematic flowchart of a video description generation method applied to a smart home device provided by an embodiment of the present application. The execution subject of the embodiment of the present application can be a smart home device that executes the video description generation method, or a processor in a smart home device that executes the video description generation method, or a video description generation service in a smart home device that executes the video description generation method. For the convenience of description, the following takes the execution subject as the processor in the smart home device as an example to introduce the specific execution process of the video description generation method.

[0139] As Figure 10 shown, the video description generation method is applied to a smart home device, and the method includes:

[0140] S1002: Collect the security video of the environment where it is located, and extract the target event video corresponding to the basic event from the security video.

[0141] Among them, the smart home device includes but is not limited to a smart monitoring device, a smart door lock with a monitoring function, and a smart doorbell with a monitoring function, etc. The smart home device specifically includes: a collection module, a storage module, a transmission module, etc.

[0142] The acquisition module is used to monitor the environment where the smart home device is located in real time and capture the monitored video data, such as the security video of the environment where the smart home device is located. The acquisition module includes, but is not limited to, a monitoring camera, etc. The storage module is a device for storing the captured monitored video data. The storage module includes, but is not limited to, a hard disk, a memory card, etc. The transmission device is used to transmit the captured monitored video data to other devices or servers. The transmission device includes, but is not limited to, a network switch, etc.

[0143] The security video is usually a continuous video of a set duration. For example, it can be a continuous video without interruption for 24 hours, that is, the smart home device generates a security video every 24 hours of recording; or it can be a continuous video without interruption for 8 hours, that is, the smart home device generates a security video every 8 hours of recording. The continuous video of the set duration here can be set accordingly based on user needs and is not limited here.

[0144] After the security video is acquired, video division processing is performed on the security video based on the basic events in the security video, so as to obtain the target event videos corresponding to the respective basic events. The basic events include, but are not limited to, unlocking, opening the door, ringing the doorbell, someone passing by, someone staying, etc. The target event video records the complete process of the corresponding basic event from occurrence to end. After video division processing is performed on the security video based on the basic events in the security video, there may be blank event videos that are not divided into target event videos, and no further processing needs to be performed on this blank event video.

[0145] S1004: Determine the video cover image corresponding to the target event video, and send the video cover image to the server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, and performs image feature extraction processing on the video cover image based on the video description large model to obtain the target image features, and performs video description generation processing on the target image features based on the video description large model to obtain the video description for the video cover image.

[0146] Among them, the relevant description of S1004 can refer to S202 to S206.

[0147] In the embodiment provided in the present application, through the cooperation of the server and the smart home device, after the server receives the video cover image corresponding to the target event video sent by the smart home device, the video description large model is used to perform image feature extraction processing on the video cover image to obtain the target image features, and then the video description large model is used to perform video description generation processing on the target image features to obtain the video description for the video cover image. By generating the video description of the video cover image, the technical problem that when the user views the video cover image, the user needs to further identify the video cover image to determine the image content corresponding to the video cover image, resulting in a low viewing efficiency of the user, is solved.

[0148] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a video description generation device applied to a server provided by an embodiment of the present application. As Figure 11 shown, the video description generation device 1100 includes:

[0149] A receiving module 1110, adapted to receive a video cover image corresponding to a target event video sent by a smart home device, where the video cover image is a cover image determined by the smart home device from the target event video corresponding to a basic event, and the target event video is an event video extracted by the smart home device from the security video collected in the environment where it is located;

[0150] An image feature extraction module 1120, adapted to perform image feature extraction processing on the video cover image based on a video description large model to obtain target image features;

[0151] A video description module 1130, adapted to perform video description generation processing on the target image features based on a video description large model to obtain a video description for the video cover image.

[0152] Optionally, the image feature extraction module 1120 includes:

[0153] A feature extraction unit, adapted to perform image character feature extraction processing on the video cover image based on a video description large model to obtain character image features, perform image event feature extraction processing on the video cover image to obtain event image features, and perform image scene feature extraction processing on the video cover image to obtain scene image features;

[0154] A target image feature determination unit, adapted to obtain target image features based on the character image features, event image features, and scene image features.

[0155] Optionally, the feature extraction unit includes:

[0156] An acquisition subunit, adapted to acquire a basic scene set for the smart home device and basic scene features corresponding to the basic scene;

[0157] A target difference feature determination subunit, adapted to perform difference feature recognition processing on the cover scene and the basic scene of the video cover image based on a video description large model to obtain target difference features;

[0158] A scene image feature determination subunit, adapted to determine scene image features based on the basic scene features and the target difference features.

[0159] Optionally, the video description module 1130 includes:

[0160] A description generation unit, adapted to perform person description generation processing on person image features through a video description large model to obtain a person description, perform event description generation processing on event image features to obtain an event description, and perform scene description processing on scene image features to obtain a scene description;

[0161] A video description generation unit, adapted to perform description association processing on the basis of the person description, event description, and scene description through a video description large model to obtain a video description of the video cover image.

[0162] Optionally, the video description generation device 1100 further includes:

[0163] An associated video description determination module, adapted to obtain an associated video cover image corresponding to the video cover image, and determine an associated video description corresponding to the associated video cover image based on the video description large model; wherein, the associated video cover image is the cover image of the associated event video, and the associated event video is determined from all event videos according to the time relevance of the target event video;

[0164] A video associated description determination module, adapted to perform video description association processing on the video description based on the associated video description to obtain a video associated description for the video cover image.

[0165] Optionally, the video description generation device 1100 further includes:

[0166] A creation module, adapted to create an initial video description large model based on the basic large model;

[0167] A labeling module, adapted to obtain a sample video cover image and label a sample video description for the sample video cover image;

[0168] A model training module, adapted to input the sample video cover image into the initial video description large model for model training, perform image feature extraction processing on the sample video cover image through the initial video description large model to obtain sample image features, and perform video description generation processing on the sample image features to obtain a reference video description for the sample video cover image;

[0169] A model parameter adjustment module, adapted to adjust the model parameters of the initial video description large model based on the sample video description and the reference video description during the model training process to obtain the video description large model.

[0170] Optionally, the video description generation device 1100 further includes:

[0171] A display module, adapted to determine a target event dynamic bar corresponding to the target event video in the smart home management interface, and display the video cover image and the video description in the target event dynamic bar;

[0172] A response module, adapted to respond to a video viewing request for a dynamic bar of a target event and display a target event video corresponding to a video cover image.

[0173] Please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a video description generation device applied to a smart home device provided by an embodiment of the present application. As Figure 12 shown, the video description generation device 1200 includes:

[0174] An acquisition module 1210, adapted to acquire a security video of the surrounding environment and extract a target event video corresponding to a basic event from the security video;

[0175] A sending module 1220, adapted to determine a video cover image corresponding to the target event video, send the video cover image to the server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, and perform image feature extraction processing on the video cover image based on a video description large model to obtain target image features, and perform video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image.

[0176] In an embodiment of the present application, there is also provided a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and when the one or more programs are executed by a processor, the methods described in any one of the above are implemented.

[0177] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a terminal provided by an embodiment of the present application. As Figure 13 shown, the terminal 1300 may include: at least one processor 1301, at least one network interface 1304, a user interface 1303, a memory 1305, and at least one communication bus 1302.

[0178] Among them, the communication bus 1302 is used to implement connection communication between these components.

[0179] Among them, the user interface 1303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1303 may further include a standard wired interface and a wireless interface.

[0180] Among them, the network interface 1304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0181] Among them, the processor 1301 may include one or more processing cores. The processor 1301 connects various parts within the entire terminal 1300 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1305, and by calling the data stored in the memory 1305, it executes various functions of the terminal 1300 and processes data. Optionally, the processor 1301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1301 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 1301 and may be implemented separately by a single chip.

[0182] Among them, the memory 1305 may include random access memory (RAM) and may also include read-only memory (ROM). Optionally, the memory 1305 includes a non-transitory computer-readable storage medium. The memory 1305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1305 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1305 may also be at least one storage device located far from the aforementioned processor 1301. As Figure 13 shown, the memory 1305, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a video description generation program applied to a server and / or a video description generation program applied to a smart home device.

[0183] In Figure 13In the terminal 1300 shown, the user interface 1303 is mainly used to provide an interface for the user to input and obtain the data input by the user; the processor 1301 can be used to call the video description generation program stored in the memory 1305 and applied to the server, and specifically perform the following operations:

[0184] Receive the video cover image corresponding to the target event video sent by the smart home device, where the video cover image is the cover image determined by the smart home device from the target event video corresponding to the basic event, and the target event video is the event video extracted by the smart home device from the security video collected in the environment where it is located;

[0185] Perform image feature extraction processing on the video cover image based on the video description large model to obtain the target image features;

[0186] Perform video description generation processing on the target image features based on the video description large model to obtain the video description for the video cover image.

[0187] Optionally, when the processor 1301 performs image feature extraction processing on the video cover image based on the video description large model to obtain the target image features, it specifically performs:

[0188] Perform image character feature extraction processing on the video cover image based on the video description large model to obtain character image features, perform image event feature extraction processing on the video cover image to obtain event image features, and perform image scene feature extraction processing on the video cover image to obtain scene image features;

[0189] Obtain the target image features based on the character image features, event image features, and scene image features.

[0190] Optionally, when the processor 1301 performs image scene feature extraction processing on the video cover image to obtain scene image features, it specifically performs:

[0191] Obtain the basic scene set for the smart home device and the basic scene features corresponding to the basic scene;

[0192] Perform differential feature recognition processing on the cover scene of the video cover image and the basic scene based on the video description large model to obtain the target differential features;

[0193] Determine the scene image features based on the basic scene features and the target differential features.

[0194] Optionally, when the processor 1301 performs video description generation processing on the target image features based on the video description large model to obtain the video description for the video cover image, it specifically performs:

[0195] The large model for video description performs person description generation processing on person image features to obtain a person description, performs event description generation processing on event image features to obtain an event description, and performs scene description processing on scene image features to obtain a scene description;

[0196] The large model for video description performs description association processing based on the person description, event description, and scene description to obtain a video description of the video cover image.

[0197] Optionally, after the processor 1301 performs video description generation processing on the target image features based on the large model for video description to obtain a video description for the video cover image, it specifically further performs:

[0198] Obtain an associated video cover image corresponding to the video cover image, and determine an associated video description corresponding to the associated video cover image based on the large model for video description; wherein, the associated video cover image is the cover image of the associated event video, and the associated event video is determined from all event videos according to time relevance by the target event video;

[0199] Perform video description association processing on the video description based on the associated video description to obtain a video association description for the video cover image.

[0200] Optionally, the processor 1301 is further adapted to perform:

[0201] Create an initial large model for video description based on the basic large model;

[0202] Obtain a sample video cover image, and label a sample video description for the sample video cover image;

[0203] Input the sample video cover image into the initial large model for video description for model training. Through the initial large model for video description, perform image feature extraction processing on the sample video cover image to obtain sample image features, and perform video description generation processing on the sample image features to obtain a reference video description for the sample video cover image;

[0204] During the model training process, adjust the model parameters of the initial large model for video description based on the sample video description and the reference video description to obtain the large model for video description.

[0205] Optionally, the processor 1301 is further adapted to perform:

[0206] Determine a target event dynamic bar corresponding to the target event video in the smart home management interface, and display the video cover image and the video description in the target event dynamic bar;

[0207] In response to a video viewing request for the target event dynamic bar, display the target event video corresponding to the video cover image

[0208] In Figure 13 In the terminal 1300 shown in Figure 13 , the user interface 1303 is mainly used to provide an interface for the user to input and obtain the data input by the user. The processor 1301 can be used to call the video description generation program stored in the memory 1305 and applied to the smart home device, and specifically perform the following operations:

[0209] Collect the security video of the surrounding environment and extract the target event video corresponding to the basic event from the security video;

[0210] Determine the video cover image corresponding to the target event video, and send the video cover image to the server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, and performs image feature extraction processing on the video cover image based on the video description large model to obtain the target image features, and performs video description generation processing on the target image features based on the video description large model to obtain the video description for the video cover image.

[0211] In several embodiments provided in the embodiments of the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical or other form.

[0212] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0213] In addition, in each embodiment of the embodiments of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0214] When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0215] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the embodiments of the present application.

[0216] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0217] The above is the description of a video description generation method, device, terminal, and computer-readable storage medium provided by the embodiments of the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the embodiments of the present application.

Claims

1. A video description generation method, applied to a server, wherein: The method comprises: Receiving a video cover image corresponding to a target event video sent by a smart home device, wherein the video cover image is a cover image determined by the smart home device from a target event video corresponding to a basic event, and the target event video is an event video extracted by the smart home device from a security video collected by the smart home device in the environment in which the event is located; Performing image feature extraction processing on the video cover image based on the video description large model to obtain target image features; Performing video description generation processing on the target image features based on the video description macro model to obtain a video description for the video cover image; The step of performing image feature extraction processing on the video cover image based on the video description large model to obtain target image features includes: Based on the video description big model, the video cover image is subjected to image character feature extraction processing to obtain character image features, the video cover image is subjected to image event feature extraction processing to obtain event image features, and the video cover image is subjected to image scene feature extraction processing to obtain scene image features; Obtaining target image features based on the person image features, the event image features and the scene image features; The step of performing image scene feature extraction processing on the video cover image to obtain scene image features includes: Acquire a basic scene set for the smart home device and a basic scene feature corresponding to the basic scene; Based on the video description large model, the cover scene and the basic scene of the video cover image are processed for distinguishing features to obtain target distinguishing features; A scene image feature is determined based on the basic scene feature and the target distinguishing feature.

2. The method according to claim 1, wherein: The step of performing video description generation processing on the target image features based on the video description macro model to obtain a video description for the video cover image includes: The character image features are processed by the video description big model to generate character descriptions to obtain character descriptions, the event image features are processed to generate event descriptions to obtain event descriptions, and the scene image features are processed to generate scene descriptions to obtain scene descriptions; The video description of the video cover image is obtained by performing description association processing based on the character description, the event description and the scene description through the video description big model.

3. The method according to claim 1, wherein: After performing video description generation processing on the target image features based on the video description large model to obtain a video description for the video cover image, the method further includes: Obtain an associated video cover image corresponding to the video cover image, and determine an associated video description corresponding to the associated video cover image based on the video description macromodel; wherein the associated video cover image is a cover image corresponding to an associated event video, and the associated event video is the target event video determined from all event videos according to time relevance; The video description is subjected to video description association processing based on the associated video description to obtain a video associated description for the video cover image.

4. The method according to claim 1, wherein: The method further comprises: Creating an initial video description big model based on the basic big model; Obtaining a sample video cover image, and annotating the sample video cover image with a sample video description; Inputting the sample video cover image into the initial video description large model for model training, performing image feature extraction processing on the sample video cover image through the initial video description large model to obtain sample image features, and performing video description generation processing on the sample image features to obtain a reference video description for the sample video cover image; During the model training process, the model parameters of the initial video description large model are adjusted based on the sample video description and the reference video description to obtain a video description large model.

5. The method according to claim 1, wherein: The method further comprises: Determine a target event dynamic bar corresponding to the target event video in a smart home management interface, and display the video cover image and the video description in the target event dynamic bar; In response to a video viewing request for the target event dynamic column, the target event video corresponding to the video cover image is displayed.

6. A video description generation method, applied to smart home devices, wherein: The method comprises: Collect security video of the environment, and extract target event video corresponding to the basic event from the security video; Determine a video cover image corresponding to the target event video, send the video cover image to a server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, and performs image feature extraction processing on the video cover image based on a video description big model to obtain target image features, and performs video description generation processing on the target image features based on the video description big model to obtain a video description for the video cover image; The step of performing image feature extraction processing on the video cover image based on the video description large model to obtain target image features includes: Based on the video description big model, the video cover image is subjected to image character feature extraction processing to obtain character image features, the video cover image is subjected to image event feature extraction processing to obtain event image features, and the video cover image is subjected to image scene feature extraction processing to obtain scene image features; Obtaining target image features based on the person image features, the event image features and the scene image features; The step of performing image scene feature extraction processing on the video cover image to obtain scene image features includes: Acquire a basic scene set for the smart home device and a basic scene feature corresponding to the basic scene; Based on the video description large model, the cover scene and the basic scene of the video cover image are processed for distinguishing features to obtain target distinguishing features; A scene image feature is determined based on the basic scene feature and the target distinguishing feature.

7. A video description generating device, applied to a server, wherein: The device comprises: A receiving module, adapted to receive a video cover image corresponding to a target event video sent by a smart home device, wherein the video cover image is a cover image determined by the smart home device from a target event video corresponding to a basic event, and the target event video is an event video extracted by the smart home device from a security video collected by the smart home device in the environment in which the event is located; An image feature extraction module, adapted to perform image feature extraction processing on the video cover image based on the video description macro model to obtain target image features; A video description module, adapted to perform video description generation processing on the target image features based on the video description macro model to obtain a video description for the video cover image; Wherein, the image feature extraction module includes: A feature extraction unit, adapted to perform image character feature extraction processing on the video cover image based on the video description macro model to obtain character image features, perform image event feature extraction processing on the video cover image to obtain event image features, and perform image scene feature extraction processing on the video cover image to obtain scene image features; a target image feature determination unit, adapted to obtain target image features based on the character image features, the event image features and the scene image features; Wherein, the feature extraction unit comprises: An acquisition subunit, adapted to acquire a basic scene set for the smart home device, and a basic scene feature corresponding to the basic scene; a target distinguishing feature determination subunit, adapted to perform distinguishing feature recognition processing on the cover scene and the basic scene of the video cover image based on the video description macro model to obtain target distinguishing features; The scene image feature determination subunit is adapted to determine the scene image feature based on the basic scene feature and the target distinguishing feature.

8. A video description generating device, applied to a smart home device, wherein: The device comprises: A collection module, adapted to collect security videos of the environment, and extract target event videos corresponding to basic events from the security videos; A sending module, adapted to determine a video cover image corresponding to the target event video, send the video cover image to a server, so that the server receives the video cover image corresponding to the target event video sent by the smart home device, and performs image feature extraction processing on the video cover image based on a video description macro model to obtain target image features, and performs video description generation processing on the target image features based on the video description macro model to obtain a video description for the video cover image; The step of performing image feature extraction processing on the video cover image based on the video description large model to obtain target image features includes: Based on the video description big model, the video cover image is subjected to image character feature extraction processing to obtain character image features, the video cover image is subjected to image event feature extraction processing to obtain event image features, and the video cover image is subjected to image scene feature extraction processing to obtain scene image features; Obtaining target image features based on the person image features, the event image features and the scene image features; The step of performing image scene feature extraction processing on the video cover image to obtain scene image features includes: Acquire a basic scene set for the smart home device and a basic scene feature corresponding to the basic scene; Based on the video description large model, the cover scene and the basic scene of the video cover image are processed for distinguishing features to obtain target distinguishing features; A scene image feature is determined based on the basic scene feature and the target distinguishing feature.

9. A terminal, wherein: The terminal includes: Processor; and A memory arranged to store computer executable instructions which, when executed, cause the processor to perform a method according to any one of claims 1 to 5 or 6.

10. A computer-readable storage medium, wherein: The computer-readable storage medium stores one or more programs, which, when executed by a processor, implement the method of any one of claims 1 to 5 or 6.

Citation Information

Patent Citations

  • Video title generation method and device, equipment, storage medium and program product

    CN116958866A

  • Video description text generation method and device, electronic equipment and storage medium

    CN118075551A

  • Non-transitory computer-readable recording medium, abnormality transmission method, and information processing apparatus

    US20240071082A1