Video generation method, electronic equipment and computer readable storage medium

By identifying the person information in the user's instructions and obtaining its confidence, it is determined whether to enter the person confirmation process, and the portrait clustering information is used to obtain accurate materials from the gallery. This solves the problem of inaccurate materials in the existing technology and achieves the effect of generating videos that better meet user expectations.

CN120676218AActive Publication Date: 2025-09-19HONOR DEVICE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202410284137.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-19
Estimated Expiration
2044-03-12

AI Technical Summary

Technical Problem

In the prior art, the materials obtained by the electronic device from the gallery application according to the user's instructions may not be accurate, resulting in the generated video not meeting the user's expectations.

Method used

By identifying the person information in the user's instructions and obtaining its confidence, it is determined whether to enter the person confirmation process, and the portrait clustering information is used to obtain accurate materials from the gallery to generate a video.

Benefits of technology

The accuracy of the material is improved, making the generated video more in line with user expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676218A_ABST
    Figure CN120676218A_ABST
Patent Text Reader

Abstract

The invention relates to the field of terminals, in particular to a video generation method, electronic equipment and a computer readable storage medium. The method comprises the steps that when a user instruction is monitored, first character information contained in the user instruction is recognized; obtaining a confidence coefficient corresponding to the first character information, wherein the confidence coefficient is used for representing a credibility degree of an identified character; if it is determined that a figure confirmation process is executed according to the confidence coefficient corresponding to the first figure information, confirming a figure corresponding to the first figure information to obtain a confirmation result; and generating a video according to the confirmation result. Through the method, the problem that the generated video does not conform to user expectation due to inaccurate material acquisition in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminals, and in particular to a video generation method, an electronic device, and a computer-readable storage medium. Background Art

[0002] Currently, many electronic devices are equipped with cameras, allowing users to take photos, videos, and more. Electronic devices also have communication capabilities, allowing users to download images or videos from the internet or retrieve images or videos sent by other electronic devices. Users can edit images and / or videos stored on their electronic devices to create a new video from the selected images and / or videos. For example, a user can input user commands into the electronic device, which then retrieves the corresponding footage from a gallery application and creates a new video from the captured footage.

[0003] However, in actual applications, the material that an electronic device retrieves from a gallery application based on user instructions may not be accurate. For example, a user may enter the user instruction "Help me generate a video of my daughter" into an electronic device. However, the material retrieved from the gallery application based on this user instruction may include other people besides my daughter. In this case, if the electronic device generates a video based on the retrieved material, the generated video will not meet the user's expectations. Summary of the Invention

[0004] The present application provides a video generation method, an electronic device, and a computer-readable storage medium, which solve the problem in the prior art of inaccurate material acquisition, resulting in the generated video not meeting user expectations.

[0005] To achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, a video generation method is provided, comprising:

[0007] When a user instruction is detected, identifying first person information contained in the user instruction;

[0008] Obtaining a confidence level corresponding to the first person information, where the confidence level is used to indicate the credibility of the identified person;

[0009] If it is determined to execute the character confirmation process according to the confidence level corresponding to the first character information, the character corresponding to the first character information is confirmed to obtain a confirmation result;

[0010] A video is generated according to the confirmation result.

[0011] In an embodiment of the present application, after recognizing the user instruction, it is possible to determine whether to confirm the person based on the confidence level of the identified person information. This allows further confirmation of the identified person when the credibility level of the identified person is low, thereby facilitating the acquisition of more accurate material and making the generated video more in line with user expectations.

[0012] In an implementation of the first aspect, the first character information includes a character tag and character attributes of each character;

[0013] The obtaining of the confidence level corresponding to the first character information includes:

[0014] Obtaining a first character list from an electronic device, the first character list including learning information, the learning information being information obtained by the electronic device through classification processing of images in a preset gallery, each set of learning information including a character number, a character label, a character attribute, and a confidence level of a character;

[0015] The confidence level corresponding to the first character information is obtained according to the first character list.

[0016] In an implementation of the first aspect, obtaining the confidence level corresponding to the first character information according to the first character list includes:

[0017] Searching for target information in the first character list, where the target information is learning information that matches the character tag of the first character information;

[0018] If the target information is found in the first character list, the confidence level corresponding to the first character information is set as the confidence level of the target information;

[0019] If the target information is not found in the first character list, the confidence level corresponding to the first character information is set to null.

[0020] In an implementation of the first aspect, the method further includes:

[0021] If the confidence level corresponding to the first character information is empty or less than a preset threshold, determining to execute the character confirmation process;

[0022] If the confidence level corresponding to the person information is greater than a preset threshold, it is determined that the person confirmation process will not be executed.

[0023] In this embodiment of the present application, whether to proceed to the character verification process is determined based on the confidence level of the character information. A lower confidence level indicates a lower degree of credibility of the learned character information. In the case of a lower confidence level, the character verification process is initiated. This ensures that the acquired material is more consistent with the user's expectations, thereby improving the quality of the generated video.

[0024] In an implementation of the first aspect, confirming the person corresponding to the first person information to obtain a confirmation result includes:

[0025] Obtaining portrait clustering information from a preset gallery, wherein the portrait clustering information is information obtained by performing clustering processing on images in the preset gallery by a gallery application, and the portrait clustering information includes a character number, a character label, and a character attribute of each clustered character;

[0026] detecting whether a second character list is empty, the second character list including first character information existing in the first character list;

[0027] If the second character list is not empty, obtaining a first result matching the second character list from the preset gallery, and the confirmation result is the first result;

[0028] If the second character list is empty, the portrait clustering information is filtered according to the first character information to obtain a second result, and the confirmation result is the second result.

[0029] In the embodiment of the present application, when the character information in the user instruction is not learned, images that match the character attributes in the user instruction can be filtered from the image library, which is conducive to obtaining more accurate materials, so that the generated video better meets the user's expectations. When the character information in the user instruction is learned, images that match the character information in the user instruction can be obtained from the image library, thereby achieving accurate acquisition of materials, so that the generated video better meets the user's expectations.

[0030] In an implementation of the first aspect, obtaining a first result matching the second character list from the preset gallery includes:

[0031] For the second person information, determining whether the person number of the second person information exists in the portrait cluster information, wherein the second person information is person information in the second person list;

[0032] If the character number of the second character information exists in the portrait cluster information, storing the character number corresponding to the second character information into a third character list;

[0033] If there is no untraversed second character information in the current second character list, determining whether the third character list is empty;

[0034] If the third person list is empty, obtaining at least one set of third person information according to the portrait clustering information, and using the at least one set of third person information as the first result;

[0035] If the third character list is not empty, fourth character information matching the character number in the third character list is obtained from the preset gallery, and the fourth character information is used as the first result.

[0036] In an embodiment of the present application, when the character information in the user instruction is learned, images matching the character information in the user instruction can be obtained from the gallery, thereby achieving accurate acquisition of materials and making the generated video more in line with user expectations.

[0037] In an implementation of the first aspect, the portrait clustering information further includes a category number of each clustered character, where one category number corresponds to at least one character number;

[0038] The determining whether the character number of the second character information exists in the portrait cluster information includes:

[0039] Converting the character number of the second character information into a category number;

[0040] Determining whether the category number of the second person information exists in the portrait cluster information;

[0041] If so, determining whether the character number of the second character information exists in the portrait cluster information;

[0042] If not, it is determined that the character number of the second character information does not exist in the portrait cluster information.

[0043] In some application scenarios, if a user modifies or adds a person's label, the person's category number will be modified accordingly. This may result in the same category number corresponding to multiple different person numbers. If a second person's information is searched for in the portrait cluster information based on the person number, this may lead to inaccurate detection results. The above method can effectively improve the accuracy of detection results.

[0044] In an implementation of the first aspect, the filtering the portrait cluster information according to the first person information to obtain a second result includes:

[0045] Determine whether the first character list is empty;

[0046] If the first character list is empty, at least one set of third character information is obtained according to the portrait clustering information, and the at least one set of third character information is used as the second result.

[0047] In an embodiment of the present application, when the memory of the character information in the user instruction is not learned, images that meet the character attributes in the user instruction can be filtered out from the gallery, which is conducive to obtaining more accurate materials, so that the generated video is more in line with user expectations.

[0048] In an implementation of the first aspect, after determining whether the first character list is empty, the method further includes:

[0049] If the first character list is not empty, obtaining fifth character information, where the fifth character information is character information that has not been traversed in the portrait cluster information;

[0050] Determine whether the fifth person information exists in the first person list;

[0051] If the fifth character information exists in the first character list, determining whether the character attributes of the fifth character information match the character attributes of the first character information;

[0052] If the character attributes of the fifth character information match the character attributes of the first character information, the fifth character information is stored in the fourth character list, and the next fifth character information is obtained until all the portrait cluster information has been traversed;

[0053] When all the portrait cluster information has been traversed, determining whether the fourth person list is empty;

[0054] If the fourth person list is empty, obtaining at least one set of third person information according to the portrait clustering information, and using the at least one set of third person information as the second result;

[0055] If the fourth character list is not empty, obtaining sixth character information that matches the character number in the fourth character list from the preset gallery, and using the sixth character information as the second result.

[0056] In an embodiment of the present application, when the character information in the user instruction is not learned, images that meet the character attributes in the user instruction can be filtered out from the gallery, which is conducive to obtaining more accurate materials, so that the generated video is more in line with user expectations.

[0057] In an implementation of the first aspect, the method further includes:

[0058] The method further comprises:

[0059] If the portrait clustering information is not obtained from the preset gallery, a third result is sent, and the third result does not contain an error prompt information.

[0060] In an implementation of the first aspect, the method further includes:

[0061] If it is determined not to execute the person confirmation process based on the confidence level corresponding to the first person information, a video is generated based on the first person information.

[0062] In a second aspect, an electronic device is provided, the electronic device comprising one or more processors and a memory;

[0063] The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to perform the method as described in any one of the first aspects.

[0064] According to a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes instructions. When the instructions are executed on an electronic device, the electronic device executes the method according to any one of the first aspects.

[0065] In a fourth aspect, a chip system is provided, which is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of the first aspects.

[0066] In a fifth aspect, a computer program product is provided. When the computer program product is run on an electronic device or a wireless router, the electronic device can implement the method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0068] Figure 2 is a software structure block diagram of the electronic device 100 provided in an embodiment of the present application;

[0069] Figure 3 This is a schematic diagram of an interface for generating a video through a voice assistant provided in an embodiment of the present application;

[0070] Figure 4 This is a schematic diagram of the interface for character confirmation provided in an embodiment of the present application;

[0071] Figure 5 This is a schematic diagram of a character confirmation interface provided by another embodiment of the present application;

[0072] Figure 6 This is a schematic diagram of a character confirmation interface provided by another embodiment of the present application;

[0073] Figure 7This is a schematic diagram of the judgment interaction process of the character confirmation process provided by the embodiment of the present application;

[0074] Figure 8 This is a schematic diagram of the character confirmation process provided in an embodiment of the present application;

[0075] Figure 9 Schematic diagram of the process of obtaining the first result provided in the embodiment of the present application;

[0076] Figure 10 Schematic diagram of the process of obtaining the second result provided in the embodiment of the present application;

[0077] Figure 11 This is a flow chart of character material confirmation provided in an embodiment of the present application. DETAILED DESCRIPTION

[0078] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details.

[0079] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0080] It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more than two; "and / or" describes the relationship between associated objects, indicating that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0081] In addition, in the description of this application specification and the appended claims, the terms "first", "second", "third", "fourth", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0082] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0083] The present invention provides a method for generating a video. The method can be applied to an electronic device, such as a tablet computer, a mobile phone, a wearable device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or the like. The present invention does not limit the specific type of electronic device.

[0084] Figure 1 The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0085] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0086] The processor 110 may include one or more processing units, for example: the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. For example, the processor 110 is used to execute the video generation method in the embodiment of the present application.

[0087] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0088] The internal memory 121 can be used to store computer executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and an application required for at least one function (such as an image playback function, etc.). The touch sensor 180K is also called a "touch panel". The touch sensor 180K can be set on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be set on the surface of the electronic device 100, which is different from the position of the display screen 194.

[0089] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0090] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1. For example, in the embodiments of the present application, Figure 3-Figure 6 The interfaces shown are all displayed by the monitor.

[0091] The above is a specific description of the embodiments of the present application using the electronic device 100 as an example. It should be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. The electronic device 100 may have more or fewer components than shown in the figure, may combine two or more components, or may have a different component configuration. The various components shown in the figure may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.

[0092] In addition, an operating system runs on top of the above components. The operating system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100.

[0093] See also Figure 2 , is a software structure block diagram of the electronic device 100 provided in an embodiment of the present application. The layered architecture of the electronic device 100 divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime (Android Runtime) and system library, and the kernel layer.

[0094] The application layer can include a series of application packages, such as Figure 2 As shown, the application package can include applications such as a gallery, a voice assistant, a quick application engine, and a light editing service. For example, the voice assistant can parse the conversation content entered by the user and, based on the conversation parsing results, call other applications or middleware to generate videos. The quick application engine can provide a JS (JavaScript) card download service. The light editing service can generate videos based on searched static image materials and / or video materials.

[0095] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer, including various components and services to support developers' Android development. The application framework layer includes some predefined functions. Figure 2 As shown, the application framework layer may include a window manager, a content provider, a notification manager, a resource manager, a media processing middleware, a learning memory library, and a visual image module, etc.

[0096] The learning memory store contains information about people in images. The visual image module determines the video theme for video generation based on the content characteristics of the image material. The media processing platform is responsible for executing the character verification process and searching for image materials that match the conversation content entered by the user.

[0097] Image materials can include static image materials and dynamic image materials. Static image materials can be images that remain unchanged within a preset time period. Such image materials are usually captured or created in an instant and do not exhibit dynamic characteristics that change over time. Static image materials can be, for example, photographs, paintings, charts, icons, etc. Dynamic image materials (such as videos or animations) contain information about changes in the time dimension and can present movement or changes through continuous image frames.

[0098] It should be noted that the image material involved in the embodiments of the present application can be photographed by an electronic device, downloaded by the electronic device from a server, or received by the electronic device from other electronic devices. The embodiments of the present application do not limit this.

[0099] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.

[0100] Content providers are used to store and retrieve data and make it accessible to applications. Data can include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0101] The resource manager can provide various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0102] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.

[0103] The system runtime layer includes the Android Runtime and system libraries. The Android Runtime includes the core library and virtual machine. The Android Runtime is responsible for scheduling and management of the Android system.

[0104] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0105] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0106] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0107] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0108] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0109] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0110] A 2D graphics engine is a drawing engine for 2D drawings.

[0111] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0112] It is understandable that Figure 2 The components included in the illustrated system framework layer, system library, and runtime layer do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or split certain components, or arrange the components differently.

[0113] The embodiments of the present application do not particularly limit the specific structure of the execution subject of a video generation method; as long as the code recording the video generation method of the embodiments of the present application can be executed to communicate according to the video generation method provided by the embodiments of the present application. For example, the execution subject of the video generation method provided by the embodiments of the present application can be a functional module in an electronic device that can call and execute a program, or a communication device used in an electronic device, such as a chip.

[0114] Currently, many electronic devices are equipped with cameras, allowing users to take photos and videos. Electronic devices also have communication functions, allowing users to download images and videos from the Internet or retrieve images and videos sent by other electronic devices. Users can edit images and / or videos stored on their electronic devices to create new videos from selected images and / or videos.

[0115] As an example scenario, a user inputs a user instruction to the electronic device. In response to the user instruction, the electronic device obtains corresponding materials from a gallery application according to the user instruction and generates a new video using the obtained materials.

[0116] However, in actual applications, the material that an electronic device retrieves from a gallery application based on user instructions may not be accurate. For example, a user may enter the user instruction "Help me generate a video of my daughter" into an electronic device. However, the material retrieved from the gallery application based on this user instruction may include other people besides my daughter. In this case, if the electronic device generates a video based on the retrieved material, the generated video will not meet the user's expectations.

[0117] Based on this, an embodiment of the present application provides a video generation method. In the embodiment of the present application, after recognizing a user instruction, the recognized person information can be confirmed, which is conducive to obtaining more accurate materials, so that the generated video better meets the user's expectations.

[0118] The following first describes an application scenario of the video generation method provided by an embodiment of the present application, which is introduced by taking the example of a voice assistant of an electronic device calling a video editing application. For the sake of convenience, the function of the electronic device generating a video according to the user's instructions in the embodiment of the present application is referred to as "smart filming".

[0119] See also Figure 3 , is a schematic diagram of the interface for generating a video through a voice assistant provided in an embodiment of the present application. Figures 4 to 6 This is a schematic diagram of the interface for character confirmation provided by the embodiment of this application. Figures 3 to 6 Introduce the application scenarios of video generation provided by the embodiments of the present application.

[0120] Reference Figure 3 As shown in (a) of the figure, the user speaks the voice "yoyo" within the voice detection range of the electronic device to wake up the voice assistant of the electronic device. After the electronic device detects the voice "yoyo", it responds with a voice "I am listening, please speak" and displays Figure 3 The interface shown in (a) in the Figure 3In the interface shown in (a), a voice assistant window 31 is displayed, which includes the text message "I'm listening, please speak." Of course, the purpose of this voice response and text message is to remind the user that the voice assistant has been awakened and to ask the user to speak the voice command. In actual applications, other voice response content and text message content can also be set according to the situation to serve the same reminder function.

[0121] The user responds to the voice "I'm listening, please speak" and / or the text message "I'm listening, please speak", and speaks the voice "Help me make a video of my child dancing from childhood to adulthood" within the voice monitoring range of the electronic device. After the electronic device detects the voice "Help me make a video of my child dancing from childhood to adulthood" (user instruction), the user's voice message "Help me make a video of my child dancing from childhood to adulthood" is displayed in the voice assistant window 31. Figure 3 As shown in (b) in FIG. , the voice assistant recognizes the user instruction to identify the first person information in the user instruction, and determines whether to enter the person confirmation process based on the recognized first person information.

[0122] If there is no need to enter the character confirmation process, the electronic device obtains the corresponding material from the gallery according to the first person information identified, and generates a video based on the obtained material; after the video is successfully generated, a voice message "A video of a child dancing from childhood to adulthood has been generated for you" is issued, and at the same time, the following is displayed: Figure 3 The interface shown in (c) is shown in Figure 3 As shown in (c) in the figure, the text "A video of a child dancing from childhood to adulthood has been generated for you" and a cover of the generated video are displayed in the voice assistant window. The cover displays a play control 32 and a video duration of "00:30". The video duration of "00:30" indicates that the generated video is 30 seconds long. In actual application, the user can click the play control 32 to trigger the electronic device to play the video. This application will not use diagrams as examples.

[0123] If you need to enter the character confirmation process, the electronic device will execute the character confirmation process, obtain the confirmation result, and display the confirmation result in the voice assistant window. Figures 8-10 Description in the Examples.

[0124] In an embodiment of the present application, the electronic device can cluster images in the image library to obtain portrait clustering information, and then perform a person identification process based on the learned portrait clustering information. In one example, the portrait clustering information may include the character number (tagid), character label (tagname), and character attributes (such as gender, age, etc.) of each character in the image.

[0125] The confirmation results may include the following situations:

[0126] Case 1: If the portrait clustering information in the gallery is not found, the confirmation result is an error message.

[0127] For example, after executing the character confirmation process, the electronic device is Figure 3 (b) in jump to Figure 4 The interface shown in the figure is as follows. Figure 4 As shown, an error message “No relevant photos found” is displayed in the voice assistant window 31.

[0128] Case 2: If the portrait cluster information in the gallery is searched, but the learning memory bank has not learned the person in the user instruction, the confirmation result is at least one group of person information that ranks high in the portrait cluster information in the gallery.

[0129] For example, after executing the character confirmation process, the electronic device is Figure 3 (b) in jump to Figure 5 The interface shown in (a) in the figure. Figure 5 As shown in (a), the voice assistant window 31 displays the top 6 groups of character information in the portrait cluster information of the gallery, including 6 character cards "Character 1", "Character 2", "Character 3", "Character 4", "Character 5", and "Character 6", as well as the character labels marked under each character card.

[0130] If the confirmed result is the person the user wants to select, the user can click the "OK" control 52 in the voice assistant window 31. In response to the user operation, the electronic device obtains the materials corresponding to the current confirmed results ("Person 1", "Person 2", "Person 3", "Person 4", "Person 5" and "Person 6") from the gallery, generates a video based on the obtained materials, and displays the video. Figure 3 The interface shown in (c) in the figure.

[0131] If the result is not the person the user wants to select, the user can click the "View All" button 51 in the voice assistant window 31. In response to the user operation, the electronic device jumps to the following screen: Figure 5 The interface shown in (b) in the figure. Figure 5 The interface shown in (b) includes a display frame 53, which displays character cards included in the portrait cluster information of the gallery. Each character card includes a selection control 54. The user can select a character card on the interface. For example, the user operates the selection control 54 corresponding to "character 12". In response to the user operation, the electronic device displays the following Figure 5 The interface shown in (c) is shown in Figure 5 As shown in (c) of FIG, the selection control 54 corresponding to "character 12" becomes selected. The user operates the confirmation control 55, and in response to the user operation, the electronic device jumps to the following Figure 5The interface shown in (d) is shown in Figure 5 As shown in (d), the character cards of "Character 4", "Character 1", "Character 2", "Character 3", "Character 4" and "Character 5" are displayed in the voice assistant window 31, and the selection control 56 of "Character 12" is selected by default.

[0132] After the user selects a character, the electronic device can generate a video based on the user's selection. Figure 5 As shown in (d) in FIG, when the user operates the "confirm" control 52 in the voice assistant window 31, in response to the user operation, the electronic device obtains the material corresponding to the current confirmation result ("person 12") from the gallery, generates a video based on the obtained material, and displays it as shown in FIG. Figure 3 The interface shown in (c) in the figure.

[0133] After the user selects a character, he can also add / modify the character's name. Figure 5 As shown in (d) in the figure, when the user operates the "Xiao Li" control 57 under the "Character 12" in the voice assistant window 31, in response to the user operation, the electronic device jumps to the following Figure 6 The interface shown in (e) is shown in Figure 6 As shown in (e) in FIG, the user enters the name of "person 12" in the input box 58 and clicks the "confirm" control 59. In response to the user operation, the electronic device jumps to the Figure 5 The interface shown in (f) is shown in the figure. Figure 5 As shown in (f), the name of "character 12" in the voice assistant window 31 is changed to "little daughter".

[0134] Case 3: If the portrait cluster information of the gallery is searched, and the first person information in the user instruction exists in the portrait cluster information of the gallery, the confirmation result is a portrait photo in the gallery that matches the first person information.

[0135] For example, after executing the character confirmation process, the electronic device is Figure 3 (b) in jump to Figure 6 The interface shown in (a) in the figure. Figure 6 As shown in (a) in the figure, the voice assistant window 31 displays portrait photos "Person 1", "Person 2" and "Person 3" in the gallery that match the first person information.

[0136] If the confirmed result is the person the user wants to select, the user can click the "OK" control 66 in the voice assistant window 31. In response to the user operation, the electronic device obtains the materials corresponding to the current confirmed results ("Person 1", "Person 2" and "Person 3") from the gallery, generates a video based on the obtained materials, and displays the video. Figure 3 The interface shown in (c) in the figure.

[0137] If the result of the confirmation is not the person the user wants to select, the user can click the "View All" button 61 in the voice assistant window 31. In response to the user operation, the electronic device jumps to the following screen: Figure 6 The interface shown in (b) in the figure. Figure 6 The interface shown in (b) includes a display box 62, in which portrait photos included in the portrait clustering information of the gallery are displayed, and each portrait photo includes a selection control 63. Among them, the selection control 63 of the portrait photo corresponding to the confirmation result ("Person 1", "Person 2" and "Person 3") obtained through the character confirmation process is selected by default. The user can select the portrait photo on this interface. For example, the user operates the selection control 63 corresponding to "Person 1", "Person 2", "Person 3" and "Person 4" respectively, and in response to the user operation, the electronic device displays the following Figure 6 The interface shown in (c) is shown in Figure 6 As shown in (c) of FIG, the selection controls 63 corresponding to "Character 1", "Character 2" and "Character 3" are respectively turned into unselected state, and the selection control 63 corresponding to "Character 4" is turned into selected state. The user operates the confirmation control 64, and in response to the user operation, the electronic device jumps to the following state: Figure 6 The interface shown in (d) is shown in Figure 6 As shown in (d) in the figure, the portrait photo of “Person 4” is displayed in the voice assistant window 31, and the portrait photos of “Person 1”, “Person 2” and “Person 3” are no longer displayed.

[0138] After the user selects a character, the electronic device can generate a video based on the user's selection. Figure 6 As shown in (d) in FIG, when the user operates the "confirm" control 66 in the voice assistant window 31, in response to the user operation, the electronic device obtains the material corresponding to the current confirmation result ("person 4") from the gallery, generates a video based on the obtained material, and displays it as shown in FIG. Figure 3 The interface shown in (c) in the figure.

[0139] After the user selects a character, he can also add / modify the character's name. Figure 6 As shown in (d) in the figure, when the user operates the "add name" control 65 under "character 4" in the voice assistant window 31, in response to the user operation, the electronic device jumps to the following Figure 6 The interface shown in (e) is shown in Figure 6 As shown in (e) in FIG, the user enters the name of "Person 4" in the input box 67 and clicks the "Confirm" control 66. In response to the user operation, the electronic device jumps to the Figure 6 The interface shown in (f) is shown in the figure. Figure 6 As shown in (f), the name of "Character 4" in the voice assistant window 31 is changed to "Little Daughter".

[0140] In the embodiment of the present application, the electronic device stores the name of the portrait photo added / modified by the user, such as in a learning memory bank. The next time the "smart photo" is created, there is no need to confirm the person added / modified by the user.

[0141] As described in the above example, the electronic device needs to determine whether to enter the person confirmation process. The following describes the judgment process of the person confirmation process.

[0142] See also Figure 7 , is a schematic diagram of the judgment interaction process of the character confirmation process provided in the embodiment of the present application. As an example and not a limitation, Figure 7 As shown, the interaction determination process may include the following steps:

[0143] S701, the voice assistant receives a user instruction.

[0144] The user instructions in the embodiment of the present application can be voice instructions or text instructions. For example, the user says "yoyo" to wake up the voice assistant of the electronic device, and after receiving the voice feedback "I am listening, please speak" from the voice assistant, the user says the voice instruction "help me make a video of my child dancing from childhood to adulthood". For another example, the user says "yoyo" to wake up the voice assistant of the electronic device, and after receiving the voice feedback "I am listening, please speak" from the voice assistant, the user says the voice instruction "help me make a video of my child dancing from childhood to adulthood". Figure 3 In the voice assistant window 31 shown in (a) of FIG, a text instruction "Help me make a video of my child dancing from childhood to adulthood" is input. For another example, the user presses and holds the power button (e.g., 1 second) to open the voice assistant application, and then enters the text instruction "Help me make a video of my child dancing from childhood to adulthood" in the voice assistant window 31.

[0145] It should be noted that user instructions can also be received with the help of other applications with voice functions, and this embodiment of the present application does not specifically limit this.

[0146] S702: The voice assistant identifies the first person information according to the user instruction.

[0147] In some implementations, a large language model can be deployed in the voice assistant. Accordingly, in S702, the voice assistant can recognize the user instruction and extract the first character information from the user instruction. The first character information may include a character tag and character attributes. For example, if the user instruction is "help me make a video of my child dancing from childhood to adulthood," the recognized first character information includes "child" (character tag) and minor (age).

[0148] In other implementations, the language model can be deployed separately or in other applications. Accordingly, in S702, the voice assistant can call the language model to recognize the user instruction.

[0149] S703: The voice assistant obtains the first character list from the learning memory library.

[0150] In some implementations, the learning memory library can classify and process the characters contained in the image library according to a preset period to obtain learning information; the learning information is stored in the learning memory library in the form of a first character list. In other implementations, such as the above Figure 5 and Figure 6 In the application scenario described in the embodiment, the character information (input information) added / modified by the user can also be stored in the learning memory library.

[0151] Among them, the learning information may include the character number (tagid), character label (tagname), character attributes (such as gender, age, etc.) and confidence of the characters contained in the image. When a group of learning information is learned from the learning memory library, the confidence of the group of learning information is the first value; when a group of learning information is stored after being added / modified by the user, the confidence of the group of learning information is the second value. The higher the confidence, the more credible the learned data. Since the character information added / modified by the user is more credible, the second value is set to be greater than the first value. For example, a confidence of 1 indicates complete credibility, a confidence less than 1 indicates incomplete credibility, and a confidence of 0 indicates complete untrustworthiness. The second value can be set to 1, and the first value can be set to a positive number less than 1.

[0152] S704: The voice assistant determines whether to enter the character confirmation process based on the identified first character information and the confidence level in the first character list obtained from the learning memory library.

[0153] If yes, execute S705; if no, execute S707.

[0154] In some embodiments, the determination method may include: if the first person list is empty, i.e., contains no learning information, then determining that the person confirmation process needs to be entered; if the first person list contains target information and the confidence level of the target information is less than a preset threshold, then determining that the person confirmation process needs to be entered; if the first person list contains target information and the confidence level of the target information is equal to or greater than a preset threshold, then determining that the person confirmation process does not need to be entered. The target information is learning information whose person tag matches the identified first person information.

[0155] For example, the first character list includes learning information 1 (character number 001, character label "son", confidence 1) and learning information 2 (character number 002, character label "daughter", confidence 0.5). If the first character information identified according to the user instruction is "son", and learning information 1 is found in the first character list as the target information, and the confidence of the target information is 1, it is determined that there is no need to enter the character confirmation process. If the first character information identified according to the user instruction is "daughter", and learning information 2 is found in the first character list as the target information, and the confidence of the target information is less than 1, it is determined that there is a need to enter the character confirmation process.

[0156] In some implementations, if the first person list contains the target information and the confidence level of the target information is less than a preset threshold, the person ID (tagid) of the target information can be stored in the second person list. The second person list can be stored in a learning memory library or in the voice assistant.

[0157] S705: If the character confirmation process needs to be entered, the voice assistant sends a character confirmation request to the media processing center.

[0158] In some examples, the character confirmation request sent by the voice assistant to the media processing center can use the following structure:

[0159]

[0160] Among them, in the person confirmation request, the value of the "command" field is a fixed value, such as "com.hihonor.intent.utilities.person.confirm". "sessionId" represents the session ID, which is used to distinguish which session the current person confirmation request belongs to. "person" represents the person information to be found. "gender" indicates that the gender of the person needs to be confirmed; for example, 0 represents male, 1 represents female, and 2 represents unknown. "ageGroup" indicates the age range of the person that needs to be confirmed; for example, "adult" represents adults, aged greater than or equal to 18. "minor" represents minors, aged less than 18. "unknow" means that the confirmed person's age is unpredictable. When the second person list is obtained, the person ID tagid of the person in the second person list will be placed in the "personTagIds" parameter.

[0161] S706: After receiving the character confirmation request, the media processing center executes the character confirmation process.

[0162] The implementation process of the character confirmation process can be found in the following Figures 8-10 Description in the Examples.

[0163] S707: After the media processing center executes the character confirmation process, it returns the confirmation result to the voice assistant.

[0164] In an example, the confirmation result returned by the media processing center to the voice assistant can use the following structure:

[0165]

[0166] The data stored in "yoyoData" is the information of the character card. "cardId" indicates the card ID of the character card. "cardUrl" indicates the URL address of the character card. "cardType" indicates the type of the character card. "packageName" indicates the package name of the character card. "sessionId" indicates the session ID of the character confirmation. "cardData" indicates the content that needs to be displayed on the character card. The content corresponding to this field is a JSON array. "tag_id" indicates the character ID of the person to be displayed. It is generated by the gallery, and each portrait has a unique tagid. "tag_name" indicates the character tag of the person to be displayed, which is consistent with the display in the gallery. "cover_data" indicates the image data of the clustered portrait. "selected" indicates that the character card is selected.

[0167] S708: The voice assistant executes the video generation process.

[0168] If the person confirmation process is executed, the voice assistant executes the video generation process according to the confirmation result of the person confirmation process. If the person confirmation process is not executed, the voice assistant executes the video generation process according to the recognized first person information.

[0169] In some implementations, the video generation process may include: the voice assistant sends a video generation instruction to the media processing center, which may include the first person information; the media processing center searches for relevant materials from the gallery based on the first person information, and sends the searched materials to the light editing service; the light editing service generates a video based on the materials and returns it to the voice assistant; the voice assistant displays the video cover in the voice assistant window.

[0170] In this embodiment of the present application, whether to proceed to the character verification process is determined based on the confidence level of the character information. A lower confidence level indicates a lower degree of credibility of the learned character information. In the case of a lower confidence level, the character verification process is initiated. This ensures that the acquired material is more consistent with the user's expectations, thereby improving the quality of the generated video.

[0171] The following describes the character confirmation process.

[0172] See also Figure 8 , is a schematic diagram of the character confirmation process provided in the embodiment of the present application. As an example and not a limitation, Figure 8 As shown, the character confirmation process may include the following steps:

[0173] S801: The media processing center receives a character confirmation request sent by the voice assistant and detects the portrait clustering information in the gallery.

[0174] In an embodiment of the present application, the gallery application can perform image clustering processing on the images in the gallery to obtain portrait clustering information. In one example, the portrait clustering information may include the character number (tagid), character label (tagname), and character attributes (such as gender, age, etc.) of each character in the image. The obtained portrait clustering information can be stored in the gallery application. Among them, the image clustering processing can adopt an existing image classification algorithm, which is not specifically limited in the embodiment of the present application.

[0175] S802: If the portrait clustering information of the gallery is not detected, the media processing center returns an error prompt message to the voice assistant (third result).

[0176] S803: If portrait clustering information of the gallery is detected, check whether the second person list is empty.

[0177] At step S804, if the second person list is not empty, the media processing center obtains a first result that matches the second person list from the image library and returns the obtained first result to the voice assistant for display. The first result includes a portrait photo, a person ID tagid, and a person tagname.

[0178] Depend on Figure 7 As can be seen from the embodiment, if the first character list contains the target information and the confidence level of the target information is less than a preset threshold, the character number (tagid) of the target information can be stored in the second character list. In other words, the second character list contains the first character information that exists in the learning memory library (i.e., exists in the first character list). If the second character list is empty, it means that the learning memory library does not have the memory of the first character information in the user instruction. If the second character list is not empty, it means that the learning memory library does have the memory of the first character information in the user instruction.

[0179] In one implementation, see Figure 9 , is a schematic diagram of the process of obtaining the first result provided in the embodiment of the present application. As an example and not a limitation, Figure 9 As shown, obtaining the first result in S804 may include:

[0180] S901: Obtain untraversed second person information from the second person list.

[0181] In the embodiment of the present application, the character information in the second character list is recorded as the second character information.

[0182] S902: Determine whether the character number of the second character information exists in the portrait cluster information.

[0183] If it exists, execute S903; if it does not exist, execute S904.

[0184] S903: If the second character information exists, store the character number of the second character information into the third character list.

[0185] Since the second character information exists in the learning memory library, that is, exists in the first character list, the character number corresponding to the second character information can be obtained from the first character list.

[0186] S904: Determine whether there is any untraversed second character information in the second character list.

[0187] If there is untraversed second character information in the second character list, execute S901.

[0188] If the untraversed second character information does not exist in the second character list, execute S905.

[0189] S905: Determine whether the third person list is empty.

[0190] S906: If the third person list is empty, obtain at least one set of third person information ranked high based on the portrait clustering information in the gallery, and use the obtained at least one set of third person information as the first result. The third person information includes a portrait photo, a person ID, and a person tag.

[0191] The portrait clustering information includes the number of photos corresponding to each person, such as Figure 5 As shown in (b), "person 1" is the eldest daughter, and the number of photos is 90. The persons in the portrait cluster information are sorted according to the number of photos corresponding to each person in the portrait cluster information, and at least one set of third person information that ranks high is used as the first result.

[0192] S907, if the third person list is not empty, obtain fourth person information that matches the person number in the third person list from the gallery, and use the obtained fourth person information as the first result, wherein the fourth person information includes a portrait photo, a person number, and a person label.

[0193] In an embodiment of the present application, when the learning memory library contains memory of the character information in the user instruction, images matching the character information in the user instruction can be obtained from the gallery, thereby achieving accurate acquisition of materials and making the generated video more in line with user expectations.

[0194] In an embodiment of the present application, after the image library performs image clustering processing, each character corresponds to a character number and a category number. When the image clustering processing is normal, the character number and category number of each character are one-to-one corresponding. For example, Xiao Wang's character number is 001, and the category number is 1. The character number here is similar to the character's ID number, and the category number is similar to the character's name, and each ID number corresponds to a name. However, in some application scenarios, abnormal image clustering processing may occur, and the same character may be clustered into two different characters. In this case, two pieces of character information will appear. For example, the character number of Xiao Wang who appears for the first time in the image is 001, and the category number is 1; the character number of Xiao Wang who appears for the second time in the image is 002, and the category number is 2. As mentioned above Figure 5 and Figure 6 As described in the embodiment, if a user modifies or adds a person's label, the person's category number will be modified accordingly. This may result in the same category number corresponding to multiple different person numbers. If a second person's information is searched for in the portrait cluster information based on the person number, this may lead to inaccurate detection results.

[0195] To solve the above problem, in one implementation, S902 may include:

[0196] Convert the character number of the second character information into a category number; determine whether the category number of the second character information exists in the portrait cluster information; if so, determine that the character number of the second character information exists in the portrait cluster information; if not, determine that the character number of the second character information does not exist in the portrait cluster information.

[0197] S805: If the second person list is empty, the media processing center filters the portrait cluster information in the image library based on the first person information to obtain a second result, and returns the second result to the voice assistant for display. The second result includes the portrait photo, its person number, and person label.

[0198] In one implementation, see Figure 10 , is a schematic diagram of the process of obtaining the second result provided in the embodiment of the present application. As an example and not a limitation, Figure 10 As shown, the process of obtaining the second result in S805 may include:

[0199] S1001, obtaining a first character list from a learning memory library.

[0200] S1002: Determine whether the first character list is empty.

[0201] S1003: If the first person list is empty, obtain at least one set of third person information ranked high according to the portrait clustering information of the gallery, and use the obtained third person information as the second result.

[0202] S1004: If the first person list is not empty, determine whether there is untraversed fifth person information in the portrait cluster information.

[0203] In the embodiment of the present application, the person information in the portrait cluster information is recorded as the fifth person information.

[0204] If it exists, execute S1005-S1008; if it does not exist, execute S1009-S1011.

[0205] S1005: If there is untraversed fifth person information in the portrait cluster information, obtain untraversed fifth person information from the portrait cluster information.

[0206] S1006: Determine whether the current fifth person information exists in the first person list.

[0207] If the current fifth person information does not exist in the first person list, continue to execute S1004.

[0208] S1007: If the current fifth person information exists in the first person list, determine whether the person attributes of the fifth person information match the person attributes of the first person information.

[0209] S1008 , if the character attributes of the fifth character information match the character attributes of the first character information, store the character number of the fifth character information into the fourth character list, and continue to execute S1004 .

[0210] If the character attributes of the fifth character information do not match the character attributes of the first character information, the process continues with S1004 .

[0211] S1009: If the untraversed fifth person information does not exist in the portrait cluster information, determine whether the fourth person list is empty.

[0212] If the fourth character list is empty, execute S1003.

[0213] S1011, if the fourth person list is not empty, obtain the sixth person information that matches the person number in the fourth person list from the gallery, and use the obtained sixth person information as the second result, wherein the sixth person information includes a portrait photo, its person number and person label.

[0214] For example, if the user command is "Generate a video of my son traveling in Hong Kong last year", if the learning memory bank has not learned the character "son", that is, the second character list is empty, but the character attribute in the first character information is recognized as male, then the top 6 male photos can be obtained from the gallery and sent to the voice assistant for display.

[0215] In an embodiment of the present application, when there is no memory of the character information in the user instruction in the learning memory library, images that meet the character attributes in the user instruction can be filtered out from the gallery, which is conducive to obtaining more accurate materials, so that the generated video is more in line with user expectations.

[0216] The following describes the process of determining character materials. Figure 11 , is a flow chart of character material confirmation provided in the embodiment of the present application. As an example and not a limitation, Figure 11 As shown, during the character material confirmation process, the interaction process between various modules may include:

[0217] S1101: The voice assistant sends a character confirmation request to the media processing center.

[0218] In some implementations, the voice assistant sends the character confirmation request to the AigcManager of the media processing center through the atomic service framework (appclip) (AigcManager registers the service to the atomic service framework through static registration).

[0219] S1102, the media processing center obtains the first character list from the learning memory library.

[0220] S1103: The media processing center obtains portrait clustering information from the image library.

[0221] S1104: The media processing center performs a character confirmation process based on the first character list and the portrait clustering information, and returns the confirmation result to the voice assistant.

[0222] The implementation of step S1104 can be found in Figures 8-10 Description in the Examples.

[0223] S1105: The voice assistant downloads the corresponding character confirmation card from the quick application engine according to the confirmation result and renders the card.

[0224] S1106, when the user clicks "View All", in response to the user operation, jump to the gallery more people confirmation page.

[0225] S1107, after the user selects a character, the gallery application returns the character label and character image data of the user-selected character to the voice assistant.

[0226] S1108: The voice assistant updates the character image based on the received data.

[0227] The application scenarios of steps S1106-S1108 can be found in Figure 5 (a)-(b) and Figure 6 (a)-(b) in .

[0228] S1109, when the user clicks "Add", in response to the user operation, jump to the gallery character label change page.

[0229] S1110, when the user modifies the label, the gallery updates the character label and sends the updated label to the voice assistant and the learning memory library to keep the information in sync with the learning memory library.

[0230] The application scenario of the user adding a character tag in steps S1109-S1110 can be found in Figure 5 (d)-(f) and Figure 6 (d)-(f) in the figure.

[0231] S1111, when the user clicks "Confirm", the voice assistant confirms the character tagid and character label of the currently selected character.

[0232] For application scenarios where the user clicks "Confirm", see Figure 5 (d) and Figure 3 (c) in the.

[0233] In an embodiment of the present application, after a user inputs a user command to a voice assistant, the electronic device can confirm the character based on the character information in the user command, which helps to obtain more accurate material, so that the generated video better meets the user's expectations. In addition, the user can also select a character through user operation and adjust the identified character to make the generated video more in line with the user's expectations. The user can also add / modify character tags, and the modified character tags are synchronized to the learning memory library. The next time a video is generated, there is no need to confirm the character tag.

[0234] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0235] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is run on an electronic device, it can implement the steps in the above-mentioned various method embodiments.

[0236] The embodiments of the present application further provide a computer program product. When the computer program product is run on an electronic device or a wireless router, the electronic device can implement the steps in the above-mentioned various method embodiments.

[0237] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include at least: any entity or device capable of carrying the computer program code to the first device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0238] The present application also provides a chip comprising a processor coupled to a memory, wherein the processor invokes a computer program stored in the memory to implement the steps of any method embodiment of the present application. The chip may be a single chip or a chip module composed of multiple chips.

[0239] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0240] Those skilled in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0241] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A video generation method, characterized in that: include: When a user instruction is detected, identifying first person information contained in the user instruction; Obtaining a confidence level corresponding to the first person information, where the confidence level is used to indicate the credibility of the identified person; If it is determined to execute the character confirmation process according to the confidence level corresponding to the first character information, the character corresponding to the first character information is confirmed to obtain a confirmation result; A video is generated according to the confirmation result.

2. The method according to claim 1, characterized in that The first character information includes a character tag and character attributes of each character; The obtaining of the confidence level corresponding to the first character information includes: Obtaining a first character list from an electronic device, the first character list including learning information, the learning information being information obtained by the electronic device through classification processing of images in a preset gallery, each set of learning information including a character number, a character label, a character attribute, and a confidence level of a character; The confidence level corresponding to the first character information is obtained according to the first character list.

3. The method according to claim 2, characterized in that The obtaining, according to the first character list, a confidence level corresponding to the first character information, includes: Searching for target information in the first character list, where the target information is learning information that matches the character tag of the first character information; If the target information is found in the first character list, the confidence level corresponding to the first character information is set as the confidence level of the target information; If the target information is not found in the first character list, the confidence level corresponding to the first character information is set to null.

4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: If the confidence level corresponding to the first character information is empty or less than a preset threshold, determining to execute the character confirmation process; If the confidence level corresponding to the person information is greater than a preset threshold, it is determined that the person confirmation process will not be executed.

5. The method according to any one of claims 2 to 4, characterized in that The confirming the person corresponding to the first person information to obtain a confirmation result includes: Obtaining portrait clustering information from a preset gallery, wherein the portrait clustering information is information obtained by performing clustering processing on images in the preset gallery by a gallery application, and the portrait clustering information includes a character number, a character label, and a character attribute of each clustered character; detecting whether a second character list is empty, the second character list including first character information existing in the first character list; If the second character list is not empty, obtaining a first result matching the second character list from the preset gallery, and the confirmation result is the first result; If the second character list is empty, the portrait clustering information is filtered according to the first character information to obtain a second result, and the confirmation result is the second result.

6. The method according to claim 5, characterized in that The obtaining, from the preset gallery, a first result matching the second character list, includes: For the second person information, determining whether the person number of the second person information exists in the portrait cluster information, wherein the second person information is person information in the second person list; If the character number of the second character information exists in the portrait cluster information, storing the character number corresponding to the second character information into a third character list; If there is no untraversed second character information in the current second character list, determining whether the third character list is empty; If the third person list is empty, obtaining at least one set of third person information according to the portrait clustering information, and using the at least one set of third person information as the first result; If the third character list is not empty, fourth character information matching the character number in the third character list is obtained from the preset gallery, and the fourth character information is used as the first result.

7. The method according to claim 6, characterized in that The portrait clustering information also includes the category number of each clustered character, where one category number corresponds to at least one character number; The determining whether the character number of the second character information exists in the portrait cluster information includes: Converting the character number of the second character information into a category number; Determining whether the category number of the second person information exists in the portrait cluster information; If so, determining whether the character number of the second character information exists in the portrait cluster information; If not, it is determined that the character number of the second character information does not exist in the portrait cluster information.

8. The method according to claim 5, characterized in that The filtering of the portrait cluster information according to the first character information to obtain a second result includes: Determine whether the first character list is empty; If the first character list is empty, at least one set of third character information is obtained according to the portrait clustering information, and the at least one set of third character information is used as the second result.

9. The method according to claim 8, characterized in that After determining whether the first character list is empty, the method further includes: If the first character list is not empty, obtaining fifth character information, where the fifth character information is character information that has not been traversed in the portrait cluster information; Determine whether the fifth person information exists in the first person list; If the fifth character information exists in the first character list, determining whether the character attributes of the fifth character information match the character attributes of the first character information; If the character attributes of the fifth character information match the character attributes of the first character information, the fifth character information is stored in the fourth character list, and the next fifth character information is obtained until all the portrait cluster information has been traversed; When all the portrait cluster information has been traversed, determining whether the fourth person list is empty; If the fourth person list is empty, obtaining at least one set of third person information according to the portrait clustering information, and using the at least one set of third person information as the second result; If the fourth character list is not empty, obtaining sixth character information that matches the character number in the fourth character list from the preset gallery, and using the sixth character information as the second result.

10. The method according to claim 5, characterized in that The method further comprises: If the portrait clustering information is not obtained from the preset gallery, a third result is sent, and the third result is an error prompt message.

11. The method according to claim 1, characterized in that The method further comprises: If it is determined not to execute the person confirmation process based on the confidence level corresponding to the first person information, a video is generated based on the first person information.

12. An electronic device, characterized in that: The electronic device includes one or more processors and a memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the electronic device to execute the method according to any one of claims 1 to 11.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes instructions, and when the instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 11.

14. A chip system, characterized in that: The chip system is applied to an electronic device, and the chip system includes one or more processors, and the one or more processors are used to call computer instructions so that the electronic device executes the method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Voice interaction method and apparatus, and electronic equipment

    CN111292743A

  • Video generation method and related device

    CN111669515A

  • Photo album video generation method, electronic equipment and storage medium

    CN112035685A

  • Method and device for training information fusion model and generating collective video

    CN112182297A

  • Video generation method and device, equipment and medium

    CN112541353A