Media resource editing method, electronic equipment and storage medium

Adjusting the number of media materials through voice assistant functions and dialogue inputs solves the problem of inflexible media materials adjustment in the existing technology, and improves the convenience and user experience of media resource editing.

CN119946367APending Publication Date: 2025-05-06HONOR DEVICE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311867611.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-27
Filing Date
2023-12-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing electronic devices cannot flexibly adjust media materials when editing media resources, resulting in poor user experience.

Method used

Through the voice assistant function, users can generate media resources through dialogue input instructions, and flexibly adjust media materials by adjusting the number of media materials in dialogue input.

Benefits of technology

It realizes flexible adjustment of the number of media materials, improving the convenience and user experience of media resource editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946367A_ABST
    Figure CN119946367A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of terminals, in particular to a media resource editing method, electronic equipment and a storage medium. The method can be applied to electronic equipment such as mobile phones and tablet personal computers. In such a method, after a voice assistant is awakened, in response to a first dialog input, the electronic device displays a dialog box and displays a first answer in the dialog box. Thereafter, the electronic device displays a second answer in the dialog box in response to the second dialog input. Next, in response to the third dialog input, the electronic device displays a third answer in the dialog box. In this method, the electronic device may replace the media material, e.g., the media material in a second answer, in response to a dialog indicating an increase in the media material, in a case where the media material is more, such as in a case where the media material is greater than a first number threshold. Therefore, the electronic equipment can flexibly adjust the media material.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office on October 27, 2023, with application number 202311418278.6 and invention name “A video production method and electronic device based on large models”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of terminal technology, and in particular to a media resource editing method, an electronic device, and a storage medium. Background Art

[0003] In order to meet the needs of users to record and share their lives anytime and anywhere, most electronic devices such as mobile phones and tablets are equipped with cameras. In order to further enhance the user's shooting and creation experience, electronic devices can generate media resources based on the photos and videos taken by users, and support users to edit the media resources on electronic devices.

[0004] Currently, electronic devices can generate media resources based on media materials. During the process of editing the media resources by the electronic devices, the electronic devices have a relatively fixed way of adjusting the media materials and are unable to flexibly adjust the media materials. Summary of the invention

[0005] The embodiments of the present application provide a media resource editing method, an electronic device, and a storage medium, which can flexibly adjust media materials during the process of editing media resources.

[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0007] In the first aspect, an embodiment of the present application provides a method for editing media resources, which can be applied to electronic devices such as mobile phones and tablet computers, and the electronic devices include voice assistants. The above method includes: when the voice assistant is awakened, the electronic device responds to the first dialogue input and displays a dialog box. The above first dialogue input indicates the generation of media resources, and the dialog box includes a first answer, and the first answer includes a first number of media materials, and the first number is less than the first number threshold. Next, in response to the second dialogue input, the electronic device displays the second answer in the dialog box, and the second answer includes a second number of media materials. The second number is greater than or equal to the first number threshold, and the second dialogue input indicates the addition of media materials. Afterwards, the electronic device responds to the third dialogue input and displays the third answer in the dialog box, and the third answer includes the second number of media materials, and the third dialogue input indicates the addition of media materials. The second number of media materials in the third answer is different from at least one of the second number of media materials in the second answer.

[0008] Among them, the above-mentioned electronic device including a voice assistant can be understood as the electronic device is installed with a voice assistant application, or the electronic device has a voice assistant function.

[0009] In this method, when there are more media materials, such as when the media materials are greater than a first quantity threshold, in response to a dialog indicating to increase the media materials, the electronic device can replace the media materials, such as replacing the media materials in the second answer. Thus, the electronic device can flexibly adjust the media materials, especially the quantity of the media materials.

[0010] In a possible design of the first aspect, after the third answer is displayed in the dialog box, the method further includes: in response to a fourth dialogue input, the electronic device displays a fourth answer in the dialog box, the fourth answer includes a third amount of media materials, the third amount is less than or equal to a second amount threshold, the second amount threshold is less than the first amount threshold, and the fourth dialogue input indicates a reduction in media materials. Afterwards, in response to the fourth dialogue input, the electronic device displays a fifth answer in the dialog box, the fifth answer includes a fourth amount of media materials, the fourth amount is less than the third amount, and the fourth dialogue input indicates a reduction in media materials. The difference between the third amount and the fourth amount is less than the difference between the second amount and the third amount.

[0011] It is understandable that if the amount of media material is too small, the quality of the subsequently generated media resources will be reduced. Therefore, when the amount of media material is relatively small, such as less than the second amount threshold, the electronic device reduces the amount of media material. In this way, the quality of the subsequently generated media resources can be improved.

[0012] In another possible design of the first aspect, after the third answer is displayed in the dialog box, the method further includes: after the media material is reduced to a fifth quantity, in response to a fifth dialogue input, the electronic device prompts in the dialog box that the media material cannot be further reduced, and the fifth quantity is less than or equal to the third quantity threshold.

[0013] In this design, if the number of media materials is relatively small, it will affect the quality of the subsequently generated media resources. Therefore, after the number of media materials is reduced to the fifth number, the electronic device can prompt in a dialog box that the number of media materials cannot be reduced again. Then, the electronic device can improve the quality of the subsequently generated media resources and improve the user experience.

[0014] In another possible design of the first aspect, when the second dialogue input indicates an increase in a specified amount of media material, the difference between the second quantity and the first quantity is the specified quantity; when the second dialogue input does not indicate an increase in a specified amount of media material, the difference between the second quantity and the first quantity is positively correlated with the first quantity.

[0015] In this design, if the dialogue input does not indicate the addition of a specified number of media materials, that is, the dialogue input is a voice instruction with ambiguous semantics; then, the electronic device can determine the number of media materials to be added based on the dialogue input according to the current number of media materials when receiving the second dialogue input, such as the first number. In other words, the difference between the second number and the first number is positively correlated with the first number. Therefore, in the case where the dialogue input is a voice instruction with ambiguous semantics, the electronic device can also flexibly adjust the media materials, which can improve the user experience.

[0016] In another possible design of the first aspect, the second dialogue input includes quantity degree indication information, and the higher the degree indicated by the quantity degree indication information, the greater the difference between the second quantity and the first quantity.

[0017] In this design, the electronic device can determine the amount of media material added based on the second dialogue input according to the quantity degree indication information included in the second dialogue input. That is, the difference between the second amount and the first amount is positively correlated with the degree indicated by the quantity degree indication information included in the second dialogue input. Therefore, in the case where the dialogue input includes the quantity degree indication information, the electronic device can also flexibly adjust the media material.

[0018] In another possible design of the first aspect, when the fourth dialogue input does not indicate reducing the specified amount of media materials, the difference between the second amount and the third amount is positively correlated with the amount of media materials when the fourth dialogue input is received.

[0019] In this design, if the dialogue input does not indicate reducing the specified amount of media materials, that is, the dialogue input is a voice instruction with ambiguous semantics; then, the electronic device can determine the amount of media materials to be reduced based on the dialogue input according to the current amount of media materials when receiving the fourth dialogue input, such as the second amount. In other words, the difference between the second amount and the third amount is positively correlated with the amount of media materials when the fourth dialogue input is received. Therefore, in the case where the dialogue input is a voice instruction with ambiguous semantics, the electronic device can also flexibly adjust the media materials, which can improve the user experience.

[0020] In another possible design of the first aspect, the dialog box further includes a first control, and after the third answer is displayed in the dialog box, the method further includes: in response to a triggering operation on the first control, the electronic device displays a first interface, the first interface including media materials in a selected state and media materials in a non-selected state. The media materials in the selected state are media materials included in the third answer, and the media materials in the non-selected state are media materials included in the third answer and not included in the second answer.

[0021] In this design, the electronic device displays the first interface so that the user can visually observe in the first interface that the electronic device replaces the media material in response to the third dialogue input.

[0022] In another possible design of the first aspect, after the third answer is displayed in the dialog box, the method further includes: in response to a sixth dialogue input, the electronic device displays a sixth answer in the dialog box, the sixth answer includes a first media resource, the first media resource includes the media material included in the third answer, and the sixth dialogue input indicates that the media resource is generated with the media material included in the third answer.

[0023] In this design, after displaying the sixth answer, the electronic device may also generate a media resource in response to the dialogue input.

[0024] In another possible design of the first aspect, after the sixth answer is displayed in the dialog box, the method further includes: in response to the seventh dialogue input, the electronic device displays the seventh answer in the dialog box, and the seventh answer includes a media editing control. When the seventh dialogue input indicates the specified editing content, the electronic device specifies the menu option corresponding to the editing content in the media editing control in an expanded state. When the seventh dialogue input does not indicate the specified editing content, the menu option of the electronic device in the media editing control is not expanded. The specified editing content includes: one or more of changing the template, changing the background music, and changing the duration; changing the template corresponds to the first menu option, changing the background music corresponds to the second menu option, and changing the duration corresponds to the third menu option.

[0025] In this design, the electronic device can edit the media resource by displaying the media editing control, so that the electronic device can flexibly edit the media resource.

[0026] In another possible design of the first aspect, the method further includes: in response to a seventh dialogue input, displaying a guide bubble in the dialog box, where the guide bubble is used to prompt the user to issue a voice command.

[0027] In this design, the electronic device can prompt the user to issue voice instructions through a guide bubble, prompting and guiding the user to edit the media resources.

[0028] In another possible design of the first aspect, after the third answer is displayed in the dialog box, the method further includes: in response to the eighth dialogue input, the electronic device displays the eighth answer in the dialog box, the eighth answer includes a second media resource, and the second media resource includes the same media material as the first media resource. And, in the case where the eighth dialogue input indicates a specified duration, the difference in duration between the second media resource and the first media resource is within the specified duration. In the case where the eighth object input does not indicate a specified duration, the difference in duration between the second media resource and the first media resource is within a first range.

[0029] In this design, when adjusting the duration of a media resource through a dialogue input, the electronic device can adjust the duration of the media resource to different degrees based on whether the dialogue input is clear semantics or fuzzy semantics. That is, when the eighth dialogue input indicates a specified duration, the difference in duration between the second media resource and the first media resource is the specified duration. When the eighth object input does not indicate a specified duration, the difference in duration between the second media resource and the first media resource is within a first range. Thus, the electronic device can also flexibly adjust the duration of the media resource, which can improve the user experience.

[0030] According to a second aspect, an electronic device is provided, comprising a memory and one or more processors, wherein the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code comprises computer instructions; when the computer instructions are executed by the processor, the electronic device executes the method provided by the first aspect and any possible design of the first aspect.

[0031] According to a third aspect, a computer-readable storage medium is provided, comprising computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method provided by the first aspect and any possible design of the first aspect.

[0032] According to a fourth aspect, a computer program product comprising instructions is provided. When the computer program product is run on an electronic device, the electronic device can execute the method provided by the first aspect and any possible design of the first aspect.

[0033] Among them, the technical effects brought about by any design method in the second to fourth aspects can refer to the technical effects brought about by different design methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A schematic diagram of a usage scenario provided for an embodiment of the present application;

[0035] Figure 2 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0036] Figure 3 A schematic diagram of a software architecture of an electronic device provided in an embodiment of the present application;

[0037] Figure 4 A flowchart of a media resource editing method provided in an embodiment of the present application;

[0038] Figure 5 A set of interface schematic diagrams provided for embodiments of the present application;

[0039] Figure 6 Another set of interface schematic diagrams provided for embodiments of the present application;

[0040] Figure 7 Another set of interface schematic diagrams provided for embodiments of the present application;

[0041] Figure 8 Another set of interface schematic diagrams provided for embodiments of the present application;

[0042] Fig. 9 Another set of interface schematic diagrams provided for embodiments of the present application;

[0043] Fig.10 Another set of interface schematic diagrams provided for embodiments of the present application;

[0044] Fig.11 Another set of interface schematic diagrams provided for embodiments of the present application;

[0045] Fig.12 Another set of interface schematic diagrams provided for embodiments of the present application;

[0046] Fig.13 Another set of interface schematic diagrams provided for embodiments of the present application;

[0047] Fig.14 Another set of interface schematic diagrams provided for embodiments of the present application;

[0048] Fig.15 A schematic diagram of the hardware structure of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Among them, in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in the present application is only a kind of association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. And, in the description of the embodiments of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or its similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference.

[0050] Meanwhile, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.

[0051] In the technical solutions disclosed in this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the relevant laws and regulations and do not violate public order and good morals.

[0052] With the development of terminal technology, users are using electronic devices more and more frequently. In order to meet the needs of users to record and share their lives anytime and anywhere, most electronic devices such as mobile phones and tablets are equipped with cameras, and electronic devices can take photos, videos, etc. based on the camera. In order to further enhance the user's shooting and creation experience, electronic devices can generate media resources based on photos, videos, etc. taken by users, and electronic devices can edit media resources. For example, electronic devices can change the background music used by media resources, adjust the duration of media resources, change the template used by media resources, adjust the number of media materials that make up media resources, and so on.

[0053] Currently, users need to perform complicated operations to edit media resources on electronic devices, which makes it inconvenient and inconvenient for users to edit media resources.

[0054] In view of this, an embodiment of the present application provides a method for editing media resources, in which an electronic device can edit media resources based on a user's voice command. In this way, the convenience of electronic devices for editing media resources can be improved, and users can edit media resources conveniently and quickly, which can improve the user experience.

[0055] It should be noted that the media resource may include a video or an image. The media resources (such as videos or images) edited by the electronic device are not limited in the embodiments of the present application. In the subsequent embodiments, the type of media resource edited is a video as an example for schematic illustration. The media material can be understood as the photos or videos that constitute the media resources.

[0056] For example, if photo A, photo B, and video C are edited to get video D, then photo A, photo B, and video C are all media materials, and video D is a media resource. For another example, if photo E is edited to get video F, then photo E is a media material, and video F is a media resource.

[0057] For example, see Figure 1 The technical solution provided in the embodiment of the present application can be applied to the process in which a user uses the electronic device 100 to edit media resources. It is particularly applicable to the process in which a user generates media resources through the voice assistant function of the electronic device and edits the media resources. The voice assistant function of the electronic device can be implemented through technologies such as voice recognition. For the voice assistant function on the electronic device, see the following introduction, which will not be repeated here.

[0058] The electronic device 100 may also be referred to as a terminal, terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), etc. The electronic device 100 may be a mobile phone, tablet computer, wearable device, smart screen, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., which have a display screen; it may also be a vehicle-mounted device with a display screen such as a vehicle-mounted computer, a vehicle-mounted computer, etc.; it may also be some smart watches, smart bracelets, etc., which are IoT devices with a display screen. The embodiments of the present application do not impose any restrictions on the product form of the electronic device.

[0059] Next, the hardware structure and software architecture of the electronic device provided in the embodiments of the present application are introduced.

[0060] Figure 2 The hardware structure diagram of the electronic device 100 is shown, and the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, a display screen 194, an audio module 170, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a touch sensor 180K, etc.; and the audio module 170 may include a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, etc.

[0061] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0062] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0063] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0064] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0065] The MIPI interface can be used to connect the processor 110 with peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate via the DSI interface to implement the display function of the electronic device 100.

[0066] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present application is only a schematic illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0067] The USB interface 130 is an interface that complies with the USB standard specification, and specifically can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transfer data between the electronic device 100 and a peripheral device. It can also be used to connect headphones to play audio through the headphones. The interface can also be used to connect other electronic devices, such as AR devices, etc.

[0068] The electronic device 100 implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.

[0069] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0070] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A.

[0071] In some embodiments, touch operations acting on the same touch position but with different touch operation strengths may correspond to different operation instructions. For example, when a touch operation with a touch operation strength less than a first pressure threshold acts on a short message application icon, an instruction to view a short message is executed. When a touch operation with a touch operation strength greater than or equal to the first pressure threshold acts on a short message application icon, an instruction to create a new short message is executed.

[0072] The touch sensor 180K is also called a "touch panel". The touch sensor 180K can be set on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be set on the surface of the electronic device 100, which is different from the position of the display screen 194.

[0073] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function, such as storing music, video and other files in the external memory card.

[0074] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0075] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.

[0076] The speaker 170A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.

[0077] The receiver 170B, also called a "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or voice message, the voice can be received by placing the receiver 170B close to the human ear.

[0078] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to microphone 170C to input the sound signal into microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, realize directional recording function, etc.

[0079] The earphone interface 170D is used to connect a wired earphone and can be a USB interface 130 or a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0080] For example, the electronic device 100 may acquire a wake-up word through the microphone 170C, and in response to acquiring the wake-up word, the electronic device 100 activates the voice assistant function.

[0081] Among them, the voice assistant function of the electronic device can be understood as the electronic device 100 collecting the user's voice through a microphone, analyzing and identifying the user's voice, and executing the instructions corresponding to the user's voice, so that the user can control the terminal device 100 through voice.

[0082] In different devices, the voice control function may have different names, such as "voice control", "intelligent voice", "voice assistant", "see and speak", "voice command", "free command", "intelligent AI", etc. The specific implementation of voice control functions with different names may also be different.

[0083] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with other devices through wireless communication technology and network.

[0084] For example, the electronic device 100 may obtain media resources generated by other devices through the antenna 1 and the mobile communication module 150 , or through the antenna 2 and the wireless communication module 160 , and edit the media resources on the electronic device 100 .

[0085] Next, the software architecture of the electronic device 100 is introduced. The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 100.

[0086] Figure 3 1 is a schematic diagram of the software structure of the electronic device 100 of the embodiment of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely, the application layer, the application framework layer, the Android runtime (Android runtime) and the system library, and the kernel layer.

[0087] The application layer can include a series of application packages.

[0088] like Figure 3 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, Bluetooth, music, video, short message, etc. Among them, the gallery is used to provide functions such as viewing photos and videos, and searching for photos and videos.

[0089] The above-mentioned application layer also includes a voice assistant application, which is used to provide a voice assistant function for the electronic device 100.

[0090] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0091] like Figure 3 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0092] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0093] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0094] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images.

[0095] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including connecting, hanging up, etc.).

[0096] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0097] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message is displayed in the status bar, a prompt sound is emitted, an electronic device vibrates, an indicator light flashes, etc.

[0098] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for scheduling and management of the Android system.

[0099] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.

[0100] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.

[0101] The system library may include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.

[0102] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.

[0103] The media library supports playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0104] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0105] A 2D graphics engine is a drawing engine for 2D drawings.

[0106] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.

[0107] The electronic device is a mobile phone, which has the above Figure 2 The hardware structure and Figure 3 Taking the software architecture shown as an example, the interface display method provided in the embodiment of the present application is introduced.

[0108] For example, see Figure 4 The media resource editing method provided in the embodiment of the present application may include steps S400-S401.

[0109] S400. Mobile phone obtains media resources.

[0110] In some embodiments, the mobile phone can filter the photos and / or videos to obtain media materials. Afterwards, the mobile phone edits the media materials to obtain media resources.

[0111] As a possible implementation, the mobile phone may filter photos and / or videos according to target conditions to obtain media materials.

[0112] Among them, the above-mentioned target conditions may include one or more conditions among shooting time conditions, shooting location conditions and shooting content conditions. The shooting time condition includes that the shooting time of the photo and / or video is close to the target time. The shooting location condition includes that the shooting location of the photo and / or video is close to the target location. The shooting content condition includes that the shooting content of the photo and / or video includes target content, and the target content includes target people, target animals, target actions, target actions of target people, etc. Among them, the shooting content of the photo can be understood as the people, animals, actions of people, actions of animals, etc. included in the photo; similarly, the shooting content of the video can be understood as the people, animals, actions of people, actions of animals, etc. included in the video.

[0113] Specifically, the mobile phone can perform image recognition on photos to obtain the content of the photos, and the mobile phone can perform image recognition on videos to obtain the content of the videos. For the process of image recognition of photos / videos by mobile phones, please refer to the relevant technology, and this application will not go into details here. In addition, the above-mentioned photos and / or videos can be stored locally on the mobile phone or stored in other devices (such as the cloud), and can be set according to actual usage requirements.

[0114] It can be understood that the above-mentioned target time, target location, target content, etc. can all be specified by the user or generated by the mobile phone.

[0115] In one example, the mobile phone may analyze shooting locations of photos and / or videos stored locally in the mobile phone, and select shooting locations with a relatively large number, such as more than 10, as target locations.

[0116] It should be understood that the mobile phone can also have other ways to generate the target location, target content, and target time; specifically, the way in which the mobile phone generates the target location, target content, and target time can be designed according to actual usage needs, and the embodiments of the present application do not impose any restrictions on this.

[0117] In another example, the mobile phone can obtain the target location, target content, target time, etc. specified by the user through voice instructions through the voice assistant function based on the user's voice instructions.

[0118] In some other embodiments, the mobile phone may also obtain media resources generated on other devices (such as the cloud). Alternatively, the mobile phone may also obtain media materials from other devices, generate media resources locally on the mobile phone, and so on. Specifically, the design may be made according to actual usage requirements, and the embodiments of the present application do not impose any restrictions on this.

[0119] Optionally, in the process of filtering photos and / or videos to obtain media materials according to the target condition, the mobile phone can filter media materials within a specified number range. The specified number range of media materials can be less than or equal to 30 media materials.

[0120] As a possible implementation method, the mobile phone can edit the media material by adding filters to the media material, adding background music to the media material, adding text to the media material, adding special effects to the media material, etc. After the mobile phone edits the media material, the mobile phone obtains the media resource. Similarly, the process of editing the media material can also be executed by other devices (such as the cloud). Specifically, it can be designed according to actual usage requirements, and the embodiments of the present application do not impose any restrictions on this.

[0121] In the following embodiments of the present application, the technical solution provided in the embodiments of the present application will be introduced in detail by taking the example of a mobile phone obtaining media materials from the local mobile phone based on the user's voice command and generating media resources.

[0122] Since the mobile phone needs to use the voice assistant function of the mobile phone to obtain the media material from the local mobile phone based on the user's voice command, the voice assistant function of the mobile phone is briefly described below.

[0123] In response to obtaining the wake-up word, the voice assistant function of the mobile phone is awakened, and the mobile phone displays the voice assistant card. After the voice assistant function is awakened, the user can instruct the mobile phone to perform some functions through voice commands. In response to obtaining the voice command, the mobile phone recognizes and analyzes the voice command to obtain the recognition result and analysis result of the voice command. Next, the mobile phone displays the recognition result in the voice assistant card and executes the function indicated by the voice command based on the analysis result.

[0124] For example, see Figure 5 , the mobile phone displays the desktop, and the user inputs the wake-up word into the mobile phone voice, such as "Hello yoyo". In response to obtaining the wake-up word, the voice assistant function of the mobile phone is awakened, and the mobile phone displays the voice assistant card 500. Next, the user inputs a voice command into the mobile phone voice, such as "How is the weather today". The mobile phone recognizes the voice command and obtains the recognition result of the voice command, such as "How is the weather today". And, the mobile phone parses the recognition result and obtains the parsing result of the voice command, such as querying today's weather. After that, the mobile phone executes the parsing result, such as displaying the text 503 of today's weather in the voice assistant card 520. It should be understood that the mobile phone can also display the recognition result 502 of the voice command in the voice assistant card 510. And, the mobile phone can also display a prompt 501 in the voice assistant card 510, such as "Querying today's weather for you". And, for the voice assistant card, it can have different display sizes, which can be set according to the actual use requirements of the voice assistant card, and the embodiment of the present application does not impose any restrictions on this. Among them, the above-mentioned voice assistant card 500 can also include keyboard controls, AI option controls, etc. The keyboard control is used to trigger text input in the voice assistant card, and the AI ​​option control is used to trigger switching to the recommendation menu in the voice assistant card.

[0125] In some embodiments, the voice assistant card can be referred to as a dialog box, the voice command can be referred to as a dialog input, and in the voice card, the interface element displayed by the mobile phone in response to the dialog input can be referred to as an answer. It should be understood that in some embodiments, text input in the voice assistant card through the keyboard control can also be referred to as a dialog input.

[0126] It is understandable that voice commands may include voice commands with fuzzy semantics or voice commands with precise semantics. Among them, voice commands with precise semantics can be understood as the mobile phone being able to obtain a unique and definite parsing result based on the recognition result. Voice commands with fuzzy semantics can be understood as the mobile phone being unable to obtain a parsing result based on the recognition result; or obtaining multiple parsing results based on the recognition result.

[0127] For example, see again Figure 5 After the voice assistant function of the mobile phone is awakened, the user inputs a voice command to the mobile phone, such as "How is the weather?" The mobile phone recognizes and analyzes the voice command, and obtains that the voice command is a voice command with fuzzy semantics. The mobile phone displays a prompt 504 in the voice assistant card 530, such as "Which day do you want to know the weather?" Among them, the prompt 504 is used to prompt the user to enter a voice command with precise semantics. Next, the user inputs a voice command to the mobile phone, such as "How is the weather tomorrow?" The mobile phone recognizes and analyzes the voice command, and obtains that the voice command is a voice command with precise semantics, and displays the text 505 of tomorrow's weather in the voice assistant card 540.

[0128] It should be pointed out that Figure 5 In the scene shown, while the mobile phone displays the prompt 504, tomorrow's weather text 505, prompt 501 and today's weather text 503, the speaker of the mobile phone can also send corresponding sounds. Specifically, this can be designed according to actual usage requirements.

[0129] It should be understood that in the following embodiments of the present application, the mobile phone implements the voice assistant function and the above Figure 5 The corresponding process is similar. For the implementation of the voice assistant function on the mobile phone, please refer to Figure 5 In addition, the following description will focus on the execution process after the mobile phone responds to the voice command, and will not elaborate on the process of the user inputting the voice command to the mobile phone.

[0130] Next, the process of obtaining media materials from the local mobile phone based on the user's voice command and generating media resources is introduced.

[0131] It should be understood that the user can indicate one or more of the target location, the target location, and the target content through voice instructions. The mobile phone can collect the voice instructions and obtain the above-mentioned user can indicate one or more of the target location, the target location, and the target content through voice instructions based on the voice instructions.

[0132] For example, after the voice assistant function of the mobile phone is awakened, in response to obtaining the user's voice command, the mobile phone obtains one or more of the target location, target time, and target content from the voice command. Since voice commands can be divided into voice commands with fuzzy semantics and voice commands with precise semantics, the processing process of the mobile phone when facing the two is slightly different, and the two will be introduced separately in different scenarios.

[0133] In some scenarios, after the voice assistant function of the mobile phone is awakened, the mobile phone obtains a voice command with precise semantics, and the mobile phone obtains media materials based on the voice command.

[0134] For example, see Figure 6 In response to receiving the voice command "Generate Jesse's video", the mobile phone parses the voice command to obtain the target content as "Jesse", obtains the media material based on the target content, and displays the media material 601 in the grid layout in the voice assistant card 600. It should be understood that the user can adjust the order of the media materials in the grid layout, view the media materials in large images, and other operations. In addition, the user can also slide the media materials in the grid layout to switch to display more media materials. In some other examples, the user can also check or uncheck the media material 601 in the grid layout.

[0135] It should be noted that the above-mentioned grid layout can be a 3*3 grid layout or a 4*4, 2*2, etc. grid layout. Figure 6 The grid layout is shown in the form of 3*3. In actual use, it can be set according to actual usage requirements. This application does not impose any restrictions on the form of the grid layout.

[0136] In some embodiments, while the mobile phone displays the media material in a grid layout, the mobile phone can also display a first optional item in the voice assistant card, where the first optional item is used to generate a media resource based on the media material.

[0137] For example, see again Figure 6 , the voice assistant card 600 may include a first optional item 602. In response to the triggering operation of the first optional item 602, the mobile phone displays a voice assistant card 610. The voice assistant card 610 includes a media resource 611, which is generated by the media material in the media material 601 of the grid layout.

[0138] It should be understood that the triggering operation on the first optional item may include: obtaining a voice instruction of "generate video" or a click operation on the first optional item 602. Also, the media resource 611 may be in a playing state or in a waiting state.

[0139] In some embodiments, the voice assistant card 610 also includes some other controls for the media resource 611, such as a "like" control, a "dislike" control, a "save" control, a "share" control, a "delete" control, and the like.

[0140] In some other embodiments, while the mobile phone displays the media material in a grid layout, the mobile phone can also display a second optional item in the voice assistant card, and the second optional item is used to adjust the media material.

[0141] For example, see again Figure 6 In response to the triggering operation of the second option 603, the mobile phone displays a media material detail interface 620. The media material detail interface 620 is used to provide more media materials for the user to view in one interface compared to the media materials in the above-mentioned grid layout. At the same time, the media materials can also be dragged and sorted in the media material detail interface 620.

[0142] Furthermore, the media material details interface 620 can also be used to provide a function of adding media materials. The media material details interface 620 can include an add button control 621. In response to the triggering operation of the add button control 621, the mobile phone displays a search interface 630 of the album application. The user can view, check, uncheck, and so on the media material through the search interface 630. The search interface 630 can include a material check box 631. The material check box 631 includes media materials that are already in a checked state. At the same time, the user can also be provided with an uncheck operation on the media material in the material check box 631, such as clicking a delete control corresponding to the media material. It should be understood that if the user triggers a check operation on a certain media material, the media material can be displayed in the above-mentioned material check box 631. Furthermore, if the user triggers an uncheck operation on a certain media material, the media material is not displayed in the above-mentioned material check box 631.

[0143] In addition, the search interface 630 may also include a trigger button 633. The trigger button 633 is used to provide a user with a trigger to view more other photos and / or videos in the album application. For example, in response to a click operation on the trigger button 633, the mobile phone displays an interface 640 of the album application. For the interface 640 of the album application, please refer to the introduction of the relevant technology, and the embodiment of the present application does not impose any limitation on this.

[0144] Furthermore, the search interface 630 may further include an add control 632. The add control is used to save the user's addition operation on the media material in the current search interface 630, that is, the check operation. In response to the triggering operation on the add control 632, the mobile phone returns to display the above media material details interface 660. In the media material details interface 660, more media materials are displayed than in the above interface 620. Furthermore, the media material details interface 660 also includes a prompt 661, which is used to prompt that the media material has been sorted.

[0145] It should be understood that, considering that when a mobile phone is faced with a large number of media materials, it will take a long time to generate media resources from the media materials in the future, so the mobile phone can pre-configure the upper limit of the number of media materials. If the number of media materials exceeds the above threshold, the mobile phone will display a relevant prompt to prompt the user that the number of media materials exceeds the threshold; at the same time, the media materials will not be added. In this way, the time for generating media resources from media materials in the future can be reduced, and the user experience can be improved. Among them, the number of media materials can be understood as the number of photos and / or videos included in the media materials.

[0146] The upper limit of the number may be 50, 60, 100, etc., and the following description will be made by taking the upper limit of 50 as an example. In some embodiments, the upper limit of the number may be referred to as a first quantity threshold.

[0147] For example, see again Figure 6 In the case where the number of media materials exceeds the upper limit, in response to the user's operation of checking the media material, the mobile phone displays a prompt bubble 651 in the interface 650 to prompt the user that the number of media materials exceeds the upper limit, and the mobile phone does not add the media material.

[0148] In some scenarios, after the voice assistant function of the mobile phone is awakened, the mobile phone receives a voice command with ambiguous semantics; the mobile phone displays a prompt for the voice command to clarify the voice command with ambiguous semantics.

[0149] For example, see Figure 7 In response to obtaining the voice command "generate a video of a classmate", the mobile phone recognizes and parses the voice command. Since the mobile phone does not obtain the accurate target content after recognizing and parsing the voice command, the mobile phone confirms that the voice command is an ambiguous voice command. In response to obtaining the above voice command, the mobile phone displays a voice assistant card 700. The voice assistant card 700 includes a prompt 701 for the voice command. And, the voice assistant card 700 also includes a selection box 702. The selection box 702 is used to provide multiple options for the user to choose from, and each option corresponds to the target content. The user can select the target content by triggering the above options.

[0150] Next, in response to the triggering operation of the option to be selected, the mobile phone takes the option to be selected as the target content. After that, the mobile phone searches for media materials based on the target content and displays the media materials in a grid layout in the voice assistant card.

[0151] For example, see again Figure 7 In response to the trigger operation of the treatment option "Lucy", the mobile phone uses Lucy as the target content. Afterwards, the mobile phone displays the media material 711 of the grid layout in the voice assistant card 710. For a detailed description of the media material 711 of the grid layout, please refer to the above Figure 6 The corresponding introduction is not repeated here.

[0152] In some embodiments, after the mobile phone obtains media materials through photos and / or videos, the mobile phone can also adjust the media materials based on voice commands. It should be understood that the above adjustment of media materials can be understood as: adjusting the amount of media materials, such as adding photos and / or videos to media materials (hereinafter referred to as adding media materials), replacing photos and / or videos in media materials (hereinafter referred to as replacing media materials), reducing photos and / or videos in media materials (hereinafter referred to as reducing media materials), etc.

[0153] As a possible implementation method, the mobile phone can make different adjustments to the media material according to voice commands with ambiguous semantics, voice commands with precise semantics, and the like.

[0154] It should be understood that in the scenario of adjusting the media material, a precise voice instruction can be understood as a voice instruction indicating a certain number, and a vague voice instruction can be understood as a voice instruction that does not indicate a certain number. In other words, the mobile phone can determine a precise voice instruction or a vague voice instruction based on the words indicating the number in the voice instruction.

[0155] For example, "add a lot of Jesse's videos" is a voice command with ambiguous semantics.

[0156] For example, "Add some Johnny's material" is a voice command with ambiguous semantics.

[0157] For another example, "delete 5 photos of Lucy at the AA location" is a semantically precise voice command.

[0158] For example, "delete all Jesse's materials" is a voice command with precise semantics.

[0159] Since the above-mentioned fuzzy semantic voice instructions do not indicate a definite quantity, the mobile phone can specify a target quantity for these fuzzy semantic voice instructions. In other words, the mobile phone can give these fuzzy semantic voice instructions corresponding to the target quantity.

[0160] For a voice instruction for indicating a certain quantity, if the voice instruction includes quantity level indication information, the target quantity can be obtained based on the quantity level indication information and the current quantity of the media material.

[0161] The quantity degree indication information can be understood as information indicating the quantity degree. For example, the quantity degree indicated by "many" is higher than the quantity degree indicated by "some", and the quantity degree indicated by "a large amount" is higher than the quantity degree indicated by "a small amount".

[0162] That is, when the current quantity of the media material is determined, the higher the degree indicated by the quantity degree indication information, the larger the target quantity will be.

[0163] Exemplarily, referring to the following Table 1, the target quantity can be obtained by the mobile phone through the current quantity of the media material.

[0164] Table 1

[0165] Semantic Quantity Condition 1 Semantic Quantity Condition 2 X≤N*30% X≤N*50%

[0166] Wherein, X represents the target quantity, and X is an integer, and N represents the current quantity of the media material. The semantic quantity condition 1 includes: a voice instruction indicating a relatively small (few) quantity of fuzzy semantics, and a voice instruction indicating a fuzzy semantics that does not indicate a quantity. The semantic quantity condition 2 includes: a voice instruction indicating a relatively large (many) quantity of fuzzy semantics.

[0167] The voice instruction indicating the fuzzy semantics of a relatively small quantity can be understood as a voice instruction including the fuzzy semantics of quantifiers such as "a small amount", "a small part", "some", "a small part", "a small amount" and the like.

[0168] The voice instruction indicating the fuzzy semantics of a relatively large quantity can be understood as a voice instruction including the fuzzy semantics of quantifiers such as "a large amount", "most of", "a lot", "most of", "a large amount" and the like.

[0169] The fuzzy semantic voice instruction that does not indicate quantity can be understood as the fuzzy semantic voice instruction that does not include a quantifier.

[0170] For example, “add a large number of Jesse’s videos” is a voice command indicating a relatively large number of ambiguous semantics.

[0171] For another example, "increase Johnny's material" is a voice instruction with ambiguous semantics that does not indicate the quantity.

[0172] For another example, “add some pictures of AA location” is a voice command indicating a relatively small number of ambiguous semantics.

[0173] In some embodiments, with respect to adding media material, the mobile phone may adjust the media material according to a rule for adding media material.

[0174] The above rules for adding media materials include:

[0175] If the current number of media materials does not exceed the above-mentioned upper limit, if the voice command is a voice command with precise semantics, the mobile phone increases the number of media materials indicated by the voice command. Also, if the current number of media materials exceeds the above-mentioned upper limit, if the voice command is a voice command with precise semantics, the mobile phone replaces the number of media materials indicated by the voice command. In other words, the mobile phone will perform the increase or replacement operation based on whether the current number of media materials reaches the upper limit.

[0176] For example, the voice command is "add 5 photos of Lucy at AA location". In response to the voice command, when the current number of media materials does not exceed 50, the mobile phone searches the album for 5 photos and / or videos whose target location is AA location and whose target content is Lucy, and adds these 5 photos and / or videos to the media materials.

[0177] In another example, the voice command is "add 5 photos of Lucy at location AA". In response to the voice command, when the current number of media materials exceeds 50, the mobile phone searches for 5 photos and / or videos whose target location is location AA and whose target content is Lucy from the album, and replaces these 5 photos and / or videos with the five photos and / or videos in the media materials. Among them, the mobile phone can replace the photos and / or videos in the media materials by balanced replacement or aesthetic score replacement. It should be understood that the mobile phone can also replace the photos and / or videos in the media materials by other more replacement methods, which can be specifically set according to actual usage requirements, and the embodiments of the present application are not limited to this.

[0178] Furthermore, the above-mentioned rules for adding media materials may also include:

[0179] If the current number of media materials does not exceed the above upper limit, if the voice command is a voice command with ambiguous semantics, the mobile phone increases the target number of media materials corresponding to the voice command. Also, if the current number of media materials exceeds the above upper limit, if the voice command is a voice command with ambiguous semantics, the mobile phone replaces the target number of media materials corresponding to the voice command. For the target number corresponding to the voice command, please refer to the above description.

[0180] For example, the voice command is "add some photos of Johnny". In response to the voice command, if the current number of media materials does not exceed 50, the mobile phone searches for the target number of Johnny's photos from the album and adds the target number of Johnny's photos to the media materials.

[0181] For another example, the voice command is "add some photos of Johnny". In response to the voice command, when the current number of media materials exceeds 50, the mobile phone searches for the target number of photos of Johnny from the album, and replaces the target number of photos of Johnny with the target number of photos and / or videos in the media materials.

[0182] Furthermore, in the above media material rules, the mobile phone can also control the number of media materials added to not exceed the above upper limit. In other words, for adding media materials, the mobile phone will not increase the number of media materials to exceed the upper limit, and the maximum number of media materials is the above upper limit.

[0183] For example, the voice command is "add some photos of Johnny". In response to the voice command, when the number of media materials is 45, the mobile phone searches for 5 photos of Johnny from the album and adds the 5 photos of Johnny to the media materials.

[0184] For another example, the voice command is "add 10 photos of Johnny". In response to the voice command, when the number of media materials is 46, the mobile phone searches for 4 photos of Johnny from the album and adds the 4 photos of Johnny to the media materials.

[0185] Next, combining the above Figure 7 The scenario shown introduces the process of adjusting the media material by the mobile phone according to the rule of adding media material.

[0186] For example, see Figure 8, the mobile phone displays a voice assistant card 800, and the mobile phone obtains 30 photos and / or videos of Lucy as media materials. In response to obtaining the voice command "Add some photos of Johnny at the AA location", the mobile phone searches the album for the target number of photos, and the target content is Johnny, and the target location is the AA location, and adds these photos to the above media materials. After the mobile phone adds these photos to the above media materials, the mobile phone displays the media material 801 in a grid layout. It should be understood that in the above process, since the current number of media materials is 30, and the voice command is a voice command with fuzzy semantics, it can be seen from the description corresponding to Table 1 above that the target number is 9. Among them, the media material 801 in the above grid layout corresponds to the current number of media materials after the photos are added, that is, 39. And, while the above mobile phone displays the media material 801 in the grid layout, the mobile phone can also display related prompts 802.

[0187] For example, see again Figure 8 , after the mobile phone displays the voice assistant card 800, the user controls the mobile phone through a voice command, and the voice command is "add a lot of photos of Lucy at BB's location", which is a voice command with fuzzy semantics. The current number of media materials is 39. From the description corresponding to the above Table 1, it can be seen that the target number is 20. And because if 20 photos are added to the media material, the number of media materials will exceed the upper limit, that is, 50 photos. The mobile phone will control the number of media materials, such as setting the target number to 11. In response to obtaining the voice command, the mobile phone searches for 11 photos from the album and the target content is Lucy, and the target location is the BB location, and adds these photos to the above media materials. After the mobile phone adds these photos to the above media materials, the mobile phone displays a voice assistant card 810. The media material 811 of the grid layout is displayed in the voice assistant card 810. Among them, the media material 811 of the grid layout corresponds to the current number of media materials after the photos are added, that is, 50 photos.

[0188] For example, see again Figure 8After the mobile phone displays the voice assistant card 810, the user controls the mobile phone through a voice command, and the voice command is "add 5 more photos of Johnny at the AA location", which is a voice command with precise semantics. Since the current number of media materials has reached the upper limit, the mobile phone replaces the 5 photos and / or videos in the media materials with the 5 photos of Johnny at the AA location. After the replacement is completed, the mobile phone displays the media materials 821 in the grid layout in the voice assistant card 820. And, a prompt 822 for the replacement operation. Among them, the voice assistant card 820 also includes a second optional item 823. In response to the triggering operation of the second optional item 823, the mobile phone displays a media material details interface 830. The media material details interface 830 includes the media materials added by the above replacement process, as well as the media materials reduced by the above replacement process.

[0189] In some embodiments, the media material details interface 830 further includes an add button control 831. In response to the triggering operation of the add button control 831, the mobile phone displays a search interface 840 of the photo album application. It should be understood that the search content corresponding to the search interface 840 of the photo album application is related to the voice command obtained by the mobile phone last time. For example, the search content is a search content with the target content being "Johnny" and the target location being "BB location".

[0190] In some embodiments, with respect to reducing media material, the mobile phone may adjust the media material according to a media material reduction rule.

[0191] The above rules for reducing media materials include:

[0192] In the case where the current number of media materials is greater than the recommended number, if the voice instruction is a voice instruction with precise semantics, the mobile phone reduces the number of media materials indicated by the voice instruction, and the number of media materials after the reduction is greater than or equal to the lower limit of the number; and, in the case where the current number of media materials is greater than the recommended number, if the voice instruction is a voice instruction with fuzzy semantics, the mobile phone reduces the target number of media materials corresponding to the voice instruction, and the number of media materials after the reduction is greater than or equal to the lower limit of the number. It should be understood that since the mobile phone needs to generate media resources based on the media materials in the subsequent period, in order to enable the mobile phone to generate media resources based on the media materials, the number of media materials needs to be greater than or equal to the lower limit of the number. Exemplarily, the above lower limit of the number can be 1, 2, 3, etc. And, the lower limit of the number should be smaller than the above recommended number in terms of value, and the recommended number should be smaller than the above upper limit in terms of quantity. In the following, the lower limit of the number will be introduced as 1, and in some examples, the above lower limit of the number can also be referred to as the fifth number.

[0193] The recommended number of media resources can achieve better results when subsequently generating media resources. The recommended number can be 6, 7, 10, etc., and 6 is taken as an example for introduction below. In some embodiments, the recommended number can also be referred to as a second number threshold.

[0194] For example, the voice command is "reduce some photos of Johnny", and the current number of media materials is 30. In response to obtaining the voice command, the mobile phone reduces 9 photos of Johnny in the media materials.

[0195] For another example, the voice command is “reduce 10 photos of Johnny”, and the current number of media materials is 30. In response to obtaining the voice command, the mobile phone reduces 10 photos of Johnny in the media materials.

[0196] For another example, the voice command is "reduce 9 photos of Johnny", and the current number of media materials is 8. In response to obtaining the voice command, the mobile phone maintains the number of reduced media materials greater than or equal to 1, and the mobile phone reduces 7 photos of Johnny in the media materials.

[0197] For another example, the voice command is “reduce most of Jesse’s photos”, and the current number of media materials is 9. In response to obtaining the voice command, the mobile phone reduces 4 photos of Jesse in the media materials.

[0198] Furthermore, the above rules for reducing media materials also include:

[0199] In the case where the current number of media materials is less than or equal to the recommended number, if the voice instruction is a voice instruction with precise semantics, the mobile phone reduces the number of media materials indicated by the voice instruction, and the number of media materials after the reduction is greater than or equal to 1; and, in the case where the current number of media materials is less than or equal to the recommended number, if the voice instruction is a voice instruction with fuzzy semantics, the mobile phone reduces the number of media materials by A, and the number of media materials after the reduction is greater than or equal to 1. In some examples, the above number A can be 1, 2, 4, etc. That is, the mobile phone can make different adjustments to the number of media materials based on the numerical relationship between the current number of media materials and the recommended number.

[0200] It should be understood that when the number of media materials is the recommended number, the quality of the subsequently generated media resources is relatively high. Therefore, when the current number of media materials is less than or equal to the recommended number, the quality of the media resources can be guaranteed by reducing the number of media materials by A in response to the voice command with ambiguous semantics.

[0201] For example, the voice command is "reduce some photos of Johnny", and the current number of media materials is 6. In response to obtaining the voice command, the mobile phone reduces 1 photo of Johnny in the media material.

[0202] For another example, the voice command is “reduce 3 photos of Johnny”, and the current number of media materials is 6. In response to obtaining the voice command, the mobile phone reduces 3 photos of Johnny in the media materials.

[0203] For another example, the voice command is "reduce 10 photos of Johnny", and the current number of media materials is 5. In response to obtaining the voice command, the mobile phone maintains the number of reduced media materials greater than or equal to 1, and the mobile phone reduces 4 photos of Johnny in the media materials.

[0204] For another example, the voice command is “reduce most of Jesse’s photos”, and the current number of media materials is 6. In response to obtaining the voice command, the mobile phone reduces 1 photo of Jesse in the media material.

[0205] Next, combining the above Figure 7 The scenario shown introduces the process of adjusting the media material by the mobile phone according to the rule of reducing the media material.

[0206] For example, see Fig. 9 , the user controls the phone through voice commands, and the voice command is "delete some photos", which is a voice command with ambiguous semantics. The current number of media materials is 30. From the description corresponding to the above Table 1, it can be seen that the target number is 9. In response to the voice command, the mobile phone deletes 9 photos in the media material, and then the mobile phone displays a voice assistant card 900 including media material 901 in a grid layout. Among them, the above-mentioned media material 901 in a grid layout corresponds to the current number of media materials after the photos are reduced, that is, 21 photos.

[0207] For example, see again Fig. 9 After the mobile phone displays the voice assistant card 900, the user controls the mobile phone through a voice command, and the voice command is "help me delete some more photos", which is a voice command with ambiguous semantics. The current number of media materials is 21. From the description corresponding to the above Table 1, it can be seen that the target number is 6. In response to the voice command, the mobile phone deletes 6 photos in the media material, and then the mobile phone displays a voice assistant card 910 including media material 911 in a grid layout. Among them, the above-mentioned media material 911 in a grid layout corresponds to the current number of media materials, that is, 15.

[0208] For example, see again Fig. 9After the mobile phone displays the voice assistant card 910, the user controls the mobile phone through a voice command. The voice command is "help me delete 10 materials", which is a voice command with clear semantics and indicates 10 materials. In response to the voice command, the mobile phone deletes 10 materials from the media materials. After that, the mobile phone displays the voice assistant card 920 including the media material 921 in the grid layout. Among them, the media material 921 in the grid layout corresponds to the current number of media materials, that is, 5.

[0209] For example, see again Fig. 9 After the mobile phone displays the voice assistant card 920, the user controls the mobile phone through a voice command, and the voice command is "help me delete the material at the AA location", that is, a voice command with clear semantics, indicating the AA location. In response to the voice command, the mobile phone deletes the material at the AA location in the media material, and then the mobile phone displays the voice assistant card 930 including the media material 931 in the grid layout. Among them, the media material 931 in the grid layout corresponds to the current number of media materials, that is, 3.

[0210] For example, see again Fig. 9 After the mobile phone displays the voice assistant card 930, the user controls the mobile phone through voice commands, and the voice command is "Help me delete all materials", that is, a voice command with clear semantics. Since the user has instructed to delete all materials through voice commands, the mobile phone needs to keep the number of reduced media materials greater than or equal to 1. Therefore, the mobile phone will reduce 2 materials. In response to the voice command, the mobile phone displays a voice assistant card 940. The voice assistant card 940 includes a prompt 942, which prompts that the media material cannot be further reduced. For example, the prompt 942 can be a text description of "At least 1 material is required to generate a video. Try to keep 1 material to generate a video." In addition, the above-mentioned voice assistant card 940 can also include media material 941 in a grid layout, and the media material 941 in the grid layout corresponds to the current number of media materials, that is, 1.

[0211] For example, see again Fig. 9 After the mobile phone displays the voice assistant card 940, the user controls the mobile phone through a voice command, and the voice command is "help me delete 1 material", that is, a voice command with clear semantics. Since the current number of media materials is equal to the lower limit, the mobile phone does not reduce the media materials. In response to the voice command, the mobile phone displays the voice assistant card 950. The voice assistant card 950 includes a prompt 952, which prompts that the media materials cannot be further reduced.

[0212] In some embodiments, for replacing media material, the mobile phone may adjust the media material according to the replacement media material rule.

[0213] The above-mentioned rules for replacing media materials include:

[0214] If the voice command is a voice command with precise semantics, the mobile phone replaces the media material of the quantity indicated by the voice command. If the voice command is a voice command with ambiguous semantics, the mobile phone replaces the media material of the target quantity corresponding to the voice command.

[0215] For example, the voice command is “replace 5 photos of Johnny with photos of Lucy”. In response to acquiring the voice command, the mobile phone replaces 5 photos of Johnny in the media material with photos of Lucy.

[0216] For another example, the voice command is “replace some of Johnny's photos with Lucy's photos”, and the current number of media materials is 30. In response to obtaining the voice command, the mobile phone replaces 9 of Johnny's photos with Lucy's photos.

[0217] Next, combining the above Figure 7 The scenario shown introduces the process of the mobile phone adjusting the media material according to the rule of replacing the media material.

[0218] For example, see Fig.10 , the user controls the mobile phone through voice commands, and the voice command is "replace a large number of photos", which is a voice command with ambiguous semantics. The current number of media materials is 30. From the description corresponding to the above Table 1, it can be seen that the target number is 15. In response to the voice command, the mobile phone will replace 15 photos in the media material. Afterwards, the mobile phone displays a voice assistant card 1000 including media material 1001 in a grid layout. Among them, the above-mentioned media material 1001 in a grid layout corresponds to the current number of media materials after the photos are replaced, that is, 30 photos.

[0219] For example, see again Fig.10 , after the mobile phone displays the voice assistant card 1000, the user controls the mobile phone through voice commands, and the voice command is "delete Lucy's photos and add Johnny's photos", which is a voice command with ambiguous semantics. The current number of media materials is 30. From the description corresponding to the above Table 1, it can be seen that the target number is 9. In response to the voice command, the mobile phone will replace 9 photos in the media material, such as replacing 9 photos of Lucy with photos of Johnny. Afterwards, the mobile phone displays a voice assistant card 1010 including media material 1011 in a grid layout. Among them, the media material 1011 in the above grid layout corresponds to the current number of media materials after the photos are replaced, that is, 30.

[0220] For example, see again Fig.10After the mobile phone displays the voice assistant card 1010, the user controls the mobile phone through a voice command, and the voice command is "replace five photos of Johnny", which is a voice command with clear semantics. In response to the voice command, the mobile phone will replace the five photos in the media material, such as replacing the five photos of Johnny with photos that are not of Johnny. Afterwards, the mobile phone displays a voice assistant card 1020 including media material 1021 in a grid layout. Among them, the media material 1021 in the grid layout corresponds to the current number of media materials after the photos are replaced, that is, 30 photos.

[0221] After the mobile phone generates the media resources, that is, after the mobile phone executes step S400, the mobile phone executes step S401.

[0222] S401. Edit media resources on mobile phone.

[0223] The above-mentioned editing of media resources may include: changing the background music of the media resources, changing the editing template used by the media resources, adjusting the duration of the media resources, changing the text of the media resources, etc.

[0224] Furthermore, the above-mentioned editing of media resources may further include: adjusting the media material of the media resource. The process of the mobile phone adjusting the media material of the media resource is similar to the above-mentioned step S400, and the relevant description of the above-mentioned step S400 may be referred to, which will not be repeated here.

[0225] It should be understood that in some other embodiments, the above-mentioned editing of media resources may also include adjusting the size of the media resources, adjusting the resolution of the media resources, etc.; the specific design may be based on actual usage requirements.

[0226] In some embodiments, the mobile phone can edit the media resource based on the user's editing.

[0227] The user's editing may be an editing operation performed by the user in the editing interface of the mobile phone; the user's editing may also be a voice command of the user. In other words, the mobile phone may edit the media resource in response to the voice command. For example, the mobile phone may change the background music of the media resource in response to the voice command; for another example, the mobile phone may change the editing template used by the media resource in response to the voice command; for another example, the mobile phone may adjust the duration of the media resource in response to the voice command; for another example, the mobile phone may change the text of the media resource in response to the voice command.

[0228] Regarding the process of editing media resources by the mobile phone based on the editing operation of the user in the editing interface, please refer to the relevant technology, which will not be described here.

[0229] Next, the process of editing media resources by a mobile phone based on a user's voice command is introduced.

[0230] In response to obtaining the voice command, the mobile phone determines the editing content indicated by the voice command, and the editing content may include: one or more of template, duration, background music, and accompanying text.

[0231] For example, in combination with the above Figure 6 The scenario shown in FIG. 1 introduces the process of editing media resources based on the user's voice command on a mobile phone. Fig.11 After the mobile phone generates media resources, the mobile phone displays a voice assistant card 1100. The voice assistant card 1100 includes media resources 1101. In addition, the above-mentioned voice assistant card 1100 may also include an editing prompt control 1102; the editing prompt control 1102 corresponds to the editing content, and the editing prompt control is used to provide editing operations on the editing content, and can also be used to prompt the user of the editing content that can be executed in the current card. For example, the editing prompt control 1102 includes: a control 1102a corresponding to the text, a control 1102b corresponding to the template, a control 1102c corresponding to the background music, and a control 1102d corresponding to the duration.

[0232] It should be pointed out that Fig.11 and underlined numbers in other figures, such as Fig.11 1102a, 1102b, etc. are all figure marks and should not be understood as interface elements displayed on the mobile phone.

[0233] For example, see again Fig.11 In response to the triggering operation of the control 1102a corresponding to the text, the mobile phone displays a voice assistant card 1110. The voice assistant card 1110 includes a text editing box 1113. Among them, the text editing box 1113 is used to provide editing operations for the text of the media resource. And the voice assistant card 1110 also includes an editing prompt control 1112. The editing prompt control includes a control 1112b corresponding to the template, a control 1112c corresponding to the background music, and a control 1112d corresponding to the duration; the editing prompt control does not include a control corresponding to the text.

[0234] It should be understood that the triggering operation of the control 1102a corresponding to the above-mentioned text may be a click operation on the control 1102a corresponding to the text, or may be a voice command on the control corresponding to the text, such as “smart text”.

[0235] For example, see again Fig.11After the mobile phone displays the voice assistant card 1100, in response to the triggering operation of the control 1102b corresponding to the template, the mobile phone displays the voice assistant card 1120. The voice assistant card 1120 includes: an edit menu control 1123. The "Change Template" menu of the edit menu control 1123 is in an expanded state. The "Change Template" menu includes: a "Cute" template option, a "Movie Feel" template option, a "Texture" template option, and a "Simple" template option.

[0236] Afterwards, in response to the selection operation of the "texture" template option, the mobile phone changes the template used by the media resource to the "texture" template, and the mobile phone displays the voice assistant card 1130; the voice assistant card 1130 includes an edit menu control 1133. The "Change Template" menu of the edit menu control 1133 is in an expanded state, and the "texture" template option is in a selected state. In addition, the edit menu control 1133 may also include an unexpanded "Change Music" menu, an unexpanded "Change Duration" menu, an expansion control corresponding to the "Change Duration" menu, and an expansion control corresponding to the "Change Music" menu.

[0237] The above-mentioned selection operation of the "texture" template option may be a click operation on the "texture" template option, or may be a voice command, such as "select a texture template" and the like.

[0238] Next, in response to the expansion operation of the "Change Music" menu, the mobile phone displays a voice assistant card 1140; the voice assistant card 1140 includes an edit menu control 1143, and the "Change Music" menu of the edit menu control 1143 is in an expanded state. The "Change Music" menu includes a "Light" music option, a "Romantic" music option, a "Dynamic" music option, and a "Flow" music option.

[0239] Among them, the above-mentioned expansion operation of the "Change Music" menu can be a click operation on the expansion control corresponding to the "Change Music" menu, or it can be a voice command, such as, "Expand the Change Music menu" and so on.

[0240] Afterwards, in response to the selection operation of the "romantic" music option, the mobile phone displays a voice assistant card 1150; the voice assistant card 1150 includes an edit menu control 1153, and the "romantic" music option in the edit menu control 1153 is in a selected state.

[0241] The selection operation of the "romantic" music option may be a click operation on the "romantic" music option, or may be a voice command, such as "select the romantic music option" and the like.

[0242] Then, in response to the expansion operation of the "Change Duration" menu, the mobile phone displays a voice assistant card 1160; the voice assistant card 1160 includes an edit menu control 1163, and the "Change Duration" menu of the edit menu control 1163 is in an expanded state. The "Change Duration" menu includes a "55s" duration option, a "60s" duration option, a "70s" duration option, and a "Custom" duration option.

[0243] Among them, the above-mentioned expansion operation of the "change duration" menu can be a click operation on the expansion control corresponding to the "change duration" menu, or it can be a voice command, such as "expand the change duration menu" and so on.

[0244] Next, in response to the selection operation of the "custom" duration option, the mobile phone displays a voice assistant card 1170. The voice assistant card 1170 includes a custom duration pop-up window 1174. The custom duration pop-up window 1174 is used to provide the user with a custom change in the duration of the media resource. For example, the user can customize the duration of the media resource by dragging the option bar in the custom duration pop-up window 1174.

[0245] Optionally, the above-mentioned custom change of the duration of the media resource is changed within a duration range, and the duration range is related to the media material included in the media resource, or the duration range may also be preset.

[0246] Exemplarily, the start point and the end point of the duration range may be calculated by the following expression.

[0247] T1=a*X+b*YExpression 1

[0248] T2=c*X+d*YExpression 2

[0249] Among them, the above T1 is the starting point of the duration range, T2 is the end point of the duration range, a is the first coefficient, b is the second coefficient, c is the third coefficient, d is the fourth coefficient, X is the number of photos included in the media material, and Y is the total duration of the video included in the media material.

[0250] For example, a may be 0.5, b may be 1, c may be 3, and d may be 1.5.

[0251] For another example, the starting point of the duration range may be calculated by the above expression 1, and the end point of the duration range may be preset to be, for example, 90 seconds.

[0252] Among them, the above-mentioned selection operation of the "custom" duration option can be a click operation on the "custom" duration option, or it can be a voice command, such as "select the custom duration option" and so on.

[0253] Afterwards, in response to the selection operation of the "custom" duration option, the mobile phone adjusts the duration of the media resource to 65 seconds, and the mobile phone displays the voice assistant card 1180. The voice assistant card 1180 includes an edit menu control 1183, and the "custom" duration option in the edit menu control 1183 is selected. In addition, the edit menu control 1183 also includes an "OK" option control.

[0254] Next, in response to the triggering operation of the "OK" option control, the mobile phone displays a voice assistant card 1190, which includes some prompts. In addition, the voice assistant card also includes a media resource 1193, the template used by the media resource 1193 is the "texture" template, the background music is the "romantic" background music, and the duration is 65 seconds.

[0255] As another possible design, see Fig.12 After the mobile phone generates a video, in response to receiving a voice command, such as changing music, the mobile phone displays a voice assistant card 1200; the voice assistant card 1200 includes an edit menu control 1203. The "Change Music" menu of the edit menu control 1203 is in an expanded state. In other words, the expansion state of the menu in the edit menu control is related to the voice command. The menu corresponding to the editing content indicated by the voice command will be expanded in the edit menu control.

[0256] In some other possible designs, such as the above Fig.11 As shown in the voice assistant card 1100 in FIG. 1 , the voice assistant card 1100 includes an edit prompt control 1102. After the mobile phone displays the voice assistant card 1100, in response to obtaining a voice instruction, such as changing music, the mobile phone displays a voice assistant card 1210; the voice assistant card does not include an edit prompt control for changing music.

[0257] In some other possible designs, the mobile phone displays the above Fig.11 After the voice assistant card 1100 is shown, in response to obtaining a voice command, such as changing music, the mobile phone displays a voice assistant card 1220. The voice assistant card includes a prompt 1221, which is used to prompt the trigger operation of the current editing content. It should be understood that the mobile phone can pre-configure multiple prompts for editing content, and after the user edits the media resource, the prompts are displayed in rotation.

[0258] In some other possible designs, the mobile phone displays the above Fig.11After the voice assistant card 1100 is shown, in response to obtaining a voice instruction, such as changing the duration, the mobile phone displays a voice assistant card 1230; the voice assistant card 1230 includes an edit menu control 1233, and the "Change Duration" menu of the edit menu control 1233 is in an expanded state. In addition, the "Change Duration" menu includes an option bar, which is used to provide a user with a custom adjustment of the duration of the media resource.

[0259] In some embodiments, for some voice instructions indicating editing content; if the mobile phone can match the editing content indicated by the voice instruction, the mobile phone edits the media resource based on the editing content; if the mobile phone cannot match the editing content indicated by the voice instruction, the mobile phone displays an editing menu control.

[0260] For example, see Fig.13 , the voice command is "change the movie-feeling template", and the mobile phone can match the movie-feeling template. In response to obtaining the voice command, the mobile phone changes the movie-feeling template for the media resource, and the mobile phone displays a voice assistant card 1300; the voice assistant card includes a media resource 1301 using the movie-feeling template.

[0261] Again, see Fig.13 , the voice command is "change dopamine template", the mobile phone can recognize from the voice command that the voice command indicates to change the template, but the mobile phone cannot match the dopamine template. In response to obtaining the voice command, the mobile phone displays a voice assistant card 1310; the voice assistant card 1310 includes an edit menu control 1313. Among them, the "change template" menu of the edit menu control 1313 is in an expanded state.

[0262] For example, see again Fig.13 , the voice command is "change to light music", and the mobile phone can match light music. In response to obtaining the voice command, the mobile phone changes the background music of the media resource to light music, and the mobile phone displays a voice assistant card 1320; the voice assistant card includes a media resource 1321 whose background music is light music.

[0263] Again, see Fig.13 , the voice command is "change elegant music", the mobile phone can recognize from the voice command that the voice command indicates to change music, but the mobile phone cannot match elegant music. In response to obtaining the voice command, the mobile phone displays a voice assistant card 1330; the voice assistant card 1330 includes an edit menu control 1333. Among them, the "change music" menu of the edit menu control 1333 is in an expanded state.

[0264] For example, see again Fig.13, the voice command is "add champagne to the video", and the mobile phone fails to recognize the editing content from the semantic command. In response to obtaining the voice command, the mobile phone displays a voice assistant card 1340; the voice assistant card 1340 includes an editing menu control 1343. Among them, the menus included in the editing menu control 1343 are all in a folded state.

[0265] In some embodiments, the duration range of the media resource is related to the media material included in the media resource, and the mobile phone obtains the duration indicated by the voice instruction; if the duration indicated by the voice instruction is within the above duration range, the mobile phone adjusts the duration of the media resource based on the duration indicated by the voice instruction; if the duration indicated by the voice instruction is not within the above duration range, the mobile phone prompts that the duration indicated by the voice instruction exceeds the duration range; and if the voice instruction does not indicate the duration and is a voice instruction with ambiguous semantics, the mobile phone randomly selects a duration from the first range to adjust the duration of the media resource, and at the same time, the adjusted duration of the media resource is within the above duration range. The first range can be: (5, 10), (6, 14), (8, 15), etc.

[0266] For example, in Figure 6 In the scenario shown, the media resource includes 30 media materials, the duration of the media resource ranges from [15, 90], and when the mobile phone generates the media resource, the duration of the media resource is 30 seconds. Fig.14 , the voice instruction is "shorten the duration a little bit", which is a voice instruction with ambiguous voice and does not indicate the value of the duration adjustment. In response to obtaining the voice instruction, the mobile phone shortens the duration of the media resource by 5 seconds and displays a voice assistant card 1400; the voice assistant card 1400 includes a media material 1401, and the duration of the media material 1401 is 25 seconds.

[0267] Again, see Fig.14 , after the mobile phone displays the voice assistant card 1400, the voice command is "shorten by 2 seconds". In response to receiving the voice command, the mobile phone shortens the duration of the media resource by 2 seconds and displays the voice assistant card 1410; the voice assistant card 1410 includes the media material 1411, and the duration of the media material 1411 is 23 seconds.

[0268] For example, see again Fig.14 After the mobile phone displays the voice assistant card 1410, the voice command is "adjust to 20 seconds". In response to receiving the voice command, the mobile phone adjusts the duration of the media resource to 20 seconds and displays the voice assistant card 1420; the voice assistant card 1420 includes the media material 1421, and the duration of the media material 1421 is 20 seconds.

[0269] Again, see Fig.14, after the mobile phone displays the voice assistant card 1420, the voice command is "adjust to 1 hour". In response to receiving the voice command, since 1 hour exceeds the above duration range, the mobile phone adjusts the duration of the media resource to 90 seconds and displays the voice assistant card 1430; the voice assistant card 1430 includes the media material 1431, and the duration of the media material 1431 is 1 minute and 30 seconds, that is, 90 seconds.

[0270] For example, see again Fig.14 , after the mobile phone displays the voice assistant card 1430, the voice command is "shorten to 2 seconds". In response to receiving the voice command, since 2 seconds exceeds the above duration range, the mobile phone adjusts the duration of the media resource to 15 seconds and displays the voice assistant card 1440; the voice assistant card 1440 includes the media material 1441, and the duration of the media material 1441 is 15 seconds.

[0271] It should be noted that the personal information used in the technical solution of this application is limited to information for which the individual’s separate consent has been obtained, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) and sign the agreement (authorization) including authorization of relevant user information before the user uses the function.

[0272] In combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to be beyond the scope of the present application.

[0273] In this embodiment, the electronic device can be divided into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of modules in this embodiment is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0274] The present application also provides an electronic device, such as Fig.15 As shown, the electronic device may include one or more processors 1601 , a memory 1602 and a communication interface 1603 .

[0275] The memory 1602 and the communication interface 1603 are coupled to the processor 1601. For example, the memory 1602, the communication interface 1603 and the processor 1601 may be coupled together via a bus 1604.

[0276] The communication interface 1603 is used for data transmission with other devices. The memory 1602 stores computer program code. The computer program code includes computer instructions. When the computer instructions are executed by the processor 1601, the electronic device executes the relevant method steps in the above method embodiment of the present application.

[0277] Among them, the processor 1601 can be a processor or a controller, for example, a central processing unit (CPU), a general processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the present disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0278] The bus 1604 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus 1604 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.15 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0279] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program code is stored. When the processor executes the computer program code, the electronic device executes the relevant method steps in the method embodiment.

[0280] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the relevant method steps in the above method embodiment.

[0281] Among them, the electronic device, computer-readable storage medium or computer program product provided in this application is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0282] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0283] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0284] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0285] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0286] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0287] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for editing media resources, characterized in that: Applied to an electronic device, wherein the electronic device includes a voice assistant, the method includes: When the voice assistant is awakened, in response to a first dialogue input, a dialog box is displayed, the first dialogue input indicates generating a media resource, the dialog box includes a first answer, the first answer includes a first quantity of media materials, and the first quantity is less than a first quantity threshold; In response to a second dialogue input, displaying a second answer in the dialog box, wherein the second answer includes a second amount of media material; the second amount is greater than or equal to the first amount threshold, and the second dialogue input indicates adding media material; In response to a third dialogue input, displaying a third answer in the dialog box, the third answer including the second amount of media materials, the third dialogue input indicating adding media materials; The second number of media materials in the third answer is different from at least one media material in the second number of media materials in the second answer.

2. The method according to claim 1, characterized in that After displaying the third answer in the dialog box, the method further includes: In response to a fourth dialogue input, displaying a fourth answer in the dialog box, the fourth answer including a third quantity of media materials, the third quantity being less than or equal to a second quantity threshold, the second quantity threshold being less than the first quantity threshold, the fourth dialogue input indicating a reduction in media materials; In response to a fourth dialogue input, displaying a fifth answer in the dialog box, the fifth answer including a fourth amount of media material, the fourth amount being less than the third amount, the fourth dialogue input indicating a reduction in media material; The difference between the third number and the fourth number is smaller than the difference between the second number and the third number.

3. The method according to claim 2, characterized in that After displaying the fifth answer in the dialog box, the method further includes: After the media material is reduced to a fifth quantity, in response to a fifth dialogue input, a prompt is given in the dialog box indicating that the media material cannot be further reduced, and the fifth quantity is less than or equal to the third quantity threshold.

4. The method according to any one of claims 1 to 3, characterized in that In the case where the second dialogue input indicates to increase a specified amount of media material, the difference between the second amount and the first amount is the specified amount; When the second dialogue input does not indicate to increase the specified amount of media material, the difference between the second amount and the first amount is positively correlated with the first amount.

5. The method according to claim 4, characterized in that The second dialogue input includes quantity degree indication information, and the higher the degree indicated by the quantity degree indication information, the greater the difference between the second quantity and the first quantity.

6. The method according to any one of claims 1 to 5, characterized in that: In a case where the fourth dialogue input does not indicate reducing a specified amount of media materials, the difference between the second amount and the third amount is positively correlated with the amount of media materials when the fourth dialogue input is received.

7. The method according to claim 1, characterized in that The dialog box further includes a first control. After displaying the third answer in the dialog box, the method further includes: In response to a triggering operation on the first control, a first interface is displayed, the first interface including media materials in a selected state and media materials in a non-selected state; the media materials in the selected state are media materials included in the third answer, and the media materials in the non-selected state are media materials included in the third answer and not included in the second answer.

8. The method according to claim 1, characterized in that After displaying the third answer in the dialog box, the method further includes: In response to a sixth dialogue input, a sixth answer is displayed in the dialog box, the sixth answer includes a first media resource, the first media resource includes the media material included in the third answer, and the sixth dialogue input indicates generating a media resource with the media material included in the third answer.

9. The method according to claim 8, characterized in that After displaying the sixth answer in the dialog box, the method further includes: In response to a seventh dialog input, displaying a seventh answer in the dialog box, the seventh answer comprising a media editing control; In the case where the seventh dialogue input indicates a designated editing content, a menu option corresponding to the designated editing content in the media editing control is in an expanded state; In the case where the seventh dialogue input does not indicate a designated editing content, the menu option in the media editing control is in an unexpanded state; Among them, the designated editing content includes: one or more of changing template, changing background music and changing duration; the changing template corresponds to the first menu option, the changing background music corresponds to the second menu option, and the changing duration corresponds to the third menu option.

10. The method according to claim 9, characterized in that The method further comprises: In response to the seventh dialogue input, a guide bubble is displayed in the dialog box, where the guide bubble is used to prompt the user to issue a voice instruction.

11. The method according to claim 8, characterized in that After displaying the third answer in the dialog box, the method further includes: In response to an eighth dialog input, displaying an eighth answer in the dialog box, the eighth answer comprising a second media resource, the second media resource comprising the same media material as the first media resource; In the case where the eighth dialogue input indicates a specified duration, the difference in duration between the second media resource and the first media resource is the specified duration; When the eighth object input does not indicate a specified duration, the difference in duration between the second media resource and the first media resource is within a first range.

12. An electronic device, characterized in that: The electronic device includes a processor and a memory; the processor is coupled to the memory; the memory is used to store computer program code; the computer program code includes computer instructions, and when the processor executes the above-mentioned computer instructions, the electronic device executes the method described in any one of claims 1-11.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer instructions, and when the computer instructions are executed on an electronic device, the electronic device executes the method according to any one of claims 1 to 11.

Citation Information

Cited By

  • Media resource editing method, electronic device, and storage medium

    WO2025086822A1