A photo processing method and electronic device
By receiving user voice commands, recognizing and generating photos and videos of the device owner, the problem of low human-computer interaction efficiency in existing technologies is solved, and user engagement and interaction efficiency are improved.
Patent Information
- Application Number
- CN202410046046.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-01-10
AI Technical Summary
In existing technologies, when electronic devices generate videos through voice interaction, the human-computer interaction efficiency is low, the user participation is insufficient, and it is difficult to efficiently identify and generate photos and videos of the device owner.
By receiving voice commands input from users, identifying and determining the owner's instructions, and generating videos based on photos of the owner included in the photo library application, the efficiency of human-computer interaction is improved, and user engagement is enhanced through confidence filtering and photo collection management.
It enables efficient identification and generation of photos and videos of the device owner in electronic devices, improving human-computer interaction efficiency and user engagement.
Smart Images

Figure CN119277004B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a photo processing method and an electronic device. Background Technology
[0002] With the development of terminal technology and the increasing maturity of speech recognition technology, voice input has become increasingly important due to its high naturalness and effectiveness in interaction. Voice interaction applications within electronic devices are also a frequently used function. Users can interact with electronic devices (such as mobile phones, tablets, and smartwatches) via voice to complete various operations such as command input, information retrieval, and voice chat. Summary of the Invention
[0003] This application provides a photo processing method and an electronic device that can determine instructions for the device owner based on user-inputted voice, and generate videos based on photos of the device owner included in a photo library application in response to the instructions, thereby improving human-computer interaction efficiency.
[0004] The embodiments of this application adopt the following technical solutions:
[0005] Firstly, a photo processing method is provided, applied in an electronic device. This method includes: the electronic device receiving first voice input from a user, and displaying first text corresponding to the first voice input on a first interactive interface. In other words, the first interactive interface of the electronic device can display text content corresponding to the user-input voice.
[0006] Based on this, the first text includes information instructing the device owner; that is, the first voice message includes information for instructing the device owner. In response to the first voice message, the electronic device displays a first target photo and a first control on the first interactive interface. The first target photo includes a first face, which is tagged with the device owner in the electronic device's gallery application.
[0007] In response to a user's first input to a first control, the electronic device displays a thumbnail of a target video; wherein the target video is generated by the electronic device based on a first target photo including a first face in a gallery application. Thus, by recognizing the user's first voice input, information indicating the device owner is determined; furthermore, by responding to this first voice, a video can be generated from a photo of the device owner in the electronic device's gallery application, improving the efficiency of human-computer interaction.
[0008] Optionally, the electronic device may receive first voice input from the user without displaying a first interactive interface. The first voice includes information to instruct the device owner, and in response to the first voice, the device retrieves photos including the device owner from a gallery application and generates a video.
[0009] In one possible implementation of the first aspect, the method further includes: the electronic device receiving a second operation from a user to delete a device owner tag in a gallery application; in response to a second voice input by the user, the electronic device displaying second text corresponding to the second voice on a first interactive interface, the second text including information indicating the device owner.
[0010] In response to a second voice command, the electronic device displays a second target photo, a third target photo, and a first control on the first interactive interface. The second target photo contains a second face, and the third target photo contains a third face. In response to a third user action, the second target photo is unchecked. In response to a fourth user action on the first control, the electronic device displays a thumbnail of the target video, which is generated by the electronic device based on the third target photo (containing the third face) in the gallery application. Thus, after the user deletes the device owner's tag in the gallery application, the electronic device can display photos of multiple device owners on the first interactive interface, with all photos of device owners selected by default. Furthermore, in response to the user's unchecking action, the selected device owner's photo (i.e., the third target photo) is retained, and a video is generated based on the selected photo, improving both human-computer interaction efficiency and user engagement.
[0011] In one possible implementation of the first aspect, the second target photo is determined by the electronic device based on a front-facing photo containing the second face from all photos included in a gallery application.
[0012] In one possible implementation of the first aspect, the number of front-facing photos containing the second face is the largest among all front-facing photos containing the face; or, the number of front-facing photos containing the second face is greater than or equal to a selfie count threshold.
[0013] In one possible implementation of the first aspect, the method further includes: an electronic device displaying a second interactive interface, the second interactive interface including a first location and a second location; wherein the first location is a home location learned by the electronic device, the second location is a company location learned by the electronic device, and a photo containing a third person's face has appeared in both the first and second locations.
[0014] In one possible implementation of the first aspect, a second control is also displayed on the first interactive interface; the method further includes: in response to a fifth operation input by the user to the second control, the electronic device expands to display multiple photos including the first face. Thus, the user can see multiple photos, including the first face, obtained by the electronic device from a gallery application, enhancing user engagement.
[0015] In one possible implementation of the first aspect, before displaying the first target photo, the method further includes: the electronic device determining that the first voice includes an instruction to the device owner, and querying the device owner's face ID; the electronic device determining that the device owner's face ID is the first face based on the confidence level corresponding to different face IDs; the confidence level of the first face is a first confidence level, which is greater than the confidence levels of other faces; the electronic device filtering photos including the first face to determine the first target photo. Thus, by querying the device owner's face ID through confidence level, photos of the first face are filtered out, improving the reliability of filtering the first face photo.
[0016] In one possible implementation of the first aspect, before displaying the second target photo and the third target photo on the first interactive interface, the method further includes: the electronic device determining that the second voice includes an instruction to the owner, and querying the owner's face ID; the electronic device determining that the owner's face ID is the second face and the third face based on the confidence level corresponding to different face IDs; and the electronic device filtering photos including the second face and the third face to determine the second target photo and the third target photo.
[0017] In one possible implementation of the first aspect, before the electronic device receives the first voice input from the user, the method further includes: the electronic device generating multiple photo sets of different people based on different face ID analysis results; the electronic device setting a device owner tag for the photo set containing the first face; and the electronic device determining the first confidence level of the first face as a first preset value.
[0018] In one possible implementation of the first aspect, before the electronic device receives the second voice input from the user, the method further includes: the electronic device acquiring the gender and age of each face among multiple faces included in all front-facing photos in a gallery application; the electronic device filtering out a first group of faces based on the gender and age of each face among the multiple faces; the gender of each face in the first group of faces is the same as the gender of a second face, and the difference between the age of each face and the age of the second face is within a preset range; the electronic device using the ratio of the number of front-facing photos of the second face to the number of front-facing photos of the first group of faces as a second confidence level; the second confidence level is less than a first preset value; the electronic device reading photos in the gallery application and acquiring the shooting location and shooting time of the photos corresponding to each face ID in all face IDs; the electronic device determining the target face ID corresponding to the photos taken simultaneously at a first location and a second location based on the shooting location and shooting time of the photos corresponding to each face ID in all face IDs; the target face ID includes a third face; the electronic device setting the third confidence level corresponding to the third face to a second preset value, the second preset value being less than the first preset value.
[0019] In one possible implementation of the first aspect, the method further includes: the electronic device determining whether the highest confidence value among the confidence values corresponding to different face IDs is a first preset value; when the highest confidence value is the first preset value, returning the face ID with the first preset confidence value as the face ID of the device owner; when the highest confidence value is not the first preset value, returning all face IDs with confidence values exceeding a preset threshold as the face ID of the device owner.
[0020] In one possible implementation of the first aspect, the electronic device determines the target face ID corresponding to photos taken simultaneously at a first location and a second location based on the shooting location and shooting time of the photos corresponding to each face ID in all face IDs. This includes: the electronic device determining that the distance between the shooting location and the first location is less than or equal to a first preset distance, and the distance between the shooting location and the second location is less than or equal to a second preset distance, based on the shooting location of the photos corresponding to each face ID in all face IDs; the electronic device determining the frequency of photos taken at the first location and the frequency of photos taken at the second location for each face ID in all face IDs; and the electronic device filtering out the target face IDs corresponding to photos taken simultaneously at the first location and the second location based on the shooting time; wherein the frequency of photos taken at the first location for the target face ID is greater than a preset frequency, and the frequency of photos taken at the second location is greater than a preset frequency.
[0021] In one possible implementation of the first aspect, before the electronic device determines the target face ID corresponding to the photos taken simultaneously at the first location and the second location based on the shooting location and shooting time of the photos corresponding to each face ID in the total face IDs, the method further includes: the electronic device determining that the distance between the first location and the second location is greater than or equal to a preset distance. Thus, when the distance between the first location and the second location is greater than or equal to the preset distance, the first location and the second location can be considered relatively accurate, facilitating the subsequent execution of the process to determine the device owner.
[0022] In one possible implementation of the first aspect, before the electronic device receives the second voice input from the user, the method further includes: the electronic device determining the age range of the device owner based on the user's usage habits. The method also includes: if the number of target face IDs is less than or equal to a preset number, the electronic device uses the face IDs within the age range from the target face IDs as the device owner's face ID.
[0023] Optionally, if the number of target face IDs exceeds a preset number, the electronic device will terminate the process of determining the owner to ensure the accuracy of the determined owner.
[0024] In a second aspect, an electronic device is provided, which has the functions described in any one of the first aspects above. These functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions.
[0025] Thirdly, an electronic device is provided, comprising: a memory and one or more processors, and a display screen; the memory stores computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method described in the first aspect or any one of the first aspects.
[0026] Fourthly, a chip system is provided, the chip system comprising: at least one processor and an interface for receiving instructions and transmitting them to the at least one processor; the at least one processor executes instructions to cause an electronic device to perform the method described in any one of the first aspects.
[0027] Fifthly, a computer-readable storage medium is provided that stores instructions which, when executed on a computer, cause the computer to perform the method described in any one of the first aspects.
[0028] In a sixth aspect, a computer program product containing instructions is provided, which, when run on a computer, enables the computer to perform the method described in any one of the first aspects above.
[0029] The technical effects of any of the implementation methods in the second to sixth aspects mentioned above can be referred to the technical effects of different implementation methods in the first aspect, and will not be elaborated here. Attached Figure Description
[0030] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0031] Figure 2A This application provides an illustration of a scenario for waking up a mobile phone. Figure 1 ;
[0032] Figure 2B A second schematic diagram illustrating a scenario for waking up a mobile phone, provided as an embodiment of this application;
[0033] Figure 2C This application provides an illustration of a scenario for waking up a mobile phone. Figure 3 ;
[0034] Figure 3 A schematic diagram of a voice interaction interface provided in an embodiment of this application. Figure 1 ;
[0035] Figure 4 A second schematic diagram of a voice interaction interface provided in an embodiment of this application;
[0036] Figure 5 A schematic diagram illustrating multiple photo collections included in a gallery application, as provided in an embodiment of this application;
[0037] Figure 6 A schematic diagram illustrating the addition of names to a collection of portrait photos in a gallery application, as provided in an embodiment of this application;
[0038] Figure 7 A schematic diagram of a process for determining the owner of a device provided in this application embodiment. Figure 1 ;
[0039] Figure 8 A second schematic diagram illustrating the process of determining the owner of a device, provided as an embodiment of this application;
[0040] Figure 9 A schematic diagram of a user trajectory provided in an embodiment of this application;
[0041] Figure 10 A schematic diagram of an interactive interface including home location and company location provided for an embodiment of this application;
[0042] Figure 11 A schematic diagram of a voice interaction interface provided in an embodiment of this application. Figure 3 ;
[0043] Figure 12 A schematic diagram of the software framework of an electronic device provided in an embodiment of this application;
[0044] Figure 13 A schematic diagram of a photo processing flow provided in an embodiment of this application;
[0045] Figure 14 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation
[0046] This application provides a photo processing method applicable to electronic devices such as mobile phones, tablets, and smartwatches that support voice interaction. The electronic device can provide a voice interaction interface to the user, which can receive voice input from the user. In response to the user's voice input, the electronic device determines an instruction for the user. Then, in response to the instruction, the electronic device retrieves photos including the user's photos from a photo library application and generates a video.
[0047] In summary, by adopting the solution of this application embodiment, the electronic device can determine the instruction for instructing the owner based on the user's voice input, and then generate a video based on the owner's photos included in the gallery application in response to the instruction for instructing the owner, thereby improving the efficiency of human-computer interaction.
[0048] For example, the electronic devices in this application embodiment may be mobile phones, tablets, desktops, laptops, handheld computers, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, cellular phones, and personal digital assistants (PDAs). Augmented reality (AR) / virtual reality (VR) devices and other electronic devices that support voice interaction functions are also included. This application embodiment does not impose special limitations on the specific form of the electronic devices.
[0049] refer to Figure 1 This is a hardware structure diagram of an electronic device 100 provided in an embodiment of this application. Figure 1 As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and subscriber identification module (SIM) card interfaces 1 to N 195, etc.
[0050] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0051] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0052] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0053] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via a USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device 100 via the power management module 141.
[0054] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.
[0055] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0056] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0057] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0058] The electronic device 100 can perform shooting functions through an ISP, a camera 193, a video codec, a GPU, a display screen, and an application processor. The ISP is used to process data fed back by the camera 193. The camera 193 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1.
[0059] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.
[0060] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the electronic device (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.
[0061] Electronic device 100 can implement audio functions through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor, such as music playback and recording.
[0062] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch buttons. Electronic device 100 can receive button inputs and generate key signal inputs related to user settings and function control. Motor 191 can generate vibration alerts; motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with electronic device 100. Electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1.
[0063] Of course, the electronic device 100 provided in this application embodiment may also include one or more devices such as a positioning module 181, a button 190, a motor 191, an indicator 192, and a SIM card interface 195, without limitation.
[0064] The methods described in the following embodiments can all be implemented in the electronic device 100 with the above-described hardware structure. It is understood that the photo processing method provided in this application can be executed by electronic devices supporting voice interaction, such as mobile phones, tablets, and smartwatches; it can also be executed by a chip, chip system, or processor capable of implementing the photo processing method; or it can be executed by a logic module or software capable of implementing all or part of the functions of the electronic device, without limitation. The following embodiments use a mobile phone as the execution subject to provide a detailed description of the solution provided in this application.
[0065] In some scenarios, many mobile phones support voice control. Specifically, the phone uses a microphone to capture the user's voice, analyzes and recognizes the voice, and executes the corresponding commands, allowing the user to control the phone via voice.
[0066] Generally, to conserve power and prevent accidental triggering, voice control needs to be enabled before use. For example, a user can wake up the phone by inputting a preset word (or wake-up word) via voice. Once awake, the phone can execute the corresponding voice command, thus activating voice control. Users can also enable or disable voice control by toggling preset switches on or off through the phone's user interface.
[0067] Voice control functions may have different names on different mobile phones, such as "smart assistant", "smart application", "voice assistant", etc. The specific implementation of voice control functions with different names may also be somewhat different.
[0068] The following section uses "Smart Assistant" as an example to illustrate several different implementations of voice control functions.
[0069] Before using the smart assistant to control their phone, users need to wake the phone by inputting a wake word via voice. Normally, before the phone is woken up, the microphone operates in a power-saving mode (e.g., searching for signals at lower power levels) to pick up ambient sounds. The microphone only detects the wake word at the kernel level; the phone's system and drivers do not activate the corresponding recording channel for the smart assistant.
[0070] In response to user actions (such as receiving a wake-up word from the user's voice), the phone is woken up, and the corresponding recording channel for the smart assistant is activated in the system and drivers. After the phone is woken up, the voice (audio stream) captured by the microphone is sent to the smart assistant application for processing through the corresponding recording channel. In this way, the phone can execute the user's voice commands, enabling the user to control the phone via voice. Optionally, the phone can also perform functions such as conversing with the user.
[0071] refer to Figure 2A It illustrates a scenario for waking up a mobile phone. For example... Figure 2A As shown, the mobile phone displays a desktop interface. The phone receives a voice wake-up word input by the user (e.g., "Hello, YOYO"), and wakes up in response to this wake-up word. For example, after being woken up, the phone displays a voice interaction interface, such as interface 101 (or the first interaction interface). Interface 101 displays the dialogue content between the phone and the user during a voice conversation.
[0072] refer to Figure 2B This illustrates another scenario for waking up a phone. For example... Figure 2B As shown, the phone displays a desktop interface. The phone receives a long press operation from the user on the power button (or power-on button), and in response to this operation, the phone is woken up. For example, after the phone is woken up, it displays a voice interaction interface, such as interface 101.
[0073] In some embodiments of this application, the smart assistant has corresponding function entry points on the mobile phone, such as the control center and settings application. Users can access the corresponding function entry points and turn the smart assistant's functions on or off by turning preset switches on or off in the mobile phone's human-computer interaction interface.
[0074] refer to Figure 2C This illustrates another scenario for waking up a phone. For example... Figure 2C As shown, in response to a user opening the phone's settings (e.g., clicking the settings app icon on the home screen), the phone displays a settings interface 102. The settings interface 102 includes a smart assistant option 103, which is used to enable or disable voice control functionality. For example, as... Figure 2C As shown, in response to the user's activation of the smart assistant option 103, the phone is woken up. For example, after the phone is woken up, a voice interaction interface is displayed, such as interface 101.
[0075] Optionally, after the phone is woken up, the user can issue a command by inputting voice, and the phone will execute the command. Once the phone has executed a command, or if it does not receive a voice command from the user within a certain period after being woken up (e.g., within 8 seconds), the phone will no longer respond to voice commands. For example, the phone will close the recording channel corresponding to the smart assistant. The user needs to wake up the phone again to issue commands by inputting voice again. In other words, after the phone is woken up, it enters a "short-reception" state, and can respond to voice commands for a short period (e.g., within 8 seconds).
[0076] Optionally, when connected to the network, the phone supports entering a continuous conversation scenario after being woken up, allowing users to engage in ongoing dialogue. After each broadcast, the phone will resume receiving audio without needing to be woken up again, until the user exits the continuous conversation using a command such as "exit."
[0077] For example, when the mobile phone is woken up and enters a continuous conversation scenario, the user can engage in a continuous conversation with the phone on interface 101. Interface 101 displays the content of the continuous conversation between the user and the phone. Optionally, interface 101 includes a prompt message to notify the user that the phone has entered a continuous conversation scenario. For example, the prompt message could be "Welcome."
[0078] Optionally, interface 101 also includes a prompt icon 10a, which is used to notify the user that the phone has entered continuous recording mode. In actual implementation, the prompt icon 10a can be a microphone recording icon, or a "robot recording icon," etc., and is not limited thereto.
[0079] After the phone is woken up, it enters continuous audio reception mode. The user can input voice commands, which the phone then parses and recognizes, executing the corresponding action. For example, the phone can receive voice commands from the user instructing the owner. In response, the phone can retrieve photos, including those of the owner, from the gallery app and generate a video. Here, "owner" can be understood as the primary user of the phone, i.e., the owner of the phone.
[0080] For example, such as Figure 3 As shown in (a), during the user's voice input, the phone displays interface 104, which includes a notification icon 10b to inform the user that the microphone is recording sound. For example, as shown in (a)... Figure 3 As shown in (a), prompt icon 10b displays a transparent circle based on prompt icon 10a. On this basis, the phone receives voice input 1 from the user, which includes instructions for the owner. For example, as... Figure 3 As shown in (a), the voice 1 (or first voice) can be "Generate a video from my photo"; where "I" in voice 1 is an instruction to the device owner. Figure 3 As shown in (b), in response to voice 1, the mobile phone displays interface 105, which includes text content 1 (or first text) that includes text instructing the user. For example, the text content 1 could be "Generate a video from my photos". For instance, the "I" in the aforementioned text content 1 is text instructing the user.
[0081] For example, the mobile phone can recognize the received user input voice 1, thereby identifying "I" as a command instructing the owner, and executing the operation corresponding to that command. Optionally, when voice 1 matches the database, the mobile phone recognizes "I" as a command instructing the owner. For example, if voice 1 includes "I" and the database includes "I," the mobile phone can recognize "I" as a command instructing the owner. It should be noted that the mobile phone can perform semantic understanding on voice 1. Voice 1 does not need to be completely consistent with the database, as long as the semantics are the same or similar. That is, voice 1 matches the database, but it does not have to be completely consistent. For example, in Figure 3 In the scenario shown, the user's voice input 1 includes "I", and the database includes "myself", "myself", and "myself". Since "I" has the same semantic meaning as "myself", "myself", and "myself", the two match. Therefore, the mobile phone can recognize that "I" is a command used to instruct the owner, thus determining that "I" is used to instruct the owner.
[0082] Alternatively, the mobile phone can recognize and parse Voice 1 to identify "I" as a command used to instruct the phone owner, thereby determining that "I" is used to instruct the phone owner. For example, after receiving Voice 1 input by the user, the mobile phone uses Automatic Speech Recognition (ASR) technology to convert the user's speech into text, and then uses Natural Language Understanding (NLU) technology to perform intent recognition on the converted text, thereby identifying "I" as a command used to instruct the phone owner, and thus determining that "I" is used to instruct the phone owner.
[0083] After the phone displays text content 1, as follows Figure 4 As shown in (a), the mobile phone displays interface 106, which includes text content 2 to indicate the response of voice 1. For example, text content 2 could be "Okay, the following materials have been found for you." In addition, the mobile phone displays message 1 on interface 106, which includes a photo of the owner (or a first target photo including a first face). Optionally, interface 106 also includes control 1 (or a second control) and control 2 (or a first control). Control 1 triggers the mobile phone to display all photos related to the owner, and control 2 triggers the mobile phone to generate a video including the owner's photos. For example, control 1 could be called a "View All" control, and control 2 could be called a "Generate Video" control.
[0084] For example, in response to a user's operation on control 1 (or the fifth operation), the phone unfolds to display multiple photos including the owner's photos (e.g., displaying multiple photos including the first person's face).
[0085] For example, in response to a user's action on control 2 (or the first action, such as a click, long press, or swipe), such as... Figure 4 As shown in (b), the mobile phone displays interface 107, which includes a thumbnail of video 1 (or target video). Video 1 is a thumbnail of the video displayed on the mobile phone. Figure 4 (a) shows the generated photo of the phone owner. Optionally, interface 107 includes control 3 for triggering the phone to play video 1. For example, the phone can receive user clicks, long presses, swipes, etc. on control 3; the following explanation will mainly focus on clicks. In response to the user's click on control 3, the phone plays video 1. The content of video 1 is not shown here; the actual implementation shall prevail.
[0086] In summary, the solution adopted in this application allows the phone to enter a continuous dialogue scenario after being woken up, enabling the user to engage in continuous conversation. Based on this, in response to the user's voice input '1', the phone determines that 'I' is used to indicate the owner and retrieves photos including the owner from the gallery application. Subsequently, the phone generates a video related to the owner based on the photos, improving the efficiency of human-computer interaction.
[0087] In some embodiments of this application, the mobile phone can first identify the owner based on photos included in the gallery application. Then, when the mobile phone recognizes a command in the user's voice indicating the owner, the mobile phone responds to the command by retrieving photos including the owner from the gallery application and generating a video based on those photos.
[0088] Optionally, the phone can identify the owner based on photos in the Gallery app while charging and the screen is off; or, the phone can identify the owner based on photos in the Gallery app while the screen is off, etc., without restriction.
[0089] As an example, a mobile phone can identify its owner based on photos in a gallery application in the following three ways. It should be noted that these three methods are merely illustrative examples and do not constitute a limitation of this application; other methods of identifying the owner's photos from a gallery application should also fall within the scope of protection of this application.
[0090] Method 1: The phone's owner is identified by the name set in the gallery application to indicate the owner.
[0091] For example, such as Figure 5As shown, the photo library app can categorize photos stored in the app into multiple photo collections based on different photo types or themes (such as time, portraits, landscapes, food, pets, etc.). Each photo collection includes photos of a specific type or theme. For example, a portrait photo collection includes photos of portraits from the app; a landscape photo collection includes photos of natural landscapes; a food photo collection includes photos of food; and a pet photo collection includes photos of pets. It is understood that different photo types or themes include, but are not limited to, the aforementioned time, portraits, landscapes, food, and pets, and can also include other types or themes, such as architecture, without restriction.
[0092] In the above embodiments, photos in the gallery application are divided into different photo collections according to different photo types or themes, which makes it convenient for users to view multiple photos of different types or themes without having to browse through all the photos, making the user operation more convenient and improving the user experience.
[0093] For example, a collection of portrait photos, such as... Figure 6 As shown, by performing facial clustering algorithms and other analyses on all photos in the image library, different photo collections of different people can be generated based on the different facial ID analysis results, such as the photo collection of person 1, the photo collection of person 2, the photo collection of person 3, etc., which will not be elaborated further.
[0094] For example, a mobile phone can accept user input to assign corresponding names to different collections of photos of people. The phone can also use a relationship analysis algorithm to assign names to different collections of photos of different people. For example... Figure 6 As shown, the photo collection for Person 1 on the phone is named "Me", the photo collection for Person 2 on the phone is named "Son", and the photo collection for Person 3 on the phone is named "Wife".
[0095] It should be noted that, Figure 6 The names corresponding to the photo collections of different people shown are merely examples and do not constitute a limitation of this application. Of course, the mobile phone can also set other names for the photo collections of different people. For example, taking the photo collection of person 1 as an example, the mobile phone can also set the name of the photo collection of person 1 as "myself" or "myself", etc., without limitation.
[0096] Understandably, names such as "Me," "Myself," or "This Person" can be pre-configured to indicate the phone owner. Therefore, based on the names "Me," "Myself," or "This Person" set in the gallery application, the phone can determine the photo collection corresponding to that name, which is the owner's photo collection, meaning the face ID corresponding to that photo collection is the owner's.
[0097] Optionally, after identifying the owner using Method 1, the phone can also record the owner's relevant information. For example, this information includes: owner A, face ID-A, confidence level A (or understood as the first level of confidence), and photo source A. Confidence level can be understood as reliability, i.e., the accuracy of the identified owner (or it can be understood as the priority of Method 1; the higher the confidence level, the higher the priority). For example, in Method 1, since the owner is identified by the phone based on the name set in the gallery application to indicate the owner, the confidence level A can be set to a value that is greater than the confidence level determined by other methods (e.g., the confidence level determined by other methods is definitely less than 100%, i.e., the confidence level is set to 1, meaning that Method 1 has the highest priority in identifying the owner. Furthermore, photo source A can be represented as: the name set in the gallery application to indicate the owner.
[0098] Method 2: The phone identifies the owner based on selfies (or front-facing camera photos) taken using the phone's gallery app. Selfies refer to photos taken with the phone's front-facing camera.
[0099] In some embodiments of this application, if the number of selfies corresponding to Face ID-1 exceeds a preset threshold (or selfie count threshold), then the face corresponding to Face ID-1 is identified as the device owner. Alternatively, if the number of selfies corresponding to Face ID-1 is the largest among all selfies included in the total number of faces, then the face corresponding to Face ID-1 is identified as the device owner.
[0100] The aforementioned preset threshold can be adjusted based on the number of photos in the gallery application, such as taking X% of the number of photos in the gallery application. The preset threshold can also be set according to other actual situations and is not limited. For example, the preset threshold can be set to other suitable values such as 5 or 10, without restriction.
[0101] Alternatively, in some other embodiments of this application, if the gallery application includes selfie photos of multiple faces, in order to improve the accuracy of identifying the owner, the mobile phone can identify the owner based on the facial information in the selfie photos of multiple faces.
[0102] For example, facial information may include selfie frequency, selfie time, face ID (or tag ID, or face identifier), face gender, face age, etc.; or, facial information may also include selfie time, etc., without limitation.
[0103] Optional, such as Figure 7As shown, determining the phone owner based on facial information may include the following steps: Step 1: The phone queries selfies in the gallery application and retrieves facial information corresponding to each face in the selfie photos from the selfie table. For example, the selfie table stores facial information for each face in the selfie photos; the selfie table is obtained by the phone's gallery application analyzing photos when the phone screen is off. Step 2: The phone determines the selfie frequency for each face ID based on the face ID in the facial information. The selfie frequency includes the number of selfies and / or the selfie frequency itself. Optionally, after performing steps 1 and 2, the process may further include step 3: The phone filters out face IDs that do not meet the criteria based on the face gender and age in the facial information (i.e., selecting the first group of faces).
[0104] For example, the mobile phone reads the user's gender and age from a user profiling table. This user profiling table is learned by the mobile phone based on records of the user's use of applications (APPs) installed on the phone, and stores the user's predicted gender and age / / age range. In some embodiments of this application, if the gender of the face in the above-mentioned facial information does not match the gender in the user profiling table, the mobile phone filters out the corresponding face ID. Alternatively, the mobile phone calculates the average age of each face ID; if the difference between the average age and the age in the user profiling table exceeds a preset value, the mobile phone filters out the corresponding face ID. The preset value is not specifically limited and is determined by actual settings. For example, the preset value can be 20, etc., and is not limited.
[0105] Optionally, assuming that in the photo library application, face ID1 includes ages 1, 2, 3, ..., n in the selfie photos, the average age of face ID1 can be expressed as (age1 + age2 + age3 + ... + agen) / n, where n is a positive integer.
[0106] Alternatively, in some other embodiments of this application, if the age of the face corresponding to the face ID differs from the age in the user habit table by more than a preset threshold or is not within the age range in the user habit table, the mobile phone filters out the face ID.
[0107] Step 4: The phone determines the owner based on the frequency of selfies taken using the Face ID. For example, assume the Face ID includes Face ID1 and Face ID2. The phone calculates the frequency percentage of Face ID1 and Face ID2 respectively, using this percentage as the confidence level (or a second confidence level). If the confidence level is greater than a confidence threshold, the Face ID corresponding to that confidence level is identified as the owner.
[0108] For example, in step 2 above, the mobile phone determines the selfie frequency of each Face ID. For instance, if Face ID1 has a selfie frequency of 10 and Face ID2 has a selfie frequency of 20, then the frequency proportion of Face ID1 is 10 / (10+20) = 0.33, meaning the confidence level of Face ID1 is 0.33; correspondingly, the frequency proportion of Face ID2 is 20 / (10+20) = 0.67, meaning the confidence level of Face ID2 is 0.67.
[0109] Optionally, if the confidence scores of both Face ID1 and Face ID2 are greater than the confidence score threshold, the phone will identify Face ID1 and Face ID2 as the owner of the phone.
[0110] To ensure the accuracy of the phone's identified owner, the confidence level can be set relatively high, for example, to 0.4. Of course, the confidence level can also be greater than 0.4; there are no restrictions.
[0111] Optionally, after the phone identifies the owner using Method 2 above, it can also record the owner's relevant information. For example, the owner's relevant information includes: owner B, face ID-B, confidence level B, and photo source B. In Method 2 above, owner B can be the user corresponding to face ID2, face ID-B can be face ID2, confidence level B can be 0.67, and photo source B can be a "selfie".
[0112] Understandably, if the phone identifies two owners using the above method two, the relevant information recorded by the phone for the owners includes: owner B, face ID-B, confidence level B, and photo source B; owner C, face ID-C, confidence level C, and photo source C.
[0113] Method 3: The phone identifies the owner based on the location and time of the photos taken in the gallery app, as well as the user's home and office locations.
[0114] For example, such as Figure 8 As shown, the process of determining the phone owner based on method three may include the following steps: Step 1: The phone reads photos from the gallery application and obtains the shooting location and time of the photo corresponding to each Face ID in the Face ID database. Step 2: The phone reads the first location (which can be understood as the location of home) and the second location (which can be understood as the location of the workplace) determined by the phone. The first and second locations are learned by the phone based on the user's movement trajectory.
[0115] Step 3: The phone calculates the distance between the first location and the second location, and determines whether the distance between the first location and the second location is greater than a preset distance. For example, if the distance between the first location and the second location is greater than the preset distance, the phone continues to Step 4. Conversely, if the distance between the first location and the second location is less than the preset distance, the phone ends the process of determining the owner. The preset distance is not specifically limited and can be set according to actual conditions. For example, the preset distance can be 3 kilometers (km), or it can be set to M times the radius of the area corresponding to the first location; or it can be set to M times the radius of the area corresponding to the second location, without limitation. Here, M is a positive integer greater than 1.
[0116] It should be noted that the distance between the first and second locations is usually quite far, generally greater than 3km. Therefore, to avoid errors in the first and second locations determined by the phone, the phone can calculate the distance between them. If the distance is greater than a preset distance, it means the determined locations are relatively accurate, and the phone can continue with the following steps. Conversely, if the distance is less than the preset distance, it means there is an error in the determined locations, and the phone can terminate the process of identifying the owner.
[0117] Step 4: The mobile phone determines the target face ID (or target face identifier) corresponding to the photos taken simultaneously at the first location and the second location. For example, based on the shooting location of the photos corresponding to each face ID, the mobile phone determines the face ID whose shooting location is the first location. Based on this, the mobile phone determines whether a photo was taken at the second location based on the face ID taken at the first location. If the face ID taken at the first location was also taken at the second location, the mobile phone uses that face ID as the target face ID (which may include multiple face IDs), meaning that the target face ID has been photographed at both the first and second locations. In other words, the mobile phone, through the method described in Step 4 above, can find the face ID (i.e., the target face ID) corresponding to the photos taken simultaneously at the first and second locations.
[0118] Optionally, if the distance between the shooting location and the first location is less than a first preset distance, the phone determines the shooting location to be the first location. Similarly, if the distance between the shooting location and the second location is less than a second preset distance, the phone determines the shooting location to be the second location. The first and second preset distances are not limited and are determined by actual settings. The first and second preset distances can be the same or different; for example, the first preset distance can be 200 meters and the second preset distance can be 500 meters.
[0119] Step 5: The mobile phone determines the first frequency of photos taken with the target face ID at the first location and the second frequency of photos taken with the target face ID at the second location. If both the first and second frequencies are greater than a preset frequency, then the face ID that meets the above conditions is identified as the owner's face ID. The preset frequency is not specifically limited and is determined by actual settings. For example, the preset frequency can be 2.
[0120] For example, suppose the target face ID includes face ID1. Face ID1 takes a photo at a first location at a first frequency of 10 times and at a second location at a second frequency of 2 times. It can be seen that both the first and second frequencies of face ID1 are greater than the preset frequency. Therefore, the mobile phone determines that face ID1 is the owner of the phone.
[0121] For example, suppose the target face ID includes face ID1 and face ID2. Face ID1 takes a photo at a first location at a first frequency of 10 and at a second location at a second frequency of 2. Face ID2 takes a photo at a first location at a first frequency of 5 and at a second location at a second frequency of 5. It can be seen that the first and second frequencies corresponding to face ID1 and face ID2 are both greater than the preset frequencies. Therefore, the mobile phone determines the owner based on face ID1 and face ID2.
[0122] In some embodiments of this application, when the number of face IDs whose photos have been taken at both the first and second locations is less than or equal to a threshold of 1, the mobile phone determines the owner based on the face's gender. For example, if the gender of the face corresponding to face ID1 is the same as the gender of the user read by the mobile phone from the user habit table, then the mobile phone determines that face ID1 is the owner. The threshold of 1 is not limited and is determined by its actual setting. For example, the threshold of 1 can be 2.
[0123] Alternatively, in some other embodiments of this application, when the number of face IDs whose photos have been taken at both the first and second locations is less than or equal to a threshold of 1, the phone determines the owner based on the face age. For example, if the face age corresponding to face ID1 is close to the user's age read by the phone from the user habit table, the phone determines face ID1 as the owner. Here, "close to the same age" can be understood as: the difference between the face age corresponding to face ID1 and the user's age is less than a preset difference; or, the user's age is within a preset range, and the face age corresponding to face ID1 is also within that preset range.
[0124] Optionally, if the number of face IDs whose photos have been taken at both the first and second locations exceeds threshold 2, the phone may abandon the process of identifying the phone owner. In this embodiment, if the phone detects a large number of face IDs whose photos have been taken at both the first and second locations, it indicates that the accuracy of the phone's identification of the phone owner is insufficient. Therefore, in this case, the phone may abandon the process of identifying the phone owner using the above-mentioned method three, thereby ensuring the accuracy of the phone's identification of the phone owner. The threshold 2 is not limited and can be set according to actual conditions. For example, threshold 2 can be 3.
[0125] Optionally, after the mobile phone identifies the owner through method three described above, it can also record the owner's relevant information. For example, the owner's relevant information includes: owner D, face ID-D, confidence level D, and photo source D. For example, in this method, the confidence level D (or understood as the third confidence level) can be defaulted to a value less than 1 (or a first preset value) (e.g., 0.96), and the specific value of the confidence level is not limited. The photo source D can be "home and company". In some embodiments, the mobile phone sets the third confidence level corresponding to the owner identified through method three to a second preset value, which is less than the first preset value.
[0126] In conclusion, combining the user's home and workplace locations to determine the phone owner allows for a wider coverage of owner identification technology.
[0127] The following provides a method for a mobile phone to determine the location of a user's home and workplace. It should be understood that the embodiments described below are merely illustrative examples of this application and do not constitute a limitation of this application.
[0128] For example, a mobile phone can calculate the user's stop points along their movement trajectory based on signaling data. Then, the phone can mark the user's daily residence and workplace based on this stop point data. Finally, based on the user's daily residence and workplace, the phone can determine the location of the user's home and company.
[0129] In some embodiments of this application, the mobile phone can calculate stop point data based on the trajectory point data of the user moving along the movement trajectory. The trajectory point data is data containing both time and spatial information. When a user carries a mobile phone while traveling, the user's trajectory point data is essentially the mobile phone's trajectory point data; the two are equivalent. The user's trajectory point data can be determined by the spatial location of the mobile phone at different points in time.
[0130] For example, trajectory point data can be obtained through GPS positioning or through information exchange between the mobile phone and the base station. For instance, the mobile phone records trajectory point data when it is turned on or off, makes calls, sends text messages, updates its location, or switches base stations. It can also be obtained through other means without restriction.
[0131] Figure 9 A schematic diagram of a movement trajectory is shown. For example, the movement trajectory could be... Figure 9 The route shown is between T1 and T22; where T1-T22 in this movement trajectory can be understood as trajectory points. For example, taking a mobile phone acquiring trajectory point data via GPS, the phone periodically acquires latitude and longitude information at time intervals. For instance, with a time interval of 1 minute, the phone acquires its current geographical location's latitude and longitude every minute, obtaining the user's daily trajectory point data.
[0132] For example, trajectory point data can be as shown in Table 1 below. Each trajectory point data includes a time point and a spatial point (such as latitude and longitude).
[0133] Table 1
[0134] trajectory points Time point spatial point T1 2022.07.01 00:00 116.671°,39.856° T2 2022.07.01 00:01 116.672°,39.857° T3 2022.07.01 00:02 116.672°,39.857° … … … T22 2022.07.01 23:59 116.78°,39.157°
[0135] For example, the mobile phone determines stop point data based on the acquired trajectory point data (as shown in Table 1). Stop points indicate where the user stays within a certain time period and their location within that time period. That is, a stop point includes both a time period and a spatial region (also called a spatial range). The spatial region can be understood as the area formed by a certain distance from the latitude and longitude information corresponding to the target trajectory point as its center. The target trajectory point can be the trajectory point located at the center of multiple trajectory points. For example, as shown... Figure 9 As shown, trajectory points T3-T8 form a resting point (which can be represented as A1), and trajectory points T12-T19 form a resting point (which can be represented as A2).
[0136] For example, the stop point data determined by the mobile phone can be as shown in Table 2 below. Each stop point data point includes a time period and a spatial region (the latitude and longitude information and spatial radius of the target trajectory point).
[0137] Table 2
[0138] Stop Time period spatial region A1 2022.07.01 00:00-2022.07.01 06:00 (116.600°, 40°), radius: 100 meters A2 2022.07.01 07:00-2022.07.01 18:00 (116.700°, 40°), radius: 130 meters A3 2022.07.01 13:00-2022.07.01 14:00 (116.601°, 40°), radius: 100 meters … … … An 2022.07.01 18:30-2022.07.01 23:59 (116.602°, 40°), radius: 90 meters
[0139] Based on Table 2 above, it can be seen that the user stayed at stop point A1 from 00:00 to 06:00 on July 1, 2022. Therefore, the mobile phone can mark the spatial area corresponding to stop point A1 as the user's residence. Similarly, the user stayed at stop point A2 from 07:00 to 18:00 on July 1, 2022. Therefore, the mobile phone can mark the spatial area corresponding to stop point A2 as the user's workplace.
[0140] In this embodiment of the application, the mobile phone can mark the user's daily residence and workplace using the above method, and determine the user's home and company locations based on the daily marked residence and workplace locations. Optionally, the mobile phone can save the determined home and company locations each time, so as to identify the phone owner later based on the saved home and company locations.
[0141] For example, such as Figure 10 As shown, the mobile phone displays interface 201 (or the second interactive interface), which includes the user's home and work locations. These locations are obtained by the mobile phone through automated learning. For example, the user's home and work locations are learned by the mobile phone using the methods described above.
[0142] Based on the above, in this embodiment, the mobile phone can determine the owner through methods one, two, and three, and record the owner's relevant information after determining the owner under each method. For example, the mobile phone can store the correspondence between the owner determined by different methods and the owner's relevant information in the form of a table or array (or a preset storage table). For example, this correspondence can be shown in Table 3.
[0143] Table 3
[0144]
[0145] In this embodiment, since different methods of determining the phone owner can lead to errors, the phone owner determined by different methods may be the same or different. Optionally, since Method 1 uses the phone's gallery application to indicate the owner's name, it can be considered that the owner determined by Method 1 is more accurate, meaning that Method 1 has the highest priority in determining the owner.
[0146] In some embodiments of this application, after receiving voice input from a user, the mobile phone, in response to the voice, determines an instruction for instructing the owner. Then, in response to the instruction for instructing the owner, the mobile phone retrieves a photo including the owner from the gallery application based on a confidence level. Since the mobile phone determines that owner A has the highest priority using method one and sets the confidence level to 1 (first preset value), the mobile phone's gallery application returns the determined face ID of owner A as the owner's face ID. For example, in conjunction with the above... Figure 4 As shown in (a), in message 1 displayed on interface 106, the owner's photo is determined by the phone based on the name set in the gallery application to indicate the owner.
[0147] In other embodiments of this application, upon receiving a user's operation (or second operation), the mobile phone deletes the name (i.e., owner tag) set in the gallery application to indicate the owner. It should be understood that since the owner tag is absent, when the owner is identified in the next photo in the gallery application, the owner cannot be determined using method one, therefore there is no owner with a confidence level of 1. Based on this, when the mobile phone retrieves photos including the owner from the gallery application according to the confidence level, it can display photos of owners with confidence levels exceeding a threshold for the user to choose from.
[0148] The mobile phone receives voice input from the user (or second voice). In response to this voice, the mobile phone determines the instruction used to instruct the owner. Then, in response to the instruction used to instruct the owner, the mobile phone retrieves photos including the owner from the gallery application. Referring to Table 3 above, assuming the mobile phone identifies the owner (determined using methods two and three) through photos in the gallery application, the owners include: owner B, owner C, and owner D, with confidence levels of 0.33, 0.67, and 0.96, respectively. Optionally, the mobile phone can display owner B, owner C, and owner D as all identified owners. As another implementation, the mobile phone can also display owner B and owner D, whose confidence levels exceed a confidence threshold (or exceed a preset threshold, such as 0.4), as identified owners. For example, the mobile phone receives voice input 2 (or second voice) from the user, which includes an instruction used to instruct the owner. For example, voice 2 could be "Generate a video from my photos"; where "I" in voice 2 is the instruction instructing the owner. Responding to voice 2, such as Figure 11 As shown in (a), the mobile phone displays interface 202 (or can be understood as the first interactive interface). Interface 202 includes text content 3 (or second text), which includes text instructing the user. For example, the text content 3 could be "Generate a video from my photos," where "I" in the text content 3 is the text instructing the user.
[0149] After the phone displays text content 2, as follows Figure 11As shown in (a), interface 202 also includes text content 4, which is used to indicate the response of voice 2. For example, text content 4 could be "Okay, the following materials have been found for you." In addition, the phone also displays a photo of owner B (or, as understood, a second target photo including a second face) and a photo of owner D (or, as understood, a third target photo including a third face) on interface 202. It should be understood that the number of owner photos displayed on interface 202 is related to the number of owners determined according to methods two and three and the confidence level setting; this solution does not impose any restrictions on this.
[0150] In some embodiments, such as Figure 11 As shown in (a), the phone defaults to selecting photos of owner B and owner D. For example, in response to a user action (or a third action), such as... Figure 11 As shown in (b), the phone unchecks the photo of the owner B. Then, in response to the user's click on the "Generate Video" control (or the first control) (or the fourth operation), the phone displays a thumbnail of Video 2, which is a video generated by the phone based on a photo of the owner D (i.e., a photo including a third person's face) in the gallery application.
[0151] It should be noted that, for Figure 11 For examples of other content on the various interfaces shown, please refer to the above. Figure 3 and Figure 4 As shown, they will not be elaborated on here.
[0152] The above embodiments mainly describe the solutions of the embodiments of this application in detail with reference to the hardware structure of the mobile phone (i.e., electronic device 100). The specific implementation process of the embodiments of this application will be described in detail below with reference to the software framework of electronic device 100.
[0153] Taking a mobile phone as an example again, the software system of a mobile phone can, for instance, adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses a layered architecture as an example to illustrate the software structure of a mobile phone.
[0154] Figure 12 This is a software structure block diagram of a mobile phone provided in an embodiment of this application. A layered architecture can divide the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the layered architecture, from top to bottom, consists of the application layer (application layer), the application framework layer (framework layer), the hardware abstraction layer (HAL), and the kernel layer (also known as the driver layer).
[0155] Application layer: This can include a series of application packages. For example, the application layer can include gallery applications and voice interaction applications. Of course, the application layer can also include applications such as camera, calendar, call, and map, which will not be elaborated further.
[0156] The gallery app is used to save photos and videos taken with the phone. Once launched, the voice interaction app allows users to interact with the phone via voice, performing various operations such as inputting commands, querying information, and voice chat.
[0157] The framework layer provides the application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. For example... Figure 13 As shown, the framework layer may include input management services, window management services, activity management services, etc. Optionally, in this embodiment, the framework layer also includes a basic judgment module for determining the device owner based on photos in the gallery application.
[0158] For example, the input management service is mainly responsible for input event listening, input event parsing, and input event distribution.
[0159] The window management service manages window applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots. The activity management service is primarily responsible for managing activities, including starting, switching, and scheduling components within the system, as well as managing and scheduling applications.
[0160] The hardware abstraction layer can include multiple library modules, such as hardware compositors (Hwcomposer, HWC), camera library modules, etc. The mobile phone's operating system can load the corresponding library modules for the device hardware, thereby enabling the application framework layer to access the device hardware.
[0161] The kernel layer is the layer between hardware and software. The kernel layer can include, but is not limited to, camera drivers, display drivers, audio drivers, etc.
[0162] Combination Figure 12 ,like Figure 13 The diagram illustrates a flowchart of a photo processing method. For example, the method may include the following steps.
[0163] Step 301: The phone's gallery application determines the Face ID corresponding to the photo used to indicate the owner's name.
[0164] Step 302: The phone's gallery application sends the corresponding face ID to the basic judgment module.
[0165] Step 303: The phone's basic judgment module determines the owner of the phone based on the corresponding face ID.
[0166] For example, steps 301-303 above are for the mobile phone to determine the owner through method one. The specific implementation method can be referred to the above embodiment, and will not be repeated here.
[0167] Step 304: The phone's gallery app obtains the facial information corresponding to the selfie photo.
[0168] The facial information can include selfie frequency, selfie time, facial ID, facial gender, facial age, etc., without any restrictions.
[0169] For example, a phone's gallery app can analyze photos stored in the gallery and generate a selfie list when the phone is screen-off (or charging with the screen off). Based on this, the phone can obtain facial information from the selfie list.
[0170] Step 305: The phone's gallery application sends the facial information corresponding to the selfie photo to the basic judgment module.
[0171] Step 306: The phone's basic judgment module determines the phone owner based on the facial information corresponding to the selfie photo.
[0172] For example, steps 304-306 above are for the mobile phone to determine the owner through method two. The specific implementation method can be referred to the above embodiment, and will not be repeated here.
[0173] Step 307: The phone's gallery application reads and retrieves photos from the gallery application, and obtains the shooting location and shooting time corresponding to each face ID in all face IDs.
[0174] Step 308: The phone's gallery application sends the shooting location and shooting time corresponding to each face ID to the basic judgment module.
[0175] Step 309: The phone's basic judgment module reads the first location and the second location.
[0176] Step 310: The phone's basic judgment module filters out face IDs that have been photographed at both the first and second locations based on the shooting location and shooting time corresponding to each face ID, in order to identify the phone owner.
[0177] For example, steps 307-310 above are for the mobile phone to determine the owner through method three. The specific implementation method can be referred to the above embodiment, and will not be repeated here.
[0178] It should be noted that the above-mentioned methods one to three are three parallel implementation methods. The order of methods one to three in this application embodiment is not limited, and the actual implementation shall prevail.
[0179] Step 311: After the mobile phone's voice interaction application determines the command to instruct the owner, it queries the basic judgment module for the owner's face ID. In this step, the owner's face ID can be determined based on the face ID and its confidence level. Specifically, the basic judgment module returns the face ID with the highest confidence level as the owner's face ID, or it can return face IDs with confidence levels exceeding a preset threshold as the owner's face ID. Optionally, the basic judgment module can determine how to return the owner's face ID based on the highest confidence level value. Specifically, it determines whether the highest confidence level value is a first preset value (e.g., the confidence level of the face ID determined in method one is 1). When the highest confidence level value is the first preset value, only face IDs with confidence levels of the first preset value can be returned as the owner's face ID; when the highest confidence level value is not the first preset value, all face IDs with confidence levels exceeding the preset threshold are returned as the owner's face ID. Step 312: The phone's voice interaction application sends an instruction message to the gallery application. The instruction message includes the owner's face ID, which is used to instruct the gallery application to filter photos that include the owner based on the owner's face ID.
[0180] Step 313: The phone's gallery app returns photos, including those of the phone owner, to the voice interaction app.
[0181] Step 314: The voice interaction application displays photos including the user's picture and generates a video based on those photos. Specifically, the voice interaction application selects one photo from the photos including the user's picture according to preset rules as the representative photo (the first target photo). In response to the user's action of viewing all photos, multiple photos of the user can be displayed. The preset rules could be based on the position of the user's face in the photo, the photo's aesthetic rating, the photo's capture time, etc., but this solution does not impose any restrictions.
[0182] In summary, by adopting the solution of this application embodiment, the basic judgment module of the mobile phone can determine the owner in three ways. The voice interaction application of the mobile phone can query the owner's face ID from the basic judgment module, and obtain the owner's photo from the gallery application based on the owner's face ID to generate a video, thereby improving the efficiency of human-computer interaction.
[0183] It should be noted that the contents described in the various embodiments of this application can be used to explain the technical solutions in other embodiments of this application. The technical features described in each embodiment can also be applied in other embodiments, and new solutions can be formed by combining the technical features in other embodiments. This application only provides an exemplary list of several embodiments for illustration and does not mean that this application is limited thereto.
[0184] This application provides an electronic device that may include a memory and one or more processors. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device performs the various functions or steps performed by the mobile phone in the above embodiment. The structure of the electronic device can be referred to the above. Figure 1 The structure of the electronic device 100 shown.
[0185] This application also provides a chip system for use in electronic devices. For example... Figure 14 As shown, the chip system 1100 includes at least one processor 1101 and at least one interface circuit 1102. The processor 1101 can be one of the embodiments described above. Figure 1 The processor 110 is shown. Based on this, the interface circuit 1102 can be, for example, an interface circuit between the processor 110 and external memory; or an interface circuit between the processor 110 and internal memory.
[0186] The processor 1101 and interface circuit 1102 described above can be interconnected via a line. For example, interface circuit 1102 can be used to receive signals from other devices (e.g., the memory of electronic device 100). As another example, interface circuit 1102 can be used to send signals to other devices (e.g., processor 1101). Exemplarily, interface circuit 1102 can read instructions stored in memory and send those instructions to processor 1101. When the instructions are executed by processor 1101, the electronic device can perform the various functions or steps performed by the mobile phone in the above embodiments. Of course, the chip system may also include other discrete components, and this application embodiment does not specifically limit this.
[0187] This application also provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.
[0188] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.
[0189] It should be noted that the terms "first" and "second," etc., in the embodiments and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more.
[0190] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0191] It should be understood that in this application, "at least one (item)" means one or more. "More than one" means two or more. "At least two (items)" means two or three or more. "And / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple. Both "...when" and "if" indicate that a corresponding action will be taken under certain objective circumstances. They are not time limits, nor do they require a judgment action to be taken when the action is taken, nor do they imply any other limitations.
[0192] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0193] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules according to the system, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0194] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0195] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0196] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional unit.
[0197] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0198] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A photo processing method, characterized in that, Applied to electronic devices, including: The electronic device receives a first voice input from a user and displays a first text corresponding to the first voice on a first interactive interface; the first text includes information indicating the owner of the device. In response to the first voice, a first target photo and a first control are displayed on the first interactive interface; wherein, the first target photo contains a first face, and the first face has a device owner tag set in the gallery application of the electronic device; the first voice is associated with the device owner tag and the first control; the device owner tag is generated based on selfie frequency, location, and user habits; wherein, the device owner tag generated based on location is determined based on the shooting location of the first target photo and the location of the user's home and company; the device owner tag generated based on user habits is determined based on the user's gender and age read from the user habit table; In response to a first operation by the user on a first control, the electronic device displays a thumbnail of a target video; the target video is generated by the electronic device based on a first target photo including the first face in a gallery application.
2. The method according to claim 1, characterized in that, The method further includes: The electronic device receives a second operation from the user to delete the main tag in the gallery application; In response to a second voice input by the user, the electronic device displays a second text corresponding to the second voice on the first interactive interface; the second text includes information indicating the owner of the device. In response to the second voice, the electronic device displays a second target photo, a third target photo, and the first control on the first interactive interface; wherein the second target photo contains a second face, and the third target photo contains a third face; In response to the user's third action, the electronic device unchecks the second target photo; In response to a fourth operation by the user on the first control, the electronic device displays a thumbnail of the target video; the target video is generated by the electronic device based on a third target photo including the third face in a gallery application.
3. The method according to claim 2, characterized in that, The second target photo is determined by the electronic device based on a front-facing photo containing the second face from all photos included in the gallery application.
4. The method according to claim 3, characterized in that, The number of front-facing photos containing the second face is the largest among all front-facing photos containing the face; or, the number of front-facing photos containing the second face is greater than or equal to the selfie count threshold.
5. The method according to any one of claims 2-4, characterized in that, The method further includes: The electronic device displays a second interactive interface, which includes a first location and a second location. The first location is the home location learned by the electronic device, the second location is the company location learned by the electronic device, and the photo containing the third person's face has appeared at both the first and second locations.
6. The method according to any one of claims 1-4, characterized in that, The first interactive interface also displays a second control; the method further includes: In response to a fifth operation input by the user to the second control, the electronic device unfolds to display multiple photos including the first face.
7. The method according to claim 2, characterized in that, Before displaying the first target image, the method further includes: The electronic device determines that the first voice message includes an instruction to the device owner, and queries the device owner's face ID; The electronic device determines the owner's face ID as the first face based on the confidence level corresponding to different face IDs; the confidence level of the first face is a first confidence level, which is greater than the confidence level of other faces; The electronic device filters photos of the first face to determine the first target photo.
8. The method according to claim 7, characterized in that, Before displaying the second target photo and the third target photo on the first interactive interface, the method further includes: The electronic device determines that the second voice includes an instruction to the device owner to query the device owner's face ID; The electronic device determines the owner's face ID as the second face and the third face based on the confidence level corresponding to different face IDs; The electronic device filters photos of the second face and the third face to determine the second target photo and the third target photo.
9. The method according to claim 7 or 8, characterized in that, The method further includes the following steps before the electronic device receives the first voice input from the user; The electronic device generates multiple photo collections of different people based on different face ID analysis results; The electronic device sets a main tag for the collection of photos containing the first face; The electronic device determines that the first confidence level of the first face is a first preset value.
10. The method according to claim 8, characterized in that, Before the electronic device receives the second voice input from the user, the method further includes: In the application of the electronic device to acquire the gallery, the gender and age of each face among the multiple faces included in all the front-facing photos; The electronic device filters out a first group of faces based on the gender and age of each face among the plurality of faces; the gender of each face in the first group of faces is the same as the gender of the second face, and the difference between the age of each face and the age of the second face is within a preset range; The electronic device uses the ratio of the number of front-facing photos of the second face to the number of front-facing photos of the first group of faces as the second confidence level; the second confidence level is less than the first preset value. The electronic device reads photos from the gallery application and obtains the shooting location and shooting time of the photos corresponding to each face ID in all face IDs; The electronic device determines the target face ID corresponding to the photo taken simultaneously at the first location and the second location based on the shooting location and shooting time of the photo corresponding to each face ID in all face IDs; the target face ID includes the third face; The electronic device sets the third confidence level corresponding to the third face to a second preset value, which is less than the first preset value.
11. The method according to claim 10, characterized in that, The method further includes: The electronic device determines whether the value of the highest confidence level among the confidence levels corresponding to different face IDs is the first preset value; When the highest confidence value is the first preset value, the face ID with the confidence value of the first preset value is returned as the face ID of the device owner; When the highest confidence value is not the first preset value, all face IDs with confidence values exceeding the preset threshold are returned as the face IDs of the device owner.
12. The method according to claim 10, characterized in that, The electronic device determines the target face ID corresponding to the photo taken simultaneously at the first location and the second location based on the shooting location and shooting time of the photo corresponding to each face ID in all face IDs, including: The electronic device determines, based on the shooting location of the photo corresponding to each face ID in all face IDs, that the distance between the shooting location and the first location is less than or equal to a first preset distance, and determines that the distance between the shooting location and the second location is less than or equal to a second preset distance; The electronic device determines the frequency of photos taken at the first location for each face ID in all face IDs, and the frequency of photos taken at the second location; The electronic device filters out the target face IDs corresponding to photos taken simultaneously at the first location and the second location based on the shooting time; The frequency of photos taken of the target face ID at the first location is greater than a preset frequency, and the frequency of photos taken at the second location is greater than the preset frequency.
13. The method according to claim 10 or 12, characterized in that, Before the electronic device determines the target face ID corresponding to the photos taken simultaneously at the first location and the second location based on the shooting location and shooting time of the photos corresponding to each face ID in the total face IDs, the method further includes: The electronic device determines that the distance between the first location and the second location is greater than or equal to a preset distance.
14. The method according to claim 13, characterized in that, Before the electronic device receives the second voice input from the user, the method further includes: The electronic device determines the age range of the device owner based on the user's usage habits; The method further includes: If the number of target face IDs is less than or equal to a preset number, the electronic device will use the corresponding face ID within the age range of the target face IDs as the face ID of the device owner.
15. An electronic device, characterized in that, include: Memory and one or more processors, display screen; The memory stores computer program code, which includes computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-14.
16. A chip system, characterized in that, include: At least one processor and an interface; The interface is used to receive instructions and transmit them to the at least one processor; The at least one processor executes the instructions to cause the electronic device to perform the method as described in any one of claims 1-14.
17. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-14.
18. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-14.
Citation Information
Patent Citations
User shooting method, device and equipment
CN108052883A
Video generation method and device
CN116437163A