A social relationship-based video generation method and electronic device

By judging and merging the facial similarity of user IDs in electronic devices, the problems of repetitive calculation and high power consumption in video generation in existing technologies are solved, achieving more efficient video generation and a better human-computer interaction experience.

CN120877341BActive Publication Date: 2026-04-14HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-04-19
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies for generating videos in electronic devices cannot effectively utilize social relationships, leading to problems such as redundant computation and high power consumption.

Method used

By determining whether user facial similarity is cached in electronic devices and predicting whether user IDs belong to the same person based on similarity, user IDs are merged to generate videos, reducing redundant calculations and lowering power consumption.

Benefits of technology

It improves the quality and performance of video generation, reduces data processing volume, lowers power consumption, and enhances human-computer interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877341B_ABST
    Figure CN120877341B_ABST
Patent Text Reader

Abstract

The application provides a video generation method based on social relations and an electronic device, relates to the technical field of terminals, and can make the overall scheme of generating a video based on social relations by the electronic device better and have better performance; the method comprises the following steps: the electronic device judges whether a first face similarity corresponding to a first user ID and a second user ID in a plurality of user IDs is cached in the electronic device; if the current judgment result indicates that the first face similarity is currently cached in the electronic device, the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the faces of the same person according to the first face similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to a video generation method and electronic device based on social relationships. Background Technology

[0002] With the development of terminal technology and the increasing maturity of speech recognition technology, voice input has become increasingly important due to its high naturalness and effectiveness in interaction. Voice interaction applications within electronic devices have also become a frequently used function. Users can interact with electronic devices (such as mobile phones, tablets, and smartwatches) via voice to complete various operations such as command input, information retrieval, and voice chat. For example, electronic devices can respond to a user's voice and generate videos from photos in their gallery applications, resulting in a better user experience and higher human-computer interaction efficiency. Summary of the Invention

[0003] This application provides a video generation method and electronic device based on social relationships, which can make the overall solution of electronic device generating videos based on social relationships more effective and perform better.

[0004] The embodiments of this application adopt the following technical solutions:

[0005] Firstly, a video generation method based on social relationships is provided. This method is applied to an electronic device, which includes multiple photos, each photo corresponding to a user ID, and each user ID corresponding to a face.

[0006] The method includes: the electronic device determining whether it has cached first face similarity scores corresponding to a first user ID and a second user ID from among multiple user IDs. If the determination result indicates that the electronic device currently has cached first face similarity scores, the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's face based on the first face similarity scores.

[0007] If the faces corresponding to the first user ID and the second user ID are the same person's face, the electronic device assigns a first fusion ID and determines the first social relationship between the device owner and the user corresponding to the first fusion ID. After the electronic device determines the first social relationship between the device owner and the user corresponding to the first fusion ID, the electronic device receives the first voice input from the user (including information instructing the device owner and information about the existence of a first social relationship with the device owner). In response to the first voice, the electronic device displays a first photo (containing the user corresponding to the first fusion ID) and a first control on the first interactive interface. In response to the user's first operation on the first control, the electronic device displays a thumbnail of a target video, which is generated by the electronic device based on the first photo of the user corresponding to the first fusion ID in the gallery application.

[0008] Based on the first aspect, the electronic device can predict whether the faces corresponding to the first user ID and the second user ID among multiple user IDs belong to the same person. If so, the electronic device can merge the first user ID and the second user ID corresponding to the same person's face into a first fused ID. Based on the first fused ID, it can predict the social relationship between the user corresponding to the first fused ID and the device owner. This avoids repeatedly predicting the social relationship between the user and the device owner corresponding to the same face, reducing data processing and achieving the goal of reducing power consumption. Furthermore, when predicting whether the faces corresponding to the first user ID and the second user ID belong to the same person, the electronic device can determine whether it has cached the first face similarity between the first user ID and the second user ID. If so, it can directly use the cached first face similarity to predict whether the first user ID and the second user ID belong to the same person, without repeated calculations, resulting in a better overall solution and improved performance.

[0009] In one possible implementation of the first aspect, before the electronic device determines whether it has cached the first face similarity, the method further includes: the electronic device periodically determining the first face similarity and caching the determined first face similarity in the electronic device. Here, the first face similarity currently cached by the electronic device, as indicated by the current determination result, is the one cached by the electronic device after the last determination of the first face similarity.

[0010] In this way, by periodically determining the first similarity between the first user ID and the second user ID, the electronic device makes the calculation results more comprehensive and accurate. This leads to better performance in predicting whether the faces corresponding to the first user ID and the second user ID belong to the same person based on the first face similarity. Furthermore, since the first face similarity currently used by the electronic device is cached after the previous determination, it can directly use the cached value when predicting the same person's face in the current iteration, without recalculating. This results in better overall performance and a more effective solution.

[0011] In one possible implementation of the first aspect, the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's faces based on the first face similarity, including: if the number of first photos is greater than or equal to a first quantity threshold, and the number of second photos is greater than or equal to a second quantity threshold, then the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's faces based on the first face similarity; wherein, the number of first photos is the number of photos of the face corresponding to the first user ID cached by the electronic device, and the number of second photos is the number of photos of the face corresponding to the second user ID cached by the electronic device.

[0012] And / or; if the number of first photos and the number of third photos are the same, and the number of second photos and the number of fourth photos are the same, then the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's face based on the first face similarity; wherein, the number of third photos is the number of photos in the image library application that include the face corresponding to the first user ID, and the number of fourth photos is the number of photos in the image library application that include the face corresponding to the second user ID.

[0013] In this way, when predicting whether the faces corresponding to the first user ID and the second user ID are the same person's faces, the electronic device can first determine whether the above conditions are met. If they are met, it means that the reliability of the first face similarity currently cached by the electronic device is high and can be used directly, thus improving the reliability of the electronic device in predicting the same person's face.

[0014] In one possible implementation of the first aspect, the electronic device determines the first facial similarity by: selecting n first target photos from photos of faces corresponding to a first user ID, and selecting m second target photos from photos of faces corresponding to a second user ID; where n and m are positive integers. Based on the n first target photos and the m second target photos, the electronic device determines whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range; if the genders of the users corresponding to the first user ID and the second user ID are the same, and the age difference is within the preset range, then the electronic device determines the first facial similarity.

[0015] Thus, when determining the first face similarity, the electronic device can first check whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range. If so, it indicates that the faces corresponding to the first user ID and the second user ID are likely to belong to the same person. Therefore, the electronic device can continue to calculate the first face similarity, resulting in a better overall solution and improved performance.

[0016] In one possible implementation of the first aspect, the method further includes: if the genders of the users corresponding to the first user ID and the second user ID are different; and / or, the age difference is outside a preset range, then the electronic device abandons determining the first face similarity.

[0017] Therefore, if the users corresponding to the first user ID and the second user ID do not meet the above conditions, it means that the faces corresponding to the first user ID and the second user ID are not the same person's faces. Therefore, we can give up calculating the similarity of the first face corresponding to the first user ID and the second user ID, thereby achieving the goal of improving the speed of calculating face similarity.

[0018] In one possible implementation of the first aspect, the electronic device determines the first face similarity by: extracting x first face features from a photograph of a face corresponding to a first user ID, and extracting y second face features from a photograph of a face corresponding to a second user ID; where x and y are positive integers. The electronic device determines the similarity between each of the x first face features and each of the y second face features, obtaining z face feature similarities; z = x × y. Furthermore, the electronic device determines the first face similarity based on the z face feature similarities.

[0019] Thus, by extracting the facial features corresponding to the first user ID and the second user ID, the accuracy of calculating the similarity of the first face can be effectively improved.

[0020] In one possible implementation of the first aspect, the method further includes: after the electronic device obtains z facial feature similarities, the electronic device caches the z facial feature similarities; before the electronic device next determines the first facial similarity, the electronic device determines whether it currently caches z facial feature similarities. If the electronic device currently caches z facial feature similarities, the electronic device performs the step of next determining the first facial similarity based on the currently cached z facial feature similarities.

[0021] In this way, after obtaining the similarity of z facial features, the electronic device can cache them for future use, resulting in a better overall solution and improved performance, thus accelerating computation.

[0022] In a second aspect, an electronic device is provided, which has the functions described in any one of the first aspects above. These functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions.

[0023] Thirdly, an electronic device is provided, comprising: a memory and one or more processors; the memory stores computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method described in any one of the first aspects.

[0024] Fourthly, a chip system is provided, the chip system comprising: at least one processor and an interface for receiving instructions and transmitting them to the at least one processor; the at least one processor executes instructions to cause an electronic device to perform the method described in any one of the first aspects.

[0025] Fifthly, a computer-readable storage medium is provided that stores instructions which, when executed on a computer, cause the computer to perform the method described in any one of the first aspects.

[0026] In a sixth aspect, a computer program product containing instructions is provided, which, when run on a computer, enables the computer to perform the method described in any one of the first aspects above.

[0027] The technical effects of any of the implementation methods in the second to sixth aspects mentioned above can be referred to the technical effects of different implementation methods in the first aspect, and will not be elaborated here. Attached Figure Description

[0028] Figure 1 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0029] Figure 2 A schematic diagram of an interface for a video generation method based on social relationships provided in an embodiment of this application;

[0030] Figure 3 A schematic diagram illustrating the prediction of social relationships provided in an embodiment of this application;

[0031] Figure 4 This is a schematic diagram illustrating an embodiment of the present application for predicting whether the faces corresponding to a first user ID and a second user ID are the same person's faces;

[0032] Figure 5 A schematic diagram illustrating a process for calculating facial similarity, provided as an embodiment of this application;

[0033] Figure 6 A schematic diagram illustrating the calculation of facial feature similarity as provided in an embodiment of this application;

[0034] Figure 7 A schematic diagram illustrating another process for calculating face similarity provided in an embodiment of this application;

[0035] Figure 8 A flowchart illustrating a process for predicting whether the faces corresponding to a first user ID and a second user ID belong to the same person, provided as an embodiment of this application.

[0036] Figure 9 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0037] The video generation method based on social relationships provided in this application can be applied to electronic devices that support voice interaction functions, such as smartphones, personal digital assistants (PDAs), laptops, tablets, handheld computers, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, cellular phones, PDAs, and augmented reality (AR) / virtual reality (VR) devices. This application does not impose any special limitations on the specific form of the electronic device.

[0038] refer to Figure 1 This is a hardware structure diagram of an electronic device provided in an embodiment of this application. Figure 1 As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a positioning module 181, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0039] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0040] Processor 110 may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0041] In some embodiments, the processor 110 can perform tasks related to identifying the device owner, voice interaction, and identifying social relationships with the device owner. In other embodiments, the processor 110 can also perform tasks related to predicting whether the faces corresponding to any two user IDs (i.e., tagids) belong to the same person. The electronic device 100 includes multiple photos, each photo corresponding to multiple user IDs, and each user ID corresponding to one face.

[0042] The charging management module 140 receives charging input from a charger, which can be either a wireless or wired charger. The power management module 141 connects to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.

[0043] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0044] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include one or more filters, switches, power amplifiers, low-noise amplifiers (LANs), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the demodulated signal and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functions of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0045] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices within one or more communication processing modules. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0046] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor that connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0047] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a Micro-LED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0048] Electronic device 100 can perform shooting functions through an ISP, camera 193, video codec, GPU, display 194, and application processor. The ISP processes data fed back from the camera 193. For example, when taking a picture, the shutter is opened, light is transmitted through the lens to the camera's photosensitive element, the light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into a visible image. The ISP can also perform algorithmic optimization of image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature. In some embodiments, the ISP can be integrated into the camera 193.

[0049] Camera 193 is used to capture images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1.

[0050] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0051] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing by running the instructions stored in internal memory 121. For example, processor 110 can display different content on display screen 194 in response to an operation to unfold display screen 194 by executing instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phone book, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0052] Electronic device 100 can implement audio functions through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor, such as music playback and recording.

[0053] Buttons 190 include a power button and volume buttons. Buttons 190 can be mechanical buttons or touch buttons. Electronic device 100 can receive button inputs and generate key signal inputs related to user settings and function control of electronic device 100. Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with electronic device 100. Electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0054] The video generation method based on social relationships provided in this application can be implemented in an electronic device 100 with the above-described hardware structure. The following description uses a mobile phone as an example to illustrate the solution of this application's embodiment.

[0055] The following describes the scenarios involved in the embodiments of this application:

[0056] In some embodiments, a mobile phone supporting voice control can capture user voice through a microphone, parse and recognize the user's voice, and execute commands corresponding to the user's voice, thereby enabling the user to control the phone via voice. Optionally, the phone can also perform functions such as conversing with the user.

[0057] It should be noted that the names of voice control functions may differ between different manufacturers' phones. For example, voice control functions may also be called "Smart Assistant," "Smart Application," "Voice Assistant," "Smart Projection," etc., without restriction. Furthermore, the specific implementation of voice control functions with different names may also vary, which will not be listed here.

[0058] To conserve power and prevent accidental activation, voice control functionality must be enabled before use. For example, the phone can respond to a preset word (or "wake-up word"; for example, "Hello, YOYO") input by the user to activate voice control. Alternatively, the phone can respond to a preset action input by the user to activate voice control. Optionally, this preset action could be turning on a preset switch in the user interface; or it could be an action on the power button (or power button) (such as a long press).

[0059] In some embodiments, after the phone is woken up, it can receive voice input from the user and execute the corresponding command in response to the voice. Once the phone has executed the command, or if the phone has been woken up for more than a preset duration (e.g., 8 seconds), the phone disables voice control and stops receiving voice input from the user. If the user wants to control the phone via voice again, they need to wake the phone up again to enable voice control, and only then can the phone receive voice input from the user again. In other words, after the phone is woken up, it enters a "short reception" state, during which it can receive voice input from the user for a period of time (e.g., 8 seconds).

[0060] In other embodiments, the phone supports entering a continuous dialogue scenario after being woken up, allowing users to engage in ongoing conversations. For example, when the phone is connected to the network, it will continue recording audio after each response to user-inputted voice, without requiring the user to repeatedly wake the phone. The continuous dialogue can only be exited by the user disabling the voice control function. This avoids the need for repeated phone wake-ups, improving both the user experience and the efficiency of human-computer interaction.

[0061] For example, after a mobile phone enters continuous audio reception mode, it can continuously receive voice input from the user and execute commands corresponding to the user's voice input. For example, in response to the user's voice input, based on information including instructions to the phone owner and information about a first social relationship with the owner, the phone can retrieve photos of target users containing these first social relationships from a photo library application, and generate a video based on these photos. This results in a video with better logic, stronger relevance, and further improved human-computer interaction efficiency.

[0062] It should be noted that in this application embodiment, social relationship can be understood as "intimate relationship," which may include lover, wife, husband, child, father, mother, etc., without limitation. Having a first social relationship with the device owner can be understood as the device owner's lover, wife, husband, child, father, or mother, etc., without limitation. It should be understood that the names of the social relationships shown in this application embodiment are merely examples and do not constitute a limitation on this application. Of course, social relationships can also have other different names; for example, a wife can also be called "wife," a husband can also be called "husband," a father can also be called "dad," etc., and a mother can also be called "mom," etc., without limitation.

[0063] For example, the mobile phone stores a mapping between multiple user IDs and multiple social relationships. This mapping can be stored in the phone as a table or array. For instance, the mapping could be shown in Table 1 below.

[0064] Table 1

[0065] social relationships User ID wife tag id1 child tag id2 Father tag id3 Mother tag id4

[0066] The aforementioned social relationships refer to the relationship between the device owner and the user corresponding to the user ID. For example, the relationship between the device owner and the user corresponding to tagid1 is that of wife; the relationship between the device owner and the user corresponding to tagid2 is that of child; the relationship between the device owner and the user corresponding to tagid3 is that of father; and the relationship between the device owner and the user corresponding to tagid4 is that of mother.

[0067] For example, a mobile phone receives a voice message 'a' input by a user. Voice message 'a' includes information instructing the phone owner and information about a social relationship 'a' (or first social relationship) with the phone owner. For instance, voice message 'a' could be "Generate a video from my father's photo." Here, "I" in voice message 'a' is the information instructing the phone owner, and "father" is the information about a social relationship 'a' with the phone owner.

[0068] Optionally, after receiving voice input 'a' from the user, the mobile phone can recognize voice 'a' to identify "I" as information indicating the phone owner and "father" as information indicating a social relationship 'a' with the phone owner. For example, when voice 'a' matches keywords included in the phone owner's database, the phone recognizes "I" as information indicating the phone owner. For instance, keywords in the phone owner's database include "I," "myself," "myself," etc., and the "I" in voice 'a' matches the keyword "I" in the phone owner's database, therefore the phone can recognize "I" as information indicating the phone owner. Similarly, when voice 'a' matches keywords included in a relationship database, the phone recognizes "father" as information indicating a social relationship 'a' with the phone owner. For instance, keywords in the relationship database include "father," "dad," "mom," etc., and the "father" in voice 'a' matches the keyword "father" in the relationship database, therefore the phone can recognize "father" as information indicating a social relationship 'a' with the phone owner.

[0069] Optionally, the phone can perform semantic understanding on voice 'a'. This voice 'a' doesn't need to be completely identical to the database, as long as the semantics are the same or close. That is, voice 'a' matches the database, but doesn't have to be completely identical. For example, voice 'a' includes the word "I," and the keywords in the phone owner's database include "myself," "myself," and "myself." Since "I" has the same semantic meaning as "myself," "myself," and "myself," the two match. Therefore, the phone can recognize that "I" is information indicating the phone owner, thus determining that "I" is used to indicate the phone owner.

[0070] Correspondingly, the voice 'a' includes 'father', and the keywords in the relational database include 'dad'; since 'father' and 'dad' have the same meaning, they match, so the phone can recognize 'father' as information that has a social relationship 'a' with the owner, thus determining that the social relationship with the owner is 'father'.

[0071] Optionally, after receiving voice input 'a' from the user, the mobile phone can recognize and analyze 'a' to identify "I" as an instruction to the phone owner and "father" as information indicating a social relationship 'a' with the phone owner. For example, after receiving voice 'a', the mobile phone uses automatic speech recognition (ASR) technology to convert the user's speech into text, and then uses natural language understanding (NLU) technology to perform intent recognition on the converted text, thereby recognizing "I" as an instruction to the phone owner and "father" as information indicating a social relationship 'a' with the phone owner.

[0072] Optionally, after the phone determines that the social relationship between the user and the owner in voice message 'a' is "father," the phone determines the target user ID corresponding to the social relationship "father" as tag id3 based on the correspondence shown in Table 1. Then, the phone retrieves a photo of the target user corresponding to tag id3 from the gallery application and generates a video based on the photo, i.e., a video of the owner's father. This results in a video with better logic, stronger relevance, and further improves the efficiency of human-computer interaction.

[0073] In some embodiments, when the phone is woken up and engaged in continuous conversation, the phone can also display the conversation content with the user through a human-computer interaction interface. For example, the phone can display... Figure 2 The interface 101 shown is used to display the conversation content between the mobile phone and the user during a continuous dialogue. Optionally, the interface 101 includes a prompt message 102, which is used to notify the user that the mobile phone has entered a continuous dialogue scenario. For example, the prompt message 102 could be "Welcome". Optionally, the interface 101 also includes a prompt icon 103, which is used to notify the user that the mobile phone has entered continuous audio recording mode. For example, the prompt icon 103 could be a microphone icon, or a "robot audio recording icon", etc., without limitation.

[0074] For example, during the process of the mobile phone receiving voice input from the user, the mobile phone displays... Figure 2 The interface 104 shown is used to notify the user that the microphone is recording sound. Optionally, such as... Figure 2 As shown, interface 104 includes a notification icon 105, which consists of gradually decreasing transparent circles displayed horizontally in the direction of notification icon 103, used to notify the user that the microphone is recording sound. Optionally, text information corresponding to voice a is also displayed above notification icon 105.

[0075] For example, the mobile phone receives a voice input 'a' from the user. In response to voice 'a', the mobile phone displays interface 106, which includes text information corresponding to voice 'a'. For example, the text information could be "Generate a video from my father's photo". Accordingly, in response to voice 'a', the mobile phone identifies the information with a social relationship 'a' with the owner as "father", and determines the target user ID corresponding to social relationship 'a' (i.e., "father") as tag id3. Then, the mobile phone retrieves the photo of the target user corresponding to tag id3 from the gallery application and displays it on interface 106. Optionally, before displaying the photo of the target user corresponding to tag id3, the mobile phone can also display the response information for voice 'a'. For example, the response information could be "Okay, the following materials have been found for you". This can further improve the user experience.

[0076] For example, interface 106 also includes control 1 and control 2. Control 1 is used to trigger the phone to display all photos related to the target user corresponding to tag id3, and control 2 is used to trigger the phone to generate a video from the photos of the target user corresponding to tag id3. Optionally, such as Figure 2 As shown, control 1 can be named "View All", and control 2 can be named "Generate Video". In response to the user's operation on control 2, the phone displays interface 107, which includes a thumbnail of video a. Video a is generated by the phone based on all photos related to the target user corresponding to tag id3. Optionally, interface 107 includes control 3, which is used to trigger the phone to play video a.

[0077] In summary, because the phone stores multiple user IDs and their corresponding social relationships, upon receiving voice input from a user, the phone identifies the primary social relationship with the user within the voice message. Then, the phone determines the target user ID corresponding to this primary social relationship, retrieves the photo of the target user from the photo library, and generates a video. This approach ensures that the video generated based on the primary social relationship has better logical coherence and stronger relevance, further improving the efficiency of human-computer interaction.

[0078] In some embodiments, the mobile phone can predict the social relationship between the user corresponding to each of multiple user IDs and the phone owner under a target state. For example, the target state could be a charging / screen-off state, or a black screen state, etc., without limitation. For instance, for each user ID, the mobile phone can... Figure 3 The process shown predicts the social relationship between the user and the device owner corresponding to each user ID.

[0079] Taking the prediction of the social relationship between a target user ID and the phone owner via mobile phone as an example, for instance, such as... Figure 3As shown, the mobile phone inputs the user information of the user corresponding to the target user ID and the user information of the phone owner into the relationship prediction model to predict the social relationship between the user corresponding to the target user ID and the phone owner.

[0080] The user information corresponding to the target user ID may include, for example, the number of photos containing the target user ID, the number of image sets containing the target user ID, and the location type of the image sets containing the target user ID, without limitation. Similarly, the user information corresponding to the device owner may include, for example, the number of photos containing the device owner, the number of image sets containing the device owner, and the location type of the image sets containing the device owner, without limitation.

[0081] Within the same photo set, the time interval between any two photos taken does not exceed a preset interval, and the distance difference between the shooting locations of any two photos does not exceed a preset distance. The location type of the photo set can be, for example, "home," "old home," "indoor," or "outdoor," and is not restricted.

[0082] Optional, such as Figure 3 As shown, the mobile phone can also input the visual relationship between the user corresponding to the target user ID and the phone owner into the relationship prediction model, which helps to improve the accuracy of predicting the social relationship between the user corresponding to the target user ID and the phone owner. The visual relationship can include, for example, closeness, back-to-back contact, proximity, or hugging, and is not limited to these.

[0083] In some embodiments, the phone's gallery application stores multiple photos, each corresponding to a different user ID, and each user ID corresponding to a face. Optionally, before predicting the social relationship between a user ID and the phone's owner, the phone can also predict whether the faces corresponding to any two user IDs belong to the same person. This allows the two user IDs corresponding to the same person's face to be merged together, and the social relationship between the user corresponding to the first merged ID and the phone's owner can be predicted based on the merged user ID (or the first merged ID). This avoids repeatedly predicting the social relationship between the same face and the phone's owner, reducing data processing and thus lowering power consumption.

[0084] For example, a mobile phone can input the facial similarity scores of every two user IDs from multiple user IDs into a prediction model for the same person, predicting whether the faces corresponding to each pair of user IDs belong to the same person. For instance, as... Figure 4As shown, the mobile phone inputs the first face similarity scores corresponding to the first user ID and the second user ID into the same person prediction model, and predicts whether the faces corresponding to the first user ID and the second user ID belong to the same person based on the first face similarity scores. Here, the first user ID and the second user ID are any two user IDs from a pool of user IDs.

[0085] Optional, such as Figure 4 As shown, the mobile phone can also input features such as the gender difference between the first user ID and the second user ID, the average age (not shown in the figure), the frequency of photo co-occurrence, and the frequency of image set co-occurrence into the same person prediction model. This can make the prediction results more accurate and reliable.

[0086] The gender differences include: both the first user ID and the second user ID correspond to male users; both the first user ID and the second user ID correspond to female users; the first user ID corresponds to a male user and the second user ID corresponds to a female user; or, the first user ID corresponds to a female user and the second user ID corresponds to a male user, without restriction.

[0087] Photo co-occurrence frequency refers to the frequency with which users corresponding to the first user ID and the second user ID appear together in the same photo. Image set co-occurrence frequency refers to the frequency with which users corresponding to the first user ID and the second user ID appear together in the same image set.

[0088] In summary, in this embodiment, when predicting whether the faces corresponding to two user IDs belong to the same person, the mobile phone needs to input the face similarity between the two user IDs into the same person prediction model. Therefore, before predicting whether the faces corresponding to two user IDs belong to the same person, the mobile phone first needs to calculate the face similarity between the two user IDs. For example, the mobile phone can calculate the face similarity between the two user IDs based on the facial features of the user corresponding to each user ID. Since the facial features of the user corresponding to each user ID are extracted by the mobile phone from multiple photos stored in the photo library application, and these photos may be updated in real time, the mobile phone can periodically calculate the face similarity between the two user IDs to make the calculation results more comprehensive and accurate, thereby improving the effectiveness of the mobile phone in predicting whether the faces corresponding to two user IDs belong to the same person based on face similarity.

[0089] For example, the mobile phone can periodically calculate the facial similarity between every two user IDs in the following ways: the mobile phone calculates the facial similarity between every two user IDs every day; or the mobile phone calculates the facial similarity between every two user IDs every week, etc., without limitation.

[0090] The following describes the specific implementation of the facial similarity calculation for each pair of user IDs involved in the embodiments of this application:

[0091] In some embodiments, before the mobile phone calculates the facial similarity between any two user IDs from the multiple user IDs, the mobile phone first selects M user IDs from the multiple user IDs. These M user IDs correspond to users whose frequency in the image set is greater than or equal to M. That is, the M users corresponding to the M user IDs are the top M users with the highest frequency in the image set. M is a positive integer.

[0092] After selecting M user IDs from multiple user IDs, the mobile phone can calculate the facial similarity between any two user IDs among the M user IDs.

[0093] It's important to note that the more frequently a user appears in the image set, the greater the likelihood that the user ID corresponds to a social relationship with the device owner. Therefore, by calculating the facial similarity between every two user IDs out of M user IDs, the reliability of subsequent social relationship prediction schemes can be improved, resulting in better overall performance.

[0094] Optionally, after selecting M user IDs from multiple user IDs, the phone selects photos from the M photos corresponding to the M user IDs where the facial angle is smaller than a preset angle, resulting in Z photos. These Z photos correspond to Z user IDs. Based on this, the phone can calculate the facial similarity between every two user IDs within the Z user IDs. Z is a positive integer, and Z ≤ M.

[0095] For example, a mobile phone can obtain the facial angle corresponding to each of the M photos from a gallery application, and select Z photos whose facial angles are smaller than a preset angle. The gallery application can recognize each photo in the gallery application and obtain the facial information corresponding to each photo when the phone is charging and the screen is off (or black). The facial information can include facial angle, user age, user gender, etc., and is not limited.

[0096] For example, the facial angle refers to the angle at which a face is turned to the side. Typically, 0° represents a frontal view, -90° represents a face turned 90° to the left, and +90° represents a face turned 90° to the side. It's important to note that a larger facial angle indicates a greater degree of side profile or back view in the photo, resulting in fewer facial features that the phone can extract, and consequently, lower accuracy in calculating facial similarity based on those features.

[0097] Thus, by selecting Z user IDs corresponding to photos with face angles smaller than a preset angle, and calculating the face similarity between every two user IDs among the Z user IDs, the accuracy of mobile phone face similarity calculation can be improved.

[0098] For example, the preset angle can be set according to specific circumstances and is not limited. For instance, the preset angle can be ±30° or ±25°, etc., and is not limited.

[0099] In this embodiment of the application, the facial similarity between any two user IDs among the Z user IDs can be calculated in the following manner. The following embodiment uses the calculation of the first facial similarity between the first user ID and the second user ID by a mobile phone as an example for illustration.

[0100] In some embodiments, such as Figure 5 As shown, the mobile phone extracts x first facial features from the photo of the face corresponding to the first user ID, and y second facial features from the photo of the face corresponding to the second user ID; x and y are positive integers. Then, the mobile phone determines the similarity between each of the x first facial features and each of the y second facial features, obtaining z facial feature similarities, and determines the first facial similarity based on the z facial feature similarities. z = x × y.

[0101] For example, such as Figure 6 As shown, assume that x first facial features include: first facial feature 1, first facial feature 2, and first facial feature 3; and y second facial features include: second facial feature 1, second facial feature 2, and second facial feature 3. It should be noted that, typically, the number of facial features that can be extracted from the facial photo corresponding to each user ID does not exceed a preset value (e.g., 50). That is, the maximum number of facial features that can be extracted from the facial photo corresponding to each user ID is 50. Currently, the preset value can also be other values, subject to actual implementation, and is not restricted here.

[0102] For example, such as Figure 6 As shown, the mobile phone determines the similarity between the first facial feature 1 and the second facial feature 1, the second facial feature 2, and the second facial feature 3, respectively, obtaining facial feature similarity 1, facial feature similarity 2, and facial feature similarity 3. Correspondingly, the mobile phone determines the similarity between the first facial feature 2 and the second facial feature 1, the second facial feature 2, and the second facial feature 3, respectively, obtaining facial feature similarity 4, facial feature similarity 5, and facial feature similarity 6. Correspondingly, the mobile phone determines the similarity between the first facial feature 3 and the second facial feature 1, the second facial feature 2, and the second facial feature 3, respectively, obtaining facial feature similarity 7, facial feature similarity 8, and facial feature similarity 9.

[0103] Furthermore, the mobile phone determines the first face similarity based on face feature similarity 1, face feature similarity 2, ..., face feature similarity 9. Optionally, the mobile phone can select the top N face feature similarities with the highest face feature similarity from the z face feature similarities and use the average of these N face feature similarities as the first face similarity. This can improve the accuracy of the first face similarity.

[0104] Optionally, the mobile phone can also use the mean of the similarity scores of z facial features as the first facial similarity score; or, the mobile phone can also use the variance (or standard deviation) of the similarity scores of z facial features as the first facial similarity score; or, the mobile phone can also use the maximum value among the similarity scores of z facial features as the first facial similarity score, etc., without restriction.

[0105] In some embodiments, such as Figure 5 As shown, before calculating the first face similarity between the first user ID and the second user ID, the phone can also determine whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range. If the genders of the users corresponding to the first user ID and the second user ID are the same, and the age difference is within the preset range, then the phone starts calculating the first face similarity between the first user ID and the second user ID. If the genders of the users corresponding to the first user ID and the second user ID are different, and / or the age difference is outside the preset range, then the phone abandons calculating the first face similarity between the first user ID and the second user ID.

[0106] The preset range can be set according to the actual implementation and is not limited. For example, the preset range can be [0, 20] years old.

[0107] Therefore, before calculating the face similarity between the first user ID and the second user ID, the system first verifies whether the users corresponding to the first user ID and the second user ID are the same person based on their age and gender. If the age and gender of the users corresponding to the first user ID and the second user ID meet the conditions, it indicates a high probability that the users are the same person, and the face similarity between the first user ID and the second user ID can be calculated. If the age and gender of the users corresponding to the first user ID and the second user ID do not meet the conditions, it indicates that the users are not the same person, and therefore the calculation of the face similarity between the first user ID and the second user ID can be abandoned, thereby improving the speed of face similarity calculation.

[0108] Optionally, the mobile phone selects n first target photos from the photos of the faces corresponding to the first user ID, and m second target photos from the photos of the faces corresponding to the second user ID; n and m are positive integers. Then, based on the n first target photos and the m second target photos, the mobile phone determines whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range.

[0109] Optionally, the mobile phone can first determine whether the number of photos of the face corresponding to the first user ID is greater than or equal to a threshold, and then determine whether the number of photos of the face corresponding to the second user ID is greater than or equal to a threshold. Here, n and m are less than or equal to the threshold. If the number of photos of the face corresponding to the first user ID is greater than or equal to the threshold, and the number of photos of the face corresponding to the second user ID is also greater than or equal to the threshold, then the mobile phone selects n first target photos from the photos of the face corresponding to the first user ID, and m second target photos from the photos of the face corresponding to the second user ID.

[0110] Optionally, if the number of photos of the face corresponding to the first user ID is less than a threshold, and the number of photos of the face corresponding to the second user ID is less than a threshold, then the mobile phone determines, based on the photos of the faces corresponding to the first user ID and the second user ID, whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range.

[0111] Taking an example where n, m, and the threshold are all 5, for instance, if the number of photos of the face corresponding to the first user ID is greater than or equal to 5, and the number of photos of the face corresponding to the second user ID is greater than or equal to 5, then the mobile phone selects 5 first target photos from the photos of the face corresponding to the first user ID, and selects 5 second target photos from the photos of the face corresponding to the second user ID. Based on the 5 first target photos and the 5 second target photos, the mobile phone determines whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range.

[0112] For example, suppose that in the 5 first target photos, the gender of the user corresponding to the first user ID in each first target photo is: female, female, female, female, male (hereinafter referred to as list1). In the 5 second target photos, the gender of the user corresponding to the second user ID in each second target photo is: female, female, female, female, female (hereinafter referred to as list2). Then, the mobile phone can select the mode of list1 and list2 for comparison. If the mode of list1 and list2 is the same, the mobile phone determines that the gender of the user corresponding to the first user ID and the second user ID is the same.

[0113] The mode refers to the value that appears most frequently in a set of data. As we can see, the mode of list1 is female, and the mode of list2 is also female.

[0114] For example, suppose that in the five first target photos, the ages of the users corresponding to the first user ID in each photo are: age1, age2, age3, age4, and age5 (hereinafter referred to as list3). And in the five second target photos, the ages of the users corresponding to the second user ID in each photo are: age6, age7, age8, age9, and age10 (hereinafter referred to as list4). Then, the phone can calculate the average (or variance, or standard deviation, etc., of list3, without restriction) as the age of the user corresponding to the first user ID, and the phone can also calculate the average (or variance, or standard deviation, etc., of list4, without restriction) as the age of the user corresponding to the second user ID. Furthermore, the phone can determine whether the age difference is within a preset range based on the ages of the users corresponding to the first and second user IDs.

[0115] In this way, the mobile phone first selects a small portion (i.e. n photos) of the first target photos from the photos of the faces corresponding to the first user ID, and selects a small portion (i.e. m photos) of the second target photos from the photos of the faces corresponding to the second user ID. It first verifies whether the genders of the users corresponding to the first user ID and the second user ID are the same and whether the age difference is within a preset range, so as to determine whether it is necessary to calculate the similarity of the first faces corresponding to the first user ID and the second user ID, thereby accelerating the calculation.

[0116] In some embodiments, such as Figure 5 As shown, after obtaining the similarity scores of z facial features, the phone caches these z facial feature similarities in the phone for use when calculating the similarity score of the first face again. For example, the phone caches the z facial feature similarities in the phone's first storage area (hereinafter referred to as storage area 1).

[0117] For example, storage area 1 caches the facial feature similarity of z individuals corresponding to the first user ID and the second user ID. For instance, the phone can cache the facial feature similarity of z individuals corresponding to the first user ID and the second user ID as a set. For example, combining... Figure 6 As shown, the z facial feature similarities corresponding to the first user ID and the second user ID cached on the mobile phone can be represented as: First user ID, second user ID = {facial feature similarity 1, facial feature similarity 2, ..., facial feature similarity 9}.

[0118] In this way, by caching the similarity of z facial features in storage area 1 of the phone, when the phone calculates the similarity of the first face again, it can directly read the similarity of z facial features from storage area 1 and calculate the similarity of the first face based on the similarity of z facial features. There is no need to repeat the calculation of the similarity of z facial features in the next calculation, which makes the overall effect of the solution better and the performance better, thus achieving the goal of accelerating the calculation.

[0119] In some embodiments, such as Figure 5 As shown, after calculating the first face similarity, the phone caches the first face similarity in the phone for use in the next prediction of whether the faces corresponding to the first user ID and the second user ID belong to the same person. For example, the phone caches the first face similarity in the phone's second storage area (hereinafter referred to as storage area 2).

[0120] For example, storage area 2 caches the correspondence between the first user ID, the second user ID, and the first face similarity. For instance, storage area 2 of the phone caches tag id1, tag id2, and sim1; where tag id1 represents the first user ID, tag id2 represents the second user ID, and sim1 represents the first face similarity.

[0121] Thus, by caching the first face similarity in the phone's storage area 2, the phone can directly read the first face similarity from storage area 2 the next time it uses the first face similarity, and predict whether the faces corresponding to the first user ID and the second user ID are the same person's face based on the first face similarity. There is no need to recalculate the first face similarity the next time, which makes the overall solution more effective and has better performance, thus achieving the goal of accelerating prediction.

[0122] Understandably, when calculating the first face similarity for the first time, the mobile phone can use the above method. Figure 5 and Figure 6 The process shown calculates the first face similarity. When the phone calculates the first face similarity a second time (and thereafter), as... Figure 7 As shown, before the phone extracts x first facial features from the photo of the face corresponding to the first user ID and y second facial features from the photo of the face corresponding to the second user ID, the phone can determine whether z facial feature similarities are cached in the phone. If the current determination result indicates that z facial feature similarities are cached in the phone, then the phone determines the first facial similarity based on the z facial feature similarities. Here, the current determination result indicates that the z facial similarities cached in the phone were cached when the phone last calculated the first facial similarity.

[0123] Optionally, if the current judgment result indicates that there are no z facial feature similarities cached in the mobile phone, then the mobile phone extracts x first facial features from the photo of the face corresponding to the first user ID, extracts y second facial features from the photo of the face corresponding to the second user ID, and calculates z facial feature similarities.

[0124] Optional, such as Figure 7 As shown, if the current judgment result indicates that the phone has z facial similarity scores cached, the phone further determines whether there are any new photos in the gallery application. If there are new photos in the gallery application, and the new photos are photos of the user corresponding to the first user ID or the user corresponding to the second user ID, the phone recalculates the first facial similarity scores corresponding to the first user ID and the second user ID. If there are no new photos in the gallery application, the phone directly uses the z facial similarity scores to calculate the first facial similarity score.

[0125] Optionally, assuming a second photo is added to the photo library application, and this second photo is a photo of the user corresponding to the first user ID, the phone extracts i third facial features from the second photo and calculates the similarity between each of the i third facial features and each of the y second facial features, obtaining i×y personal facial feature similarities. Then, the phone adds the i×y personal facial feature similarities to the set corresponding to z personal facial feature similarities. Finally, the phone calculates the first facial similarity based on the i×y personal facial feature similarities and the z personal facial feature similarities.

[0126] Optionally, assuming a second photo is added to the photo library application, and this second photo is a photo of the user corresponding to the second user ID, the phone extracts j fourth facial features from the second photo and calculates the similarity between each of the j fourth facial features and each of the x first facial features, obtaining j×x personal facial feature similarities. Then, the phone adds the j×x personal facial feature similarities to the set corresponding to z personal facial feature similarities. Finally, the phone calculates the first facial similarity based on the j×x personal facial feature similarities and the z personal facial feature similarities.

[0127] It should be noted that the specific implementation process of calculating the similarity between each of the i third face features and each of the y second face features; or the specific implementation process of calculating the similarity between each of the j fourth face features and each of the x first face features can refer to the above embodiments, and will not be repeated here.

[0128] Accordingly, the mobile phone calculates the first face similarity based on the facial feature similarity of individuals i×y and z; or, the mobile phone calculates the first face similarity based on the facial feature similarity of individuals j×x and z. The specific implementation of this method can be found in the above embodiments and will not be repeated here. Optionally, such as... Figure 7 As shown, the mobile phone caches the facial feature similarity of i×y (or j×x) individuals and the facial feature similarity of z individuals in storage area 1, and caches the calculated first facial similarity in storage area 2.

[0129] In this way, when a new photo is added to the photo library, the phone only needs to calculate the facial feature similarity based on the facial features of the new photo, without having to repeat the above steps to calculate the facial feature similarity of z people, thereby reducing the amount of computation and achieving the goal of speeding up the calculation.

[0130] In some embodiments, such as Figure 8 As shown, when the mobile phone is predicting whether the faces corresponding to the first user ID and the second user ID belong to the same person, the phone can determine whether the first face similarity is cached in the phone. If the result of this determination indicates that the first face similarity is cached in the phone, then the phone predicts whether the faces corresponding to the first user ID and the second user ID belong to the same person based on the first face similarity.

[0131] It is understandable that since the mobile phone determines the first face similarity periodically (e.g., "daily") and caches the determined first face similarity in the mobile phone, if the current judgment result indicates that the mobile phone has cached the first face similarity, then the first face similarity currently cached by the mobile phone indicated by the current judgment result is the one cached by the mobile phone after the last determination of the first face similarity.

[0132] Optional, such as Figure 8 As shown, if the current result indicates that the phone does not have the first face similarity score cached, then the phone calculates the first face similarity score. The specific implementation process for calculating the first face similarity score can be found above. Figure 5 , Figure 6 and Figure 7 The process shown will not be repeated here.

[0133] Optionally, storage area 2 of the phone also caches the number of photos of faces corresponding to the first user ID (hereinafter referred to as the first number of photos, denoted as t1) and the number of photos of faces corresponding to the second user ID (hereinafter referred to as the second number of photos, denoted as t2). The phone's gallery application includes the number of photos of faces corresponding to the first user ID (hereinafter referred to as the third number of photos, denoted as t3) and the number of photos of faces corresponding to the second user ID (hereinafter referred to as the fourth number of photos, denoted as t4).

[0134] In some embodiments, such as Figure 8 As shown, if the current judgment result indicates that the phone has cached the first face similarity, the phone further determines whether condition 1 and / or condition 2 are satisfied. Condition 1 includes: whether the number of first photos t1 is greater than or equal to a first quantity threshold (hereinafter referred to as threshold 1), and whether the number of second photos t2 is greater than or equal to a second quantity threshold (hereinafter referred to as threshold 2). Condition 2 includes: whether the number of first photos t1 and the number of third photos t3 are equal, and whether the number of second photos t2 and the number of fourth photos t4 are equal.

[0135] Optionally, if the number of first photos t1 is greater than or equal to threshold 1 and the number of second photos is greater than or equal to threshold 2, the mobile phone predicts whether the faces corresponding to the first user ID and the second user ID are the same person's faces based on the first face similarity.

[0136] The thresholds 1 and 2 are not specifically limited and are subject to actual implementation. Threshold 1 and threshold 2 can be the same or different. For example, threshold 1 and threshold 2 can be the same, such as 50.

[0137] In other words, if the number of face photos corresponding to the first user ID cached by the phone is greater than or equal to 50, and the number of face photos corresponding to the second user ID cached by the phone is also greater than or equal to 50, then the phone directly uses the previously cached first face similarity score without recalculating, thus accelerating the prediction process. This is because the phone used a sufficient number of face photos when it previously calculated the first face similarity score, indicating that the result of the previous calculation was highly reliable, and therefore can be used directly this time.

[0138] Optionally, if the number of first photos t1 equals the number of third photos t3, and the number of second photos t2 equals the number of fourth photos t4, then the mobile phone predicts, based on the first face similarity, whether the faces corresponding to the first user ID and the second user ID are the same person's faces.

[0139] It should be noted that the number of first photos t1 equals the number of third photos t3, indicating that no new photos of the face corresponding to the first user ID have been added to the phone's photo library application; similarly, the number of second photos t2 equals the number of fourth photos t4, indicating that no new photos of the face corresponding to the second user ID have been added to the phone's photo library application. Therefore, the phone can directly use the previously calculated first face similarity score without recalculating, thus accelerating the prediction process.

[0140] Optionally, if the number of first photos t1 is greater than or equal to threshold 1, and the number of second photos is greater than or equal to threshold 2; and the number of first photos t1 equals the number of third photos t3, and the number of second photos t2 equals the number of fourth photos t4, then the phone predicts, based on the first face similarity, whether the faces corresponding to the first user ID and the second user ID belong to the same person. In other words, if both conditions 1 and 2 are met simultaneously, it indicates that the credibility of the first face similarity calculated by the phone in the previous calculation is higher, and the phone can be used directly.

[0141] In some embodiments, such as Figure 8 As shown, if the mobile phone does not meet condition 1 above; or, does not meet condition 2 above, then the mobile phone calculates the first face similarity. The specific implementation process for the mobile phone to calculate the first face similarity can be referred to the above. Figure 5 , Figure 6 and Figure 7 The process shown will not be repeated here.

[0142] In summary, using the solution of this application embodiment, if the current judgment result indicates that the phone has a first face similarity score in its cache, the phone can further determine whether condition 1 and / or condition 2 are met. If condition 1 and / or condition 2 are met, it indicates that the first face similarity score cached by the phone is reliable, and the phone can use it directly without repeated calculation, thus achieving the purpose of accelerating prediction.

[0143] It should be noted that the contents described in the various embodiments of this application can be used to explain the technical solutions in other embodiments of this application. The technical features described in each embodiment can also be applied in other embodiments, and new solutions can be formed by combining the technical features in other embodiments. This application only provides an exemplary list of several embodiments for illustration and does not mean that this application is limited thereto.

[0144] This application provides an electronic device that may include a memory, one or more processors, and a display screen. The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device performs the various functions or steps performed by the mobile phone in the above embodiment. The structure of the electronic device can be referred to the above. Figure 1 The structure of the electronic device 100 shown.

[0145] This application also provides a chip system for use in electronic devices. For example... Figure 9 As shown, the chip system 1100 includes at least one processor 1101 and at least one interface circuit 1102. The processor 1101 can be one of the embodiments described above. Figure 1The processor 110 is shown. Based on this, the interface circuit 1102 can be, for example, an interface circuit between the processor 110 and external memory; or an interface circuit between the processor 110 and internal memory.

[0146] The processor 1101 and interface circuit 1102 described above can be interconnected via a line. For example, interface circuit 1102 can be used to receive signals from other devices (e.g., the memory of electronic device 100). As another example, interface circuit 1102 can be used to send signals to other devices (e.g., processor 1101). Exemplarily, interface circuit 1102 can read instructions stored in memory and send those instructions to processor 1101. When the instructions are executed by processor 1101, the electronic device can perform the various functions or steps performed by the mobile phone in the above embodiments. Of course, the chip system may also include other discrete components, and this application embodiment does not specifically limit this.

[0147] This application also provides a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.

[0148] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.

[0149] It should be noted that the terms "first" and "second," etc., in the embodiments and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. "First" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0150] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0151] It should be understood that in this application, "at least one (item)" means one or more. "More than one" means two or more. "At least two (items)" means two or three or more. "And / or" is used to describe the relationship between related objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple. Both "...when" and "if" indicate that a corresponding action will be taken under certain objective circumstances. They are not time limits, nor do they require a judgment action to be taken when the action is taken, nor do they imply any other limitations.

[0152] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.

[0153] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules according to the system, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0154] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0155] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0156] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional unit.

[0157] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0158] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A video generation method based on social relationships, characterized in that, The method is applied to an electronic device, which includes multiple photos, each photo corresponding to a user ID, and each user ID corresponding to a face; the method includes: The electronic device determines whether it has cached a first face similarity score; wherein, the first face similarity score is the face similarity score corresponding to the first user ID and the second user ID among the plurality of user IDs; If the current judgment result indicates that the electronic device currently caches the first face similarity, then the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's face based on the first face similarity. If the faces corresponding to the first user ID and the second user ID are the same person's face, the electronic device assigns a first fusion ID and determines the first social relationship between the device owner and the user corresponding to the first fusion ID; After the electronic device determines the first social relationship between the device owner and the user corresponding to the first fused ID, the electronic device receives a first voice input by the user. The first voice includes information indicating the device owner and information about the existence of the first social relationship with the device owner. In response to the first voice command, the electronic device displays a first photo and a first control on a first interactive interface; wherein the first photo contains the user corresponding to the first fusion ID; In response to a first operation input by the user to the first control, the electronic device displays a thumbnail of the target video; the target video is generated by the electronic device based on a first photo of the user corresponding to the first fusion ID in the gallery application.

2. The method according to claim 1, characterized in that, Before the electronic device determines whether it has cached the first face similarity score, the method further includes: The electronic device periodically determines the first face similarity and caches the determined first face similarity in the electronic device; The first face similarity currently cached by the electronic device, as indicated by the current judgment result, is cached by the electronic device after the first face similarity was determined last time.

3. The method according to claim 1 or 2, characterized in that, The electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's faces based on the first face similarity, including: If the number of first photos is greater than or equal to a first threshold, and the number of second photos is greater than or equal to a second threshold, then the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's face based on the first face similarity; wherein, the number of first photos is the number of photos of the face corresponding to the first user ID cached by the electronic device, and the number of second photos is the number of photos of the face corresponding to the second user ID cached by the electronic device. And / or, If the number of the first photos and the number of the third photos are the same, and the number of the second photos and the number of the fourth photos are the same, then the electronic device predicts whether the faces corresponding to the first user ID and the second user ID are the same person's face based on the first face similarity; wherein, the number of the third photos is the number of photos in the image library application that include the face corresponding to the first user ID, and the number of the fourth photos is the number of photos in the image library application that include the face corresponding to the second user ID.

4. The method according to any one of claims 1-3, characterized in that, The electronic device determines the similarity of the first face, including: The electronic device selects n first target photos from the photos of the faces corresponding to the first user ID, and selects m second target photos from the photos of the faces corresponding to the second user ID; n and m are positive integers. The electronic device determines, based on the n first target photos and the m second target photos, whether the genders of the users corresponding to the first user ID and the second user ID are the same, and whether the age difference is within a preset range. If the genders of the users corresponding to the first user ID and the second user ID are the same, and the age difference is within the preset range, then the electronic device determines the first facial similarity.

5. The method according to claim 4, characterized in that, The method further includes: If the genders of the users corresponding to the first user ID and the second user ID are different; and / or, the age difference is outside the preset range, then the electronic device abandons the determination of the first facial similarity.

6. The method according to any one of claims 2-5, characterized in that, The electronic device determines the similarity of the first face, including: The electronic device extracts x first facial features from the photo of the face corresponding to the first user ID, and extracts y second facial features from the photo of the face corresponding to the second user ID; x and y are positive integers; The electronic device determines the similarity between each of the x first facial features and each of the y second facial features, and obtains z facial feature similarities; z = x × y; The electronic device determines the first face similarity based on the similarity of the z facial features.

7. The method according to claim 6, characterized in that, The method further includes: After obtaining the z facial feature similarities, the electronic device caches the z facial feature similarities. Before the electronic device determines the similarity of the first face again, the electronic device determines whether it currently caches the similarity of the z face features; If the electronic device currently caches the similarity scores of the z facial features, then the electronic device performs the next step of determining the similarity score of the first face based on the currently cached z facial feature similarities.

8. An electronic device, characterized in that, include: Memory and one or more processors; The memory stores computer program code, the computer program code including computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-7.

9. A chip system, characterized in that, include: At least one processor and an interface; The interface is used to receive instructions and transmit them to the at least one processor; The at least one processor executes the instructions to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, include: Computer instructions; When the computer instructions are executed on the electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-7.

11. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Picture grouping method and equipment

    CN111625670A

  • Video generation method and device

    CN116437163A