Human-machine interaction method, and electronic device

By recognizing user actions and automatically playing associated image frames, this technology solves the problem of limited user interaction with images in existing technologies, enabling emotional interaction and convenient control, and enriching the interaction between users and image content.

WO2026157462A1PCT designated stage Publication Date: 2026-07-30HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-11-18
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing user-image interaction methods are simplistic, lacking emotional engagement, intelligence, and fun, making it difficult for users to easily control the switching and playback of image content.

Method used

By recognizing specific user actions, electronic devices automatically play image frames associated with those actions, including multiple frames of static or dynamic images, enabling user interaction with the image content and enhancing emotional engagement and convenience.

Benefits of technology

It enriches the interaction between users and images, enhances the intelligence and fun of the interaction, and allows users to control the changes and playback of image content without touch operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025135651_30072026_PF_FP_ABST
    Figure CN2025135651_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are a human-computer interaction method, and an electronic device. The method is applied to a first device. The method comprises: when a first device displays a first image, the first device identifying a first action of a user; and on the basis of the first action, the first device playing an image frame associated with the first action among a plurality of image frames corresponding to the first image. By means of the method and the electronic device, the interaction between a user and the content of an image displayed on a screen of the electronic device can be realized, the emotional interaction between the user and the image is enhanced, the intelligence and interestingness of the interaction between the user and the image are enhanced, and an alternative interaction manner is also provided for the user when it is inconvenient to perform touch control on the screen.
Need to check novelty before this filing date? Find Prior Art

Description

Human-computer interaction methods and electronic devices

[0001] This application claims priority to Chinese Patent Application No. 202510126274.3, filed on January 27, 2025, entitled "Method and Electronic Device for Human-Computer Interaction", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of human-computer interaction, and more specifically, to a method and electronic device for human-computer interaction. Background Technology

[0003] Currently, when browsing images on a display screen, users can interact with the images by performing corresponding touch operations. For example, multi-finger pinch or spread gestures can trigger image zooming in and out, swipe gestures can switch images, and long-press and swipe gestures or dragging a progress bar gestures can trigger the sequential playback of video content.

[0004] However, the existing forms of interaction between users and images are relatively simple. Summary of the Invention

[0005] This application provides a human-computer interaction method and an electronic device that enables interaction between the user and the content of an image displayed on the electronic device, thereby enhancing the interaction between the user and the image.

[0006] In a first aspect, a human-computer interaction method is provided, the method being applied to a first device, the method comprising: when the first device displays a first image, the first device recognizing a first action of a user; and the first device playing an image frame associated with the first action from among a plurality of image frames corresponding to the first image, based on the first action.

[0007] In this embodiment of the application, the interaction between the user and the content of the image displayed on the screen of the electronic device can be realized, which enhances the emotional interaction between the user and the image, enriches the interaction forms between the user and the image, and enhances the intelligence and fun of the interaction between the user and the image.

[0008] In conjunction with the first aspect, in one possible implementation, the plurality of image frames corresponding to the first image include one or more images in the first image, and / or include images generated based on one or more images in the first image.

[0009] In this embodiment, when the electronic device displays a first image, after recognizing a specific user action, it can display image frames associated with the user's specific action from among the multiple frames corresponding to the first image. These multiple frames corresponding to the first image can be image frames from the first image, or image frames generated based on image frames from the first image. For example, when the first image is a static image, the multiple frames corresponding to the first image can include the static image and multiple frames generated based on the static image that can present dynamic effects. This enhances the intelligence and fun of the interaction between the user and the image.

[0010] In conjunction with the first aspect, in one possible implementation, the user's first action is used to control the content in the first image.

[0011] It can be understood that, from the user's perspective, the user's actions control the changes in the image content displayed on the electronic device. For example, the user might control the subject's hair to be blown by the wind or the subject to blink. From the perspective of the electronic device's implementation, this could involve selecting the image frame corresponding to the user's action from multiple image frames and starting playback or displaying it.

[0012] In this embodiment of the application, the user can control the content of the image displayed on the electronic device to change in relation to the specific action. For example, the user can blink to control the subject in the image to blink or perform other opening and closing actions, which can further enrich the interaction between the user and the image content and enhance the user's interactive experience.

[0013] In conjunction with the first aspect, in one possible implementation, when the first image is a video, the first device plays image frames associated with the first action from among multiple image frames corresponding to the first image, according to the first action. This includes: the first device playing video segments associated with the first action from the video, according to the first action. Optionally, when the electronic device acquires a video, it can also pre-generate multiple generated image frames corresponding to a preset action based on one or more image frames in the video; subsequently, based on user operation, the electronic device can select image frames corresponding to the user action from these multiple generated image frames and image frames of the original video for display or playback.

[0014] In conjunction with the first aspect, in one possible implementation, playing the video segment associated with the first action in the video includes: the first device starting playback from a first image frame, the first image frame being the starting frame of the video segment associated with the first action in the video or the frame preceding the starting frame; the first device stopping playback when playing to a second image frame, the second image frame being the ending frame of the video segment associated with the first action in the video or the frame following the ending frame.

[0015] In this embodiment of the application, when the electronic device displays a video, the user can control the playback of video segments associated with the user's specific actions through specific actions, without requiring the user to control via touch operation. This enables interaction between the user and the video content displayed on the screen of the electronic device, enhances the emotional interaction between the user and the video, and improves the convenience of the user in using the electronic device.

[0016] In conjunction with the first aspect, in one possible implementation, when the first image is a dynamic image, the multiple image frames corresponding to the first image include multiple images in the dynamic image, and / or include at least one image generated based on at least one image in the dynamic image and associated with the first action. The dynamic image described in this application embodiment may also include video.

[0017] In this embodiment of the application, when a dynamic image is displayed on the screen of an electronic device, the user can control the playback of the image frames associated with the user's specific actions among the multiple image frames corresponding to the dynamic image through specific actions. This eliminates the need for the user to control the image through touch operation, enabling interaction between the user and the content of the dynamic image displayed on the screen of the electronic device. This enhances the emotional interaction between the user and the dynamic image and also improves the convenience of the user in using the electronic device.

[0018] In conjunction with the first aspect, in one possible implementation, when the first image is a static image, the plurality of image frames corresponding to the first image include the static image and at least one image generated based on the static image and associated with the first action.

[0019] In this embodiment, when a static image is displayed on the screen of an electronic device, the user can control the playback of multiple image frames corresponding to the static image through specific actions, which are associated with the user's specific actions. This eliminates the need for the user to control the image through touch operation, enabling the static image to present a dynamic effect. This allows for interaction between the user and the content of the static image displayed on the screen of the electronic device, enhancing the emotional interaction between the user and the static image, and improving the convenience of the user in using the electronic device.

[0020] In conjunction with the first aspect, in one possible implementation, playing the image frame associated with the first action from among a plurality of image frames corresponding to the first image includes: the first device cyclically playing the image frame associated with the first action.

[0021] In this embodiment, the electronic device can loop image frames associated with user actions, thereby improving the user's browsing experience.

[0022] In conjunction with the first aspect, in one possible implementation, when the first device displays the first image, the first device identifies the user's first action, which includes: when the first device displays the first image, the first device acquires the user's image information; and the first device identifies the user's first action based on the user's image information.

[0023] In this embodiment, when an image is displayed on the screen of an electronic device, the electronic device can detect the user's image to recognize the user's specific actions, and then control the playback of the image frames associated with the user's specific actions among multiple image frames corresponding to the image. This eliminates the need for the user to control the image through touch operation, enabling interaction between the user and the content of the image displayed on the screen of the electronic device. This enhances the emotional interaction between the user and the image and improves the convenience of the user in using the electronic device.

[0024] In conjunction with the first aspect, in one possible implementation, when the first device displays the first image, the first device identifies the user's first action, which includes: when the first device displays the first image, the first device acquires the user's voice information; and the first device identifies the user's first action based on the user's voice information.

[0025] In this embodiment, when an image is displayed on the screen of an electronic device, the electronic device can detect the user's voice to recognize the user's specific actions, and then control the playback of the image frames associated with the user's specific actions among multiple image frames corresponding to the image. This eliminates the need for the user to control the image through touch operation, enabling interaction between the user and the content of the image displayed on the screen of the electronic device. This enhances the emotional interaction between the user and the image and improves the convenience of the user in using the electronic device.

[0026] In conjunction with the first aspect, in one possible implementation, the first device acquires the user's voice information by: the first device acquiring the user's voice information via a microphone.

[0027] In conjunction with the first aspect, in one possible implementation, when the first device displays the first image, the first device identifies the user's first action, including: when the first device displays the first image, the first device acquires a touch signal on the display screen; the first device identifies the user's first action based on the touch signal.

[0028] In this embodiment of the application, when an image is displayed on the screen of an electronic device, the electronic device can identify the user's specific actions by detecting the touch signals on the display screen, and then control the playback of the image frames associated with the user's specific actions among the multiple image frames corresponding to the image. This enables the interaction between the user and the content of the image displayed on the screen of the electronic device, enhances the emotional interaction between the user and the image, and also improves the convenience of the user in using the electronic device.

[0029] In conjunction with the first aspect, in one possible implementation, when the first device displays the first image, the first device identifies the user's first action by: when the first device displays the first image, the first device acquires any two or three of the user's image information, the user's voice information, and touch signals on the display screen; the first device identifies the user's first action based on the acquired two or three of the information.

[0030] In this embodiment, when an image is displayed on the screen of an electronic device, the electronic device can identify the user's specific actions by detecting any two or more of the user's image, the user's voice information, and the touch signals on the display screen. In turn, it can control the playback of the image frames associated with the user's specific actions among the multiple image frames corresponding to the image. This enables the interaction between the user and the content of the image displayed on the screen of the electronic device, enhances the emotional interaction between the user and the image, and improves the convenience of the user in using the electronic device.

[0031] In conjunction with the first aspect, in one possible implementation, the action of the subject in the image frame associated with the first action is the same as the first action, or the action of the subject in the image frame associated with the first action satisfies a preset correspondence with the first action.

[0032] In this embodiment of the application, the user can make the electronic device play image frames associated with the specific action by performing a specific action. For example, when the electronic device is playing a video, the user can make the electronic device play video segments in the video that are associated with the user's specific action by performing a specific action. This process is fast and convenient, and can enhance the intelligence and fun of the interaction between the user and the image.

[0033] In conjunction with the first aspect, in one possible implementation, the first action is a blowing action, and the image frame associated with the first action is an image frame including an image frame showing the subject's hair being blown by the wind; or the first action is a blowing action, and the image frame associated with the first action is an image frame including an image frame showing the subject being blown by the wind; or the first action is a blowing action, and the image frame associated with the first action is an image frame including an image frame showing the subject in motion; or the first action is an opening and closing of eyes, and the image frame associated with the first action is an image frame including an image frame showing the person opening and closing their eyes; or the first action is a frowning action, and the image frame associated with the first action is an image frame including an image frame showing the subject's hair being blown by the wind. The associated image frame is an image frame that includes the subject frowning; or the first action is a smiling or laughing action, and the image frame associated with the first action is an image frame that includes the subject smiling or laughing; or the first action is a nodding or shaking head action, and the image frame associated with the first action is an image frame that includes the subject nodding or shaking head; or the first action is a jumping, arm-raising, or looking upward action, and the image frame associated with the first action is an image frame that includes the subject moving upward; or the first action is a multi-finger pinching action, and the image frame associated with the first action is an image frame that includes the subject shrinking in size.

[0034] Among them, an image frame that includes a subject moving upwards refers to an image frame in which the subject moves upwards relative to other image elements in the image frame. In other words, the visual perception given to the user is that only the subject moves upwards in the image frame, and other image elements do not move.

[0035] Similarly, an image frame that includes a subject shrinking in size refers to an image frame in which the subject shrinks relative to other image elements in the frame. In other words, the user's visual perception is that only the subject shrinks in the image frame, while the size of other image elements remains unchanged.

[0036] In conjunction with the first aspect, in one possible implementation, before the first device displays the first image, the method further includes: the first device acquiring the first image; the first device generating a plurality of generated image frames corresponding to the first image based on the first image, wherein the plurality of generated image frames corresponding to the first image are at least a portion of the plurality of image frames corresponding to the first image.

[0037] In this embodiment, after acquiring a first image, the electronic device can pre-generate multiple generated image frames corresponding to the first image. Thus, when the electronic device recognizes a user's first action, it can directly retrieve the generated image frame associated with the user's first action from the multiple generated image frames and play it.

[0038] In conjunction with the first aspect, in one possible implementation, at least one first image frame among a plurality of generated image frames corresponding to the first image is associated with the first action; or at least one first image frame among a plurality of generated image frames corresponding to the first image is associated with the first action, and at least one second image frame among a plurality of generated image frames corresponding to the first image is associated with the second action.

[0039] In this embodiment, after the electronic device acquires the first image, it can pre-generate different image frames corresponding to different user actions based on the first image. For example, it can generate an image frame of the subject being blown by the wind in the first image corresponding to a blowing action; and generate an image frame of the subject blinking in the first image corresponding to a blinking action. In this way, when the electronic device displays the first image and detects a user action, it can display an image frame associated with that user action, which can further enrich the interaction between the user and the image content.

[0040] In a second aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store computer program code, and the processor is used to execute the computer program code stored in the memory to implement the method in the first aspect or any possible implementation thereof.

[0041] Thirdly, a computer-readable storage medium is provided, which stores a computer program or instructions that, when executed, implement the method of the first aspect or any possible implementation thereof.

[0042] Fourthly, a chip is provided, wherein instructions are stored that, when executed on a device, cause the chip to perform the methods described in the first aspect or any possible implementation thereof.

[0043] Fifthly, a computer program product is provided, which stores a computer program or instructions that, when executed, implement the method in the first aspect or any possible implementation of the first aspect. Attached Figure Description

[0044] Figure 1 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application;

[0045] Figure 2 is a software structure block diagram of two electronic devices provided in the embodiments of this application;

[0046] Figure 3 is a schematic flowchart of a human-computer interaction method provided in an embodiment of this application;

[0047] Figure 4 is a schematic flowchart of another human-computer interaction method provided in an embodiment of this application;

[0048] Figure 5 is a schematic flowchart of another human-computer interaction method provided in an embodiment of this application;

[0049] Figure 6 is a schematic flowchart of another human-computer interaction method provided in an embodiment of this application;

[0050] Figure 7 is a schematic flowchart of another human-computer interaction method provided in an embodiment of this application;

[0051] Figure 8 is a schematic diagram of a human-computer interaction interface provided in an embodiment of this application;

[0052] Figure 9 is a schematic diagram of another human-computer interaction interface provided in an embodiment of this application;

[0053] Figure 10 is a schematic diagram of another human-computer interaction interface provided in an embodiment of this application;

[0054] Figure 11 is a schematic diagram of another human-computer interaction interface provided in an embodiment of this application;

[0055] Figure 12 is a schematic diagram of another human-computer interaction interface provided in an embodiment of this application;

[0056] Figure 13 is a schematic diagram of the functional modules of a human-computer interaction device provided in an embodiment of this application;

[0057] Figure 14 is a schematic diagram of the functional modules of a human-computer interaction device provided in an embodiment of this application. Detailed Implementation

[0058] The technical solutions of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments.

[0059] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "plural" or "multiple" refers to two or more than two.

[0060] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.

[0061] The terminology used in the following embodiments is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to also include expressions such as “one or more,” unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of this application, “at least one” and “one or more” refer to one, two, or more than two. The term “and / or” is used to describe the relationship between related objects, indicating that three relationships may exist; for example, A and / or B can indicate: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character “ / ” generally indicates that the preceding and following related objects are in an “or” relationship.

[0062] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "one embodiment," "some embodiments," "another embodiment," "other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0063] Currently, when browsing images on electronic devices, users can interact with the images by performing corresponding touch operations on the device's screen. For example, multi-finger pinch or spread gestures can trigger image zooming in and out, swiping can switch images, and long-press and swipe gestures or dragging a progress bar can trigger the sequential playback of video content.

[0064] However, existing forms of user-image interaction fail to meet users' expectations in terms of emotional engagement, intelligence, and fun.

[0065] In some embodiments, users can switch between images displayed on the screen of an electronic device by performing a left or right touch swipe operation on the display interface of the electronic device.

[0066] In some embodiments, when the electronic device displays image 1, the user can control image 1 to be displayed in a smaller size by performing a multi-finger pinch touch operation on the display interface of the electronic device; the user can also control image 1 to be displayed in a larger size by performing a multi-finger spread touch operation on the display interface of the electronic device.

[0067] In some embodiments, when an electronic device is playing a video, the user can control the timing of the video playback by dragging the corresponding slider on the electronic device's display interface.

[0068] In some embodiments, users can trigger the playback of dynamic images or videos by long-pressing the screen with their finger.

[0069] Therefore, in current methods of user-image interaction, the triggered interaction actions are only directed at the image, not at the image content. Different images respond the same way to the same triggered action, lacking interaction between the user's triggered action and the image content. As a result, the interaction between users and images fails to meet users' expectations in terms of emotional interaction, intelligence, and fun.

[0070] Taking video as an example, in current methods, when a user wants to switch from the currently playing screen to the target screen during video playback, the user needs to manually drag the playback progress bar to adjust the progress to the target screen, thereby controlling the target screen in the video playback. It is difficult to quickly and accurately adjust the progress to the target screen, the operation accuracy is low and the operation is complicated, and it is not convenient to switch to the playback target screen.

[0071] In view of this, embodiments of this application provide a method and electronic device for human-computer interaction. Through this method and electronic device, users can interact with the image content displayed on the electronic device, so that the interaction between users and images is no longer limited to the interaction between users and images themselves, but also realizes the interaction between users and the image content displayed on the electronic device, enriching the interaction methods between users and electronic devices, and enhancing the intelligence of the interaction between users and images.

[0072] The methods provided in this application can be applied to electronic devices such as mobile phones, tablets, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application does not impose any restrictions on the specific type of electronic device.

[0073] For example, Figure 1 shows a schematic diagram of the structure of an electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0074] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0075] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0076] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0077] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0078] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0079] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0080] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0081] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0082] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0083] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0084] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0085] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0086] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0087] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0088] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0089] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0090] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0091] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0092] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0093] The SIM card interface 195 is used to connect the SIM card.

[0094] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0095] The electronic device provided in this application embodiment can run an operating system (OS). This operating system can be various operating systems currently used in the industry, such as HarmonyOS based on OpenHarmony; or other operating systems such as Android™, iOS mobile operating system; it can also be various open-source operating systems or their derivatives, such as Linux OS, and other embedded operating systems; or it can be a future new operating system, such as an AI operating system based on artificial intelligence. An operating system is a set of interconnected system software programs that manage and control the operation of electronic devices, utilize and run hardware and software resources, and provide public services to organize user interaction. The operating system occupies a pivotal position in electronic devices, connecting to the physical devices at the hardware layer below and providing a runtime environment for application software above.

[0096] An operating system typically includes a kernel layer, a middleware layer, and an application layer. The application layer includes applications, which can include system applications and third-party applications. The middleware layer is a set of software, or frameworks, that provides various services to application developers, such as databases, multimedia, and graphics, or capabilities like distributed scheduling and system expansion. For example, the middleware layer can also be broadly divided into a framework layer and / or a system service layer. The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The system service layer includes the core capabilities of the system, providing services to applications through the framework layer. The kernel layer is the layer between hardware and software. The kernel layer can include hardware drivers and the operating system kernel. In addition to providing hardware drivers, the kernel layer also supports functions such as memory management and system process management.

[0097] The electronic devices we use in our daily lives come in various types and forms, and are applied in a wide range of scenarios. Therefore, based on the different forms and functions of electronic devices, different application scenarios, and different user needs, the operating systems used in these devices may also differ. These operating systems share commonalities but also have their own unique characteristics. Different operating systems affect user experience, application ecosystem, and system performance. The basic functions implemented by the electronic device provided in this application can be achieved using a general-purpose operating system or a dedicated operating system.

[0098] To more clearly illustrate the implementation of the embodiments of this application under a specific operating system, Figure 2(a) illustrates the architecture of HarmonyOS, and those skilled in the art can deduce the implementation of the embodiments of this application under other specific operating systems, such as the Android™ operating system.

[0099] As shown in Figure 2(a), the software architecture of an electronic device can be divided into several layers. In some embodiments, from bottom to top, these layers are: kernel layer, system service layer, framework layer, and application layer. The layers communicate with each other through software interfaces. System functions can be tailored, added, or combined at the subsystem granularity in different device deployment scenarios. Each subsystem can also be tailored, added, or combined at the functional granularity.

[0100] (1) Kernel layer

[0101] The Kernel Abstraction Layer (KAL) provides basic kernel capabilities to upper layers by shielding the differences between multiple kernels, including but not limited to process / thread management, memory management, file system, network management, and peripheral device management.

[0102] Kernel Subsystem: Supports the selection of a suitable OS kernel for different resource-constrained devices, including but not limited to Linux kernel, HarmonyOS kernel, LiteOS, etc.

[0103] Driver Subsystem: The driver framework is the foundation for the open system hardware ecosystem, providing unified peripheral access capabilities and a framework for driver development and management. The driver framework includes: display drivers, camera drivers, audio drivers, Bluetooth drivers, sensor drivers, etc.

[0104] (2) System Service Layer

[0105] The system service layer comprises the core capabilities of the system, providing services to applications through the framework layer. This layer includes, but is not limited to, the following subsystems:

[0106] The system's basic capability subsystems provide fundamental capabilities for the operation, scheduling, and migration of distributed applications across multiple devices. For example, they may include distributed soft bus, distributed data management, distributed task scheduling, and the Ark multi-language runtime. They may also include multi-modal input subsystems, graphics subsystems, security subsystems, and AI subsystems.

[0107] Basic software service subsystems: provide common and general software services; for example, event notification subsystem, telephone service subsystem, multimedia subsystem, etc.

[0108] Enhanced software service subsystem suite: Provides differentiated capability-enhancing software services for different devices; for example, it may include proprietary business subsystems for smart screens, wearable devices, and IoT devices.

[0109] Hardware service subsystem set: provides hardware services; for example, it may include location service subsystem, user IAM (Identity and Access Management) subsystem, wearable proprietary hardware service subsystem, biometric identification, IoT proprietary hardware service subsystem, etc.

[0110] Distributed task scheduling enables distributed service management (discovery, synchronization, registration, and invocation), supporting remote startup, remote invocation, remote connection, and migration of applications across devices.

[0111] Distributed data management enables data synchronization, data storage, data sharing, and data access across all scenarios and devices.

[0112] The distributed soft bus provides communication-related capabilities for seamless interconnection between multiple devices, including: WLAN service capabilities, Bluetooth service capabilities, soft bus, inter-process communication RPC (Remote Procedure Call) and other communication capabilities.

[0113] Ark Multilingual Runtime is a unified compilation runtime platform designed to support the joint compilation and execution of multiple programming languages ​​and multiple chip platforms.

[0114] (3) Framework layer

[0115] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. Examples include the ArkUI framework (which provides a complete infrastructure for UI development of system applications, including UI functionalities such as components, layouts, animations, and interactive events, as well as a real-time interface preview tool), the user application framework, and the Ability framework (an Ability is a lightweight application; the Ability framework schedules and manages the operation and lifecycle of Abilities). Different devices may run different operating systems, and therefore support different APIs.

[0116] The HarmonyOS API is a series of open capabilities provided to support HarmonyOS application development. The HarmonyOS API can be set at the framework layer or independently of the framework layer. Examples include: Audio API (audio service), Push API (push service), and Account API (account service).

[0117] (4) Application layer

[0118] Applications can include system apps and extended / third-party apps. System apps can include the desktop, control bar, settings, contacts, phone, camera, etc., while extended / third-party apps can include social apps, travel apps, etc.

[0119] Figure 2(b) is a software structure block diagram of another electronic device 100 provided in an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.

[0120] As shown in Figure 2(b), the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0121] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0122] As shown in Figure 2(b), the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.

[0123] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0124] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0125] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0126] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0127] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0128] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0129] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0130] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0131] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0132] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0133] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0134] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0135] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0136] A 2D graphics engine is a graphics engine for 2D drawing.

[0137] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0138] It should be understood that the technical solutions in the embodiments of this application can be used in systems such as Android, iOS, and HarmonyOS.

[0139] The technical solutions of this application embodiment can be applied to electronic devices with display capabilities, for example. Exemplarily, they can be applied to televisions, desktop computers, laptops, in-vehicle screens, and portable electronic devices such as mobile phones, foldable screens, tablets, cameras, camcorders, video recorders, watches, and wristbands. They can also be applied to other electronic devices with display capabilities, and further to electronic devices in future networks or in future evolved public land mobile networks (PLMNs).

[0140] For example, Figure 3 shows a schematic flowchart of a human-computer interaction method 300 provided in an embodiment of this application. The method provided in various embodiments of this application can also be referred to as an image display method. As shown in Figure 3, the method 300 includes:

[0141] S301: When the electronic device displays the first image, the electronic device recognizes the user's first action.

[0142] The embodiments of this application can be applied to scenarios where users view or browse images on electronic devices, such as, but not limited to, browsing the image gallery, wallpaper, lock screen, etc. of electronic devices.

[0143] The user's first action can be a facial expression and / or a hand gesture.

[0144] In some embodiments, a user’s facial movements may include actions such as blowing air, blinking, raising eyebrows, blinking with one eye, pouting, nodding up and down or shaking the head left and right, shaking the head, rolling the eyes, opening the mouth, chewing, staring, smiling or laughing.

[0145] In some embodiments, the user's gestures may include actions such as finger tapping, multi-finger pinching (e.g., two-finger pinching, three-finger pinching, five-finger pinching, etc.), multi-finger spreading, finger sliding, finger circling, etc., and may also include actions made by the user through the arm, elbow, etc., which are not limited in this application embodiment.

[0146] In some embodiments, when the electronic device displays the first image, the electronic device acquires the user's image information and identifies the user's first action based on the user's image information.

[0147] In one implementation, the electronic device acquires the user's image information in real time while displaying the first image.

[0148] More specifically, electronic devices can acquire user facial expression image information and / or gesture image information in real time based on the user's image information, and can further recognize the user's facial movements based on the user's facial expression image information, and / or recognize the user's hand gestures based on the user's gesture image information.

[0149] In some embodiments, an image sensor, such as a camera, is installed on the electronic device, which can acquire image information of the user in real time. The image sensor may be turned on, for example, when it detects that a first image is displayed on the screen of the electronic device.

[0150] In some embodiments, the image sensor may be, for example, a low-power image sensor placed on the front of the electronic device.

[0151] In some embodiments, when the electronic device displays the first image, the electronic device acquires the user's voice information; the electronic device identifies the user's first action based on the user's voice information.

[0152] In one implementation, the electronic device acquires the user's voice information in real time while displaying the first image.

[0153] For example, a user's voice information can be obtained through the microphone of an electronic device.

[0154] For example, when the user's voice information is the sound of blowing air, the user's first action can be identified as blowing air; when the user's voice information is the sound of coughing, the user's first action can be identified as coughing; in addition, the user's voice information can also be a sound used to indicate the first action, for example, if the content of the voice information is "look up", that is, if the user says "look up", then the user's first action is to look up; as another example, if the content of the voice information is "blow air", then the user's first action is to blow air.

[0155] In some embodiments, when the electronic device displays a first image, the electronic device acquires touch signals on the display screen of the electronic device; the electronic device identifies a user's first action based on the acquired touch signals.

[0156] In one implementation, the electronic device acquires the touch signal on the display screen, which can be done by acquiring the touch signal on the display screen in real time while displaying the first image.

[0157] For example, when the touch signal is a two-finger swipe signal, the user's first action can be recognized as a two-finger swipe; when the touch signal is a circle drawing signal, the user's first action can be recognized as a circle drawing.

[0158] In some embodiments, when the electronic device displays the first image, the electronic device acquires any two or three of the following information; the electronic device identifies the user's first action based on the acquired two or three information:

[0159] (1) User's image information.

[0160] (2) User's voice information.

[0161] (3) Touch signals on the display screen.

[0162] The electronic device displays a first image, such as playing a video on the screen; it can also display a dynamic image, such as playing a dynamically displayed theme wallpaper; it can also display a picture, such as a large image corresponding to a picture in the electronic device's gallery; it can also display a thumbnail of the first image; or, if the first image is a video or dynamic image, it can also be the first image currently in a static display, such as a frame of a video or dynamic image.

[0163] Here, "user" refers to a user of an electronic device, such as a user facing the screen of an electronic device. When the electronic device recognizes the user's first action based on the user's image information, the user's position is within the image collection range corresponding to the image sensor of the electronic device.

[0164] In some embodiments, the image sensor on the electronic device can acquire the user's image information in real time at a certain frame rate and resolution.

[0165] The first image may include, for example, a picture, a video, an animated picture, etc. The first image may be a captured image, a downloaded image, an image sent by another device, or an image generated by artificial intelligence (AI).

[0166] S302: The electronic device plays the image frame associated with the first action from among the multiple image frames corresponding to the first image, based on the user's first action.

[0167] In some embodiments, the electronic device may loop image frames associated with the first action.

[0168] In some embodiments, the plurality of image frames corresponding to the first image include one or more images in the first image, and / or include images generated based on one or more images in the first image.

[0169] In some embodiments, the first image may be, for example, an image captured by an electronic device; the first image may also be an image downloaded by the electronic device; the first image may also be an image sent by another device and received by the electronic device.

[0170] In one implementation, after acquiring the first image, the electronic device can pre-generate multiple image frames corresponding to the first image. Thus, when the electronic device recognizes the user's first action, it can directly retrieve the image frame associated with the user's first action from the multiple image frames corresponding to the first image and play it.

[0171] In some embodiments, the electronic device can preset multiple user actions, such as blowing air, blinking, frowning, smiling, shaking head, nodding, rolling eyes, pouting, etc. After the electronic device acquires the first image, it can generate multiple generated image frames corresponding to the preset actions based on the content in the first image. Different generated image frames can correspond to different preset actions. For example, blowing air corresponds to the image frame of the subject's hair being blown by the wind, and blinking corresponds to the image frame of the subject blinking.

[0172] The multiple generated image frames can be at least a portion of the multiple image frames corresponding to the first image.

[0173] In some embodiments, when the first image is a video, if the types of main actions in the multiple image frames of the video are abundant (e.g., more than 50 types), then after the electronic device acquires the first image, it does not need to generate multiple generated image frames corresponding to the preset actions based on the content of the first image. Instead, it can directly search for and play image frames corresponding to the user's actions from the multiple image frames of the video based on the user's actions. Optionally, if the types of main actions in the multiple image frames of the video are not many, after acquiring the video, the electronic device can also generate multiple generated image frames corresponding to the preset actions based on one or more image frames in the video. These multiple generated image frames are used to subsequently play image frames corresponding to the user's actions.

[0174] In some embodiments, when the first image is a video, the electronic device plays an image frame associated with the first action from among a plurality of image frames corresponding to the first image, according to the user's first action, including: the electronic device plays a video segment associated with the first action from the video according to the user's first action.

[0175] In one implementation, playing a video segment associated with the first action may include: the electronic device starting playback from a first image frame, which is the starting frame of the video segment associated with the first action, or the previous frame of the starting frame, or several frames before the starting frame; the electronic device stopping playback when playing to a second image frame, which is the ending frame of the video segment associated with the first action, or the next frame of the ending frame, or several frames after the ending frame.

[0176] In some embodiments, when the first image is a dynamic image, the plurality of image frames corresponding to the first image include a plurality of images in the dynamic image, and / or include at least one image associated with the first action generated based on at least one image in the dynamic image.

[0177] In some embodiments, when the first image is a static image, the plurality of image frames corresponding to the first image include the static image and at least one image generated based on the static image and associated with the first action.

[0178] In some embodiments, the electronic device has a preset correspondence, which is the correspondence between the user's actions and the actions of the subject in the first image. For example, it may include the correspondence between the user's facial actions and / or gesture actions and the actions of the subject in the first image. The gesture actions mentioned here can be air gesture actions relative to the screen of the electronic device, in which the user does not touch the screen of the electronic device during the execution of the gesture action, or they can be gesture operations on the screen of the electronic device, in which the user touches the screen of the electronic device during the execution of the gesture action.

[0179] In some embodiments, the action of the subject in the image frame associated with the first action is the same as or similar to the first action.

[0180] In one implementation, the electronic device determines the user's first action as the target action of the subject in the first image based on the user's first action, and then determines the image frame associated with the user's first action from multiple image frames corresponding to the first image (that is, determines the image frame corresponding to the target action of the subject in the first image). Then, the electronic device plays the image frame associated with the user's first action (that is, plays the image frame corresponding to the target action of the subject in the first image).

[0181] In some embodiments, the action of the subject in the image frame associated with the first action satisfies a preset correspondence with the first action.

[0182] In one implementation, the electronic device determines the target action of the subject in the first image based on a preset correspondence and the user's first action. Then, it determines the image frame associated with the user's first action from multiple image frames corresponding to the first image (that is, the image frame corresponding to the target action of the subject in the first image). The electronic device then plays the image frame associated with the user's first action (that is, plays the image frame corresponding to the target action of the subject in the first image).

[0183] In one example, the first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes a person's hair being blown by the wind.

[0184] In one example, the first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes the animation of the subject being blown.

[0185] In one example, the first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes a picture of the subject in motion.

[0186] In one example, the first action is the action of opening and closing the eyes, and the image frame associated with the first action is an image frame that includes the image of a person opening and closing their eyes.

[0187] In one example, the first action is the action of frowning, and the image frame associated with the first action is an image frame that includes the subject frowning.

[0188] In one example, the first action is the action of laughing, and the image frame associated with the first action is an image frame that includes the subject laughing.

[0189] In one example, the first action is a nodding or shaking motion, and the image frame associated with the first action is an image frame that includes a scene of the subject nodding or shaking their head.

[0190] In one example, the first action is the action of jumping, raising an arm, or looking upwards, and the image frame associated with the first action is an image frame that includes a picture of the subject moving upwards or flying upwards.

[0191] In this context, an image frame that includes a subject moving upwards refers to an image frame in which the subject moves upwards relative to all other image elements in the frame. In other words, the user's visual perception is that only the subject moves upwards, while other image elements remain stationary. Similarly, an image frame that includes a subject flying upwards refers to an image frame in which the subject flies upwards relative to all other image elements in the frame. In other words, the user's visual perception is that only the subject flies upwards, while other image elements remain stationary.

[0192] In one example, the first action is the action of widening the eyes, and the image frame associated with the first action is an image frame that includes a picture of the subject being frightened.

[0193] In one example, the first action is a multi-finger pinching motion, and the image frame associated with the first action is an image frame that includes a scene where the subject changes from large to small.

[0194] Among them, the image frame in which the subject changes size refers to the image frame in which the subject changes size relative to other image elements in the image frame. In other words, the visual perception given to the user is that only the subject changes size in the image frame, while the size of other image elements does not change.

[0195] In some embodiments, the user can further control the playback of image frames associated with the first action through further actions. For example, the user can control the switching between image frames associated with the first action by sliding left or right or up or down; the user can also control the zooming out of the image frames associated with the first action by pinching with multiple fingers; the user can also control the zooming out of the image frames associated with the first action by spreading multiple fingers; when the image frames associated with the first action are zoomed out, the user can also view a partial image of the zoomed-out image by smoothing left or right; the user can also control the playback sequence of the image frames associated with the first action by smoothing left or right.

[0196] In some embodiments, when the electronic device displays the second image, the electronic device recognizes the user's first action; based on the user's first action, the electronic device plays an image frame associated with the first action from among a plurality of image frames corresponding to the second image, wherein the first image and the second image are different.

[0197] Therefore, it can be understood that the response to the same action can be different for different images. For example, when the electronic device displays the first image, the response to the first action is to play the image frame associated with the first action from among the multiple image frames corresponding to the first image. For example, it could play the image frame of the subject's hair being blown by the wind. When the electronic device displays the second image, the response to the first action is to play the image frame associated with the first action from among the multiple image frames corresponding to the second image. For example, it could play the image frame of the subject moving.

[0198] In some embodiments, the user's first action is used to control the content in the first image.

[0199] It can be understood that, from the user's perspective, the user's actions control the changes in the image content displayed on the electronic device. For example, the user controls the hair of the subject in the image to be blown up by the action, or the user controls the subject in the image to blink by the action. From the perspective of the electronic device's implementation, it identifies the image frame corresponding to the user's action from multiple image frames and starts playing it.

[0200] In this embodiment, when an image is displayed on the electronic device, the electronic device can acquire the user's actions, such as facial expressions and / or gestures, in real time. Based on the recognized facial expressions and / or gestures, the electronic device can play image frames associated with the user's facial expressions and / or gestures from among the multiple frames of the image displayed on the screen. For example, it can automatically play still images, moving images, or videos, retrieve corresponding frames from moving images or videos, and play them. It can also perform zooming, shrinking, or offset operations on the main subject in the played image, enabling the user to trigger interaction with the image content displayed on the electronic device through air gestures or facial expressions. This makes the interaction between the user and the image no longer limited to the interaction between the user and the image itself, but also enables interaction between the user and the image content displayed on the electronic device. This enriches the interaction methods between the user and the electronic device, provides more alternative interaction methods for users when it is inconvenient to touch the screen, and enhances the intelligence and fun of the interaction between the user and the image.

[0201] For example, Figure 4 shows a schematic flowchart of another human-computer interaction method 400 provided in an embodiment of this application. As shown in Figure 4, the method 400 includes:

[0202] S401: When the electronic device displays the first image, the electronic device acquires the user's image information.

[0203] In one implementation, the electronic device acquires the user's image information in real time while displaying the first image.

[0204] In some embodiments, an image sensor is installed on the electronic device, which can acquire image information of the user in real time. The image sensor may be turned on, for example, when it detects that a first image is displayed on the screen of the electronic device.

[0205] In some embodiments, the image sensor may be, for example, a low-power image sensor placed on the front of the electronic device.

[0206] S402: Electronic devices recognize the user's first action based on the user's image information.

[0207] For example, an electronic device can acquire the user's facial expression image information and / or gesture image information in real time based on the user's image information, and can further recognize the user's facial movements based on the user's facial expression image information, and / or recognize the user's hand gestures based on the user's gesture image information.

[0208] In some embodiments, electronic devices may use AI algorithms, for example, to identify a user's first action based on the user's image information.

[0209] Specifically, for example, the NPU or GPU of an electronic device can use AI algorithms to recognize the user's first action based on the user's image information.

[0210] The explanation of the user's first action is described in detail in S301 of the embodiment shown in Figure 3, and will not be repeated here for the sake of brevity.

[0211] S403: The electronic device plays the image frame associated with the first action from among the multiple image frames corresponding to the first image, based on the user's first action.

[0212] The explanation of this step is the same as that of S302 in the embodiment shown in Figure 3, and will not be repeated here for the sake of brevity.

[0213] In some embodiments, the electronic device retrieves an image frame associated with the first action from a plurality of image frames corresponding to the first image based on the user's first action, and then the electronic device plays the image frame associated with the first action. For example, the electronic device jumps from playing the current image frame to playing the image frame associated with the first action.

[0214] In some embodiments, when the first image is a video, the image frames associated with the first action can form one or more video segments associated with the first action. The electronic device switching from playing the current image frame to playing the image frame associated with the first action can mean that the electronic device switches from the current screen in the video to start playing from the one or more video segments associated with the first action, or it can mean that the electronic device switches from the current screen in the video to looping the one or more video segments associated with the first action.

[0215] In one example, the user's first action is blowing air, and the electronic device jumps from the current screen of the playing video to the screen corresponding to the person's hair being blown by the wind or the main body being blown by the wind.

[0216] In a specific example, taking the first image as the first video, the main subject of the first video is a large tree, and the duration of the first video is 30 seconds. In the video segment corresponding to the time period from 1 second to 10 seconds, the large tree is in a static state. In the video segment corresponding to the time period from 11 seconds to 20 seconds, the large tree is also in a static state. In the video segment corresponding to the time period from 21 seconds to 30 seconds, the large tree is swaying in the wind. The screen of the current electronic device is playing the image corresponding to the 4th second. In this case, when the electronic device recognizes the user's blowing action, it will switch the image of the first video played by the electronic device from the image corresponding to the 4th second to start playing from the 21st second.

[0217] In one implementation, when the electronic device detects the user's blowing action, the screen of the first video played by the electronic device is switched from the screen corresponding to the 4th second to the video segment corresponding to the time period from the 21st second to the 30th second, which is played in a loop starting from the 21st second.

[0218] The explanations of the first image and the multiple image frames corresponding to the first image have been described in detail in the embodiment shown in Figure 3, and will not be repeated here for the sake of brevity.

[0219] In this embodiment of the application, when an image is displayed on the screen of an electronic device, the user can control the image frames associated with the user's actions among multiple image frames corresponding to the image by making corresponding facial expressions or air gestures. This makes the interaction between the user and the image no longer limited to the interaction between the user and the image itself, but also realizes the interaction between the user and the image content displayed by the electronic device, enhancing the emotional interaction between the user and the image, and improving the convenience of the user in using the electronic device.

[0220] For example, Figure 5 shows a schematic flowchart of another human-computer interaction method 500 provided in an embodiment of this application. As shown in Figure 5, the method 500 includes:

[0221] S501: When the electronic device displays the first image, the electronic device acquires the user's voice information.

[0222] In one implementation, the electronic device acquires the user's voice information in real time while displaying the first image.

[0223] In some embodiments, electronic devices may acquire a user's voice information via a microphone.

[0224] S502: Electronic devices recognize the user's first action based on the user's voice information.

[0225] In some embodiments, electronic devices may use AI algorithms, for example, to identify a user's first action based on the user's voice information.

[0226] Specifically, for example, the NPU or GPU of an electronic device can use AI algorithms to recognize the user's first action based on the user's voice information.

[0227] The explanation of the user's first action is described in detail in S301 of the embodiment shown in Figure 3, and will not be repeated here for the sake of brevity.

[0228] S503: The electronic device plays the image frame associated with the first action from among the multiple image frames corresponding to the first image, based on the user's first action.

[0229] The explanation of this step is the same as the explanation of S302 in the embodiment shown in Figure 3 and the explanation of S403 in the embodiment shown in Figure 4. For the sake of brevity, it will not be repeated here.

[0230] For example, Figure 6 shows a schematic flowchart of another human-computer interaction method 600 provided in an embodiment of this application. As shown in Figure 6, the method 600 includes:

[0231] S601: When the electronic device displays the first image, the electronic device acquires the touch signal on the display screen of the electronic device.

[0232] In one implementation, the electronic device acquires the touch signal on the display screen, which can be done by acquiring the touch signal on the display screen in real time while displaying the first image.

[0233] In some embodiments, the electronic device can obtain touch signals from the display screen through a touch module.

[0234] S602: Electronic devices recognize the user's first action based on touch signals on the display screen.

[0235] The explanation of the user's first action is described in detail in S301 of the embodiment shown in Figure 3, and will not be repeated here for the sake of brevity.

[0236] S603: The electronic device plays the image frame associated with the first action from among the multiple image frames corresponding to the first image, based on the user's first action.

[0237] The explanation of this step is the same as the explanation of S302 in the embodiment shown in Figure 3 and the explanation of S403 in the embodiment shown in Figure 4. For the sake of brevity, it will not be repeated here.

[0238] For example, Figure 7 shows a schematic flowchart of another human-computer interaction method 700 provided in an embodiment of this application. As shown in Figure 7, the method 700 includes:

[0239] S701: When the electronic device displays the first image, the electronic device acquires any two or three of the following: the user's image information, the user's voice information, and the touch signal on the electronic device's display screen.

[0240] S702: The electronic device recognizes the user's first action based on any two or three of the acquired image information of the user, the user's voice information, and the touch signal on the display screen of the electronic device.

[0241] The explanation of the user's first action is described in detail in S301 of the embodiment shown in Figure 3, and will not be repeated here for the sake of brevity.

[0242] S703: The electronic device plays the image frame associated with the first action from among the multiple image frames corresponding to the first image, based on the user's first action.

[0243] The explanation of this step is the same as the explanation of S302 in the embodiment shown in Figure 3 and the explanation of S403 in the embodiment shown in Figure 4. For the sake of brevity, it will not be repeated here.

[0244] In some embodiments, the electronic device may also identify the user’s first action by combining other information, such as pressure signals acting on the electronic device or on peripheral devices of the electronic device.

[0245] For example, Figure 8 shows a schematic diagram of a human-computer interaction interface provided in an embodiment of this application.

[0246] As shown in Figure 8, video 810 is playing on the display screen of electronic device 800. The duration of video 810 is 50 seconds. Video 810 includes person 1, and in one or more video clips in video 810, person 1's hair is blown by the wind.

[0247] In one example, in video 810, during the video segment from 00:34 to 00:50, Person 1's hair is blown by the wind, while during the video segment from 00:00 to 00:34, Person 1's hair is not blown by the wind.

[0248] In one scenario, the electronic device 800 can, for example, identify the user's blowing action from the user image output by the front-facing image sensor; then determine the start and end frames of the video segment corresponding to the user's blowing action from the video 810 (e.g., the start and end frames of the video segment of person 1's hair being blown by the wind); and then the electronic device can automatically play and / or loop the video segment corresponding to the user's blowing action (e.g., the video segment of person 1's hair being blown by the wind) on the display screen.

[0249] In another scenario, the electronic device 800 may, for example, identify the user's blowing action from the user image output by the front-facing image sensor; then determine the starting frame of the video segment corresponding to the user's blowing action from the video 810 (e.g., the starting frame of the video segment where person 1's hair is blown by the wind); and then the electronic device may automatically play the video segment corresponding to the user's blowing action through the display screen (e.g., jump to the starting frame of the video segment where person 1's hair is blown by the wind to start playing the video).

[0250] As shown in Figure 8(a), the electronic device 800 displays a video frame corresponding to 00:02, in which the hair of person 1 is not blown by the wind. At this time, if the electronic device 800 detects the user's blowing action, the electronic device 800 will switch the video frame displayed on the screen from the video frame corresponding to 00:02 to the video frame corresponding to 00:34, and continue to play the video segment corresponding to the duration from 00:34 to 00:50 starting from 00:34 (as shown in Figure 8(b)). That is to say, the display of the electronic device 800 switches from playing the video frame corresponding to 00:02 to playing the video segment corresponding to the duration from 00:34 to 00:50, that is, switches to playing the video segment of person 1's hair being blown by the wind.

[0251] In some examples, when the user's blowing action is detected, the display screen of the electronic device 800 switches from playing the video corresponding to 0:02 to looping the video segment corresponding to the duration from 00:34 to 00:50, that is, switching to looping the video segment of person 1's hair being blown by the wind.

[0252] In some embodiments, when there are video clips in video 810 showing multiple people 1 with their hair being blown by the wind, the display screen of electronic device 800 switches from playing the current video to playing the video clips of the multiple people 1 with their hair being blown by the wind, for example, by looping the video clips of the multiple people 1 with their hair being blown by the wind.

[0253] In some embodiments, when there is a video clip in video 810 showing the hair of person 1 being blown by the wind, and also a video clip showing the hair of person 2 being blown by the wind, the display screen of electronic device 800 can switch from playing the current video screen to playing the video clip showing the hair of person 1 being blown by the wind and / or the video clip showing the hair of person 2 being blown by the wind. For example, it can loop the video clip showing the hair of person 1 being blown by the wind and / or the video clip showing the hair of person 2 being blown by the wind.

[0254] For example, Figure 9 shows a schematic diagram of another human-computer interaction interface provided in an embodiment of this application.

[0255] As shown in Figure 9, video 910 is playing on the display screen of electronic device 900. The duration of video 910 is 30 seconds. Video 910 includes subject 1, and in one or more video segments of video 910, subject 1 is blown by the wind.

[0256] In one example, in video 910, in the video segment corresponding to the duration of 00:20 to 00:30, subject 1 is blown by the wind, while in the video segment corresponding to the duration of 00:00 to 00:20, subject 1 is not blown by the wind.

[0257] In one scenario, the electronic device 900 can, for example, identify the user's blowing action from the user image output by the front-facing image sensor; then determine the start and end frames of the video segment corresponding to the user's blowing action from the video 910 (e.g., the start and end frames of the video segment of the main body 1 being blown by the wind); and then the electronic device can automatically play and / or loop the video segment corresponding to the user's blowing action (e.g., the video segment of the main body 1 being blown by the wind) on the display screen.

[0258] In another scenario, the electronic device 900 may, for example, identify the user's blowing action from the user image output by the front-facing image sensor; then determine the starting frame of the video segment corresponding to the user's blowing action from the video 910 (e.g., the starting frame of the video segment in which the main body 1 is blown by the wind); and then the electronic device 900 may automatically play the video segment corresponding to the user's blowing action through the display screen (e.g., jump to the starting frame in which the video is played from the beginning of the frame in which the main body 1 is blown by the wind).

[0259] As shown in Figure 9(a), the electronic device 900 displays a video frame corresponding to 00:01, in which the main body 1 is not blown by the wind. At this time, if the electronic device 900 detects the user's blowing action, the electronic device 900 will switch the video frame displayed on the screen from the video frame corresponding to 00:01 to the video frame corresponding to 00:20, and continue to play the video segment corresponding to the duration of 00:20 to 00:30 starting from 00:20 (as shown in Figure 9(b)). That is to say, the display of the electronic device 900 switches from playing the video frame corresponding to 00:01 to playing the video segment corresponding to the duration of 00:20 to 00:30, that is, switches to playing the video segment of the main body 1 being blown by the wind.

[0260] In some examples, when the user's blowing action is detected, the display screen of the electronic device 900 switches from playing the video corresponding to 0:01 to looping the video segment corresponding to the duration from 00:20 to 00:30, that is, switching to looping the video segment of the main body 1 being blown by the wind.

[0261] In some embodiments, when there are multiple video clips of multiple subjects 1 being blown by the wind in the video 910, the display screen of the electronic device 900 switches from playing the current video screen to playing the video clips of the multiple subjects 1 being blown by the wind, for example, by looping the video clips of the multiple subjects 1 being blown by the wind.

[0262] In some embodiments, when there is a video clip in video 910 where the subject 1 is blown by the wind and there is also a video clip where the hair of person 1 is blown by the wind, the display screen of electronic device 900 can switch from playing the current video screen to playing the video clip where the subject 1 is blown by the wind and the video clip where the hair of person 1 is blown by the wind. For example, it can loop the video clip where the subject 1 is blown by the wind and the video clip where the hair of person 1 is blown by the wind.

[0263] For example, Figure 10 shows a schematic diagram of another human-computer interaction interface provided in an embodiment of this application.

[0264] As shown in Figure 10, video 1010 is playing on the display screen of electronic device 1000. The duration of video 1010 is 40 seconds. Video 1010 includes subject 2, and subject 2 is in a static state in one or more video segments of video 1010.

[0265] In one example, in video 1010, subject 2 is in motion in the video segments corresponding to durations of 00:10 to 00:20 and 00:25 to 00:35; and subject 2 is stationary in the video segments corresponding to durations of 00:00 to 00:10, 00:20 to 00:25, and 00:35 to 00:40.

[0266] In one scenario, the electronic device 1000 can, for example, identify the user's blowing action or eye-opening / closing action from the user image output by the front-facing image sensor; then, it can determine the start and end frames of the video segment corresponding to the user's blowing action or eye-opening / closing action from the video 1010 (e.g., the start and end frames of the video segment in motion of the main body 2); then, the electronic device 1000 can automatically play and / or loop the video segment corresponding to the user's blowing action or eye-opening / closing action (e.g., the video segment in motion of the main body 2) through the display screen.

[0267] In another scenario, the electronic device 1000 may, for example, identify the user's blowing action or eye-opening / closing action from the user image output by the front-facing image sensor; then determine the starting frame of the video segment corresponding to the user's blowing action or eye-opening / closing action from the video 1010 (e.g., the starting frame of the video segment in which the main body 2 is in motion); then the electronic device 1000 may automatically play the video segment corresponding to the user's blowing action or eye-opening / closing action through the display screen (e.g., jump to the starting frame from which the video begins to play from the starting frame in which the main body 2 is in motion).

[0268] As shown in Figure 10(a), the electronic device 1000 displays a video frame corresponding to 00:36, in which the main body 2 is stationary. At this time, as shown in Figure 10(b), if the electronic device 1000 detects the user's blowing motion or eye-opening / closing motion through image acquisition and recognition, the electronic device 1000 switches the video frame displayed on the screen from the video frame corresponding to 00:36 to the video frame corresponding to 00:10, and continues playing from 00:10 to 00:20. The video segment corresponding to the length of the video is played. That is to say, the display screen of the electronic device 1000 switches from playing the video screen corresponding to 0:36 to playing the video segment corresponding to the length of 00:10 to 00:20, that is, switching to playing the first video segment in which the main body 2 is in motion; further, as shown in Figure 10(c), after the electronic device 1000 finishes playing the video segment corresponding to the length of 00:10 to 00:20, it jumps to playing the video segment corresponding to the length of 00:25 to 00:30, that is, switching to playing the second video segment in which the main body 2 is in motion.

[0269] In some examples, when the user's blowing action is detected, the display screen of the electronic device 1000 switches from playing the video corresponding to 0:36 to looping the video segments corresponding to 00:10 to 00:20 and the video segments corresponding to 00:25 to 00:30.

[0270] For example, Figure 11 shows a schematic diagram of another human-computer interaction interface provided in an embodiment of this application.

[0271] As shown in Figure 11(a), the display screen of electronic device 1100 shows image 1110, which includes a tree and subject 3. The position of subject 3 is position A, and subject 3 can be, for example, a balloon.

[0272] As shown in Figure 11(b), when the electronic device 1100 detects the user's actions of jumping, raising their arm, or looking up, the electronic device 1100 controls the main body 3 in the image 1110 to move upward relative to other image elements in the image 1110. For example, it can move upward to position B. At this time, the electronic device 1100 can play the picture of the main body 3 in the image 1110 moving upward relative to other image elements in the image 1110.

[0273] For example, Figure 12 shows a schematic diagram of another human-computer interaction interface provided in an embodiment of this application.

[0274] As shown in Figure 12(a), the display screen of the electronic device 1200 displays an image 1210, which includes a tree and a subject 4, such as a balloon.

[0275] As shown in Figure 12(b), when the electronic device 1200 recognizes the user's multi-finger pinching action, the electronic device 1200 controls the main body 4 in the image 1210 to become smaller relative to other image elements in the image 1210. At this time, the electronic device 1200 can play a picture of the main body 4 in the image 1210 becoming smaller relative to other image elements in the image 1210.

[0276] For example, FIG13 shows a functional module schematic diagram of a human-computer interaction device 1300 provided in an embodiment of this application. As shown in FIG13, the device 1300 includes:

[0277] The motion recognition module 1310 is used to recognize the user's first action when the first image is displayed.

[0278] The user's first action can be a facial expression and / or a hand gesture.

[0279] In some embodiments, a user's facial movements may include, for example, blowing air, blinking, raising eyebrows, blinking with one eye, pouting, nodding or shaking the head, shaking the head, rolling the eyes, opening the mouth, chewing, widening the eyes, smiling or laughing, etc.

[0280] In some embodiments, the user's gestures may include, for example, finger tapping, multi-finger pinching (e.g., two-finger pinching, three-finger pinching, five-finger pinching, etc.), multi-finger spreading, finger sliding, finger circling, etc., and may also include actions made by the user through the arm, elbow, etc., which are not limited in this application.

[0281] In some embodiments, the action recognition module 1310 is used to recognize the user's first action based on the user's image information.

[0282] In some embodiments, the motion recognition module 1310 is used to recognize the user's first action based on the user's voice information.

[0283] In some embodiments, the motion recognition module 1310 is used to recognize the user's first action based on the acquired touch signal.

[0284] In some embodiments, the action recognition module 1310 is used to recognize the user's first action based on any two or three of the following information:

[0285] (1) User's image information.

[0286] (2) User's voice information.

[0287] (3) Touch signals on the display screen.

[0288] The first image may include, for example, a picture, a video, an animated image, etc. The first image may be a captured image, a downloaded image, an image sent by another device, or an image generated by AI.

[0289] The display module 1320 is used to play the image frame associated with the first action from among a plurality of image frames corresponding to the first image, based on the user's first action.

[0290] In some embodiments, the display module 1320 may be used to loop image frames associated with the first action.

[0291] In some embodiments, the plurality of image frames corresponding to the first image include one or more images in the first image, and / or include images generated based on one or more images in the first image.

[0292] In some embodiments, the first image may be, for example, an image captured by an electronic device; the first image may also be an image downloaded by the electronic device; the first image may also be an image sent by another device and received by the electronic device.

[0293] In one implementation, after acquiring the first image, the electronic device generates multiple image frames corresponding to the first image. Thus, when the electronic device recognizes the user's first action, it can directly retrieve the image frame associated with the user's first action from the multiple image frames corresponding to the first image and play it.

[0294] In some embodiments, when the first image is a video, the display module 1320 is specifically configured to play a video segment associated with the first action in the video according to the user's first action.

[0295] In one implementation, the display module 1320 is specifically configured to start playback from a first image frame, which is the starting frame of a video segment associated with the first action in the video, or the previous frame of the starting frame, or several frames before the starting frame; and to stop playback when playing to a second image frame, which is the ending frame of a video segment associated with the first action in the video, or the next frame of the ending frame, or several frames after the ending frame.

[0296] In some embodiments, when the first image is a dynamic picture, the plurality of image frames corresponding to the first image include a plurality of pictures in the dynamic picture, and / or include at least one picture generated based on at least one of the dynamic pictures and associated with the first action.

[0297] In some embodiments, when the first image is a static image, the plurality of image frames corresponding to the first image include the static image and at least one image generated based on the static image and associated with the first action.

[0298] In some embodiments, the action of the subject in the image frame associated with the first action is the same as the first action.

[0299] In some embodiments, the action of the subject in the image frame associated with the first action satisfies a preset correspondence with the first action.

[0300] In one example, the first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes a person's hair being blown by the wind.

[0301] In one example, the first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes the animation of the subject being blown.

[0302] In one example, the first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes a picture of the subject in motion.

[0303] In one example, the first action is the action of opening and closing the eyes, and the image frame associated with the first action is an image frame that includes the image of a person opening and closing their eyes.

[0304] In one example, the first action is the action of frowning, and the image frame associated with the first action is an image frame that includes the subject frowning.

[0305] In one example, the first action is the action of laughing, and the image frame associated with the first action is an image frame that includes the subject laughing.

[0306] In one example, the first action is a nodding or shaking motion, and the image frame associated with the first action is an image frame that includes a scene of the subject nodding or shaking their head.

[0307] In one example, the first action is the action of jumping, raising an arm, or looking upwards, and the image frame associated with the first action is an image frame that includes a picture of the subject moving upwards or flying upwards.

[0308] In this context, an image frame that includes a subject moving upwards refers to an image frame in which the subject moves upwards relative to all other image elements in the frame. In other words, the user's visual perception is that only the subject moves upwards, while other image elements remain stationary. Similarly, an image frame that includes a subject flying upwards refers to an image frame in which the subject flies upwards relative to all other image elements in the frame. In other words, the user's visual perception is that only the subject flies upwards, while other image elements remain stationary.

[0309] In one example, the first action is the action of widening the eyes, and the image frame associated with the first action is an image frame that includes a picture of the subject being frightened.

[0310] In one example, the first action is a multi-finger pinching motion, and the image frame associated with the first action is an image frame that includes a scene where the subject changes from large to small.

[0311] Among them, the image frame in which the subject changes size refers to the image frame in which the subject changes size relative to other image elements in the image frame. In other words, the visual perception given to the user is that only the subject changes size in the image frame, while the size of other image elements does not change.

[0312] In this embodiment, when an image is displayed on the electronic device, the electronic device can acquire the user's facial expressions and / or gestures in real time. Based on the recognized facial expressions and / or gestures, it plays image frames associated with the user's facial expressions and / or gestures from among the multiple frames of the image displayed on the screen. For example, it can automatically play still images, moving images, or videos, retrieve corresponding frames from moving images or videos for playback, and perform zooming, scaling, or offset operations on the played images. This allows the user to trigger interaction with the images displayed on the electronic device through air gestures or facial expressions, enriching the interaction methods between the user and the electronic device, providing alternative interaction methods when it is inconvenient for the user to touch the screen, and enhancing the intelligence and fun of the interaction between the user and the image.

[0313] Furthermore, it enables interaction between the user and the image displayed on the screen of the electronic device, as well as interaction between the user and the content of the image displayed on the screen of the electronic device, thus enhancing the emotional interaction between the user and the image.

[0314] For example, FIG14 shows a functional module schematic diagram of another human-computer interaction device 1400 provided in an embodiment of the present application.

[0315] As shown in Figure 14, the device 1400 includes an image acquisition module 1410, an action recognition module 1420, and a display module 1430, specifically:

[0316] The image acquisition module 1410 is used to acquire the user's image information when the electronic device displays the first image.

[0317] In some embodiments, the image acquisition module 1410 may be turned on, for example, when the electronic device displays the first image, and begin to acquire the user's image information in real time.

[0318] In some embodiments, the image acquisition module 1410 may include, for example, a low-power image sensor located on a front-mounted electronic device.

[0319] The first image can be, for example, a video, a moving image, or a still image.

[0320] In some embodiments, the image acquisition module 1410 can acquire the user's image information in real time at a certain frame rate and resolution.

[0321] The motion recognition module 1420 is used to recognize the user's first action based on the user's image information.

[0322] In some embodiments, the action recognition module 1420 may include, for example, a neural network computing unit.

[0323] In some embodiments, the user's first action may be a facial expression and / or a hand gesture.

[0324] In some embodiments, the motion recognition module 1420 acquires the user's facial expression image information and / or gesture image information in real time based on the user's image information, and can further recognize the user's facial movements based on the user's facial expression image information, and / or recognize the user's hand gestures based on the user's gesture image information.

[0325] In some embodiments, a user's facial movements may include, for example, blowing air, blinking, raising eyebrows, blinking with one eye, pouting, nodding or shaking the head, shaking the head, rolling the eyes, opening the mouth, chewing, widening the eyes, smiling or laughing, etc.

[0326] In some embodiments, the user's gestures may include, for example, finger tapping, multi-finger pinching (e.g., two-finger pinching, three-finger pinching, five-finger pinching, etc.), multi-finger spreading, finger sliding, finger circling, etc., and may also include actions made by the user through the arm, elbow, etc., which are not limited in this application.

[0327] The display module 1430 is used to play the image associated with the first action from among multiple image frames corresponding to the first image, based on the user's first action.

[0328] In some embodiments, the electronic device further includes a determining module for determining an image associated with the first action from a plurality of image frames corresponding to the first image.

[0329] In one example, the actions of the subject in the image frame associated with the first action satisfy a preset correspondence. Based on this preset correspondence, the determining module can determine the image associated with the first action from multiple image frames corresponding to the first image.

[0330] In some embodiments, when the method provided in this application is applied to the display control of dynamic wallpapers, for example, the dynamic switching of wallpapers can be controlled by the method provided in this application.

[0331] For example, users can control multiple wallpapers under a wallpaper theme to be displayed sequentially according to the order in which the user makes different facial expressions.

[0332] One or more modules or units described herein can be implemented in software, hardware, or a combination of both. When any of the above modules or units are implemented in software, the software exists as computer program instructions and is stored in memory. A processor can be used to execute the program instructions and implement the above method flow. The processor can include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., and various computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor can be built into a SoC (System-on-a-Chip) or an application-specific integrated circuit (ASIC), or it can be a separate semiconductor chip. In addition to the cores for executing software instructions to perform calculations or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), PLDs (programmable logic devices), or logic circuits that implement dedicated logic operations.

[0333] When the modules or units described herein are implemented in hardware, the hardware may be any one or any combination of a CPU, microprocessor, DSP, MCU, artificial intelligence processor, ASIC, SoC, FPGA, PLD, application-specific digital circuit, hardware accelerator, or non-integrated discrete device, which may run the necessary software or perform the above method flow independently of software.

[0334] When the modules or units described herein are implemented using software, they can be implemented, in whole or in part, in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0335] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0336] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0337] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0338] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0339] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0340] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0341] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for human-computer interaction, characterized in that, The method is applied to a first device, and the method includes: When the first device displays the first image, the first device recognizes the user's first action; The first device plays the image frame associated with the first action from among the multiple image frames corresponding to the first image, based on the first action.

2. The method according to claim 1, characterized in that, The plurality of image frames corresponding to the first image include one or more images in the first image, and / or include images generated based on one or more images in the first image.

3. The method according to claim 1 or 2, characterized in that, When the first image is a video, the first device plays the image frame associated with the first action from among multiple image frames corresponding to the first image, including: The first device plays a video segment in the video associated with the first action, based on the first action.

4. The method according to claim 3, characterized in that, Playing the video segment associated with the first action in the video includes: The first device starts playing from the first image frame, which is the starting frame of the video segment associated with the first action in the video or the frame preceding the starting frame; The first device stops playing when it reaches the second image frame, which is the end frame of the video segment associated with the first action or the frame following the end frame.

5. The method according to claim 1 or 2, characterized in that, When the first image is a dynamic image, the multiple image frames corresponding to the first image include multiple images in the first image, and / or include at least one image generated based on at least one image in the first image and associated with the first action.

6. The method according to claim 1 or 2, characterized in that, When the first image is a static image, the multiple image frames corresponding to the first image include the static image and at least one image generated based on the static image that is associated with the first action.

7. The method according to any one of claims 1 to 6, characterized in that, The image frames associated with the first action among the multiple image frames corresponding to the first image played include: The first device continuously plays the image frames associated with the first action.

8. The method according to any one of claims 1 to 7, characterized in that, When the first device displays the first image, the first device recognizes the user's first action, including: When the first device displays the first image, the first device acquires the user's image information; The first device identifies the user's first action based on the user's image information.

9. The method according to any one of claims 1 to 7, characterized in that, When the first device displays the first image, the first device recognizes the user's first action, including: While the first device displays the first image, the first device acquires the user's voice information; The first device identifies the user's first action based on the user's voice information.

10. The method according to claim 9, characterized in that, The first device acquires the user's voice information, including: The first device acquires the user's voice information through a microphone.

11. The method according to any one of claims 1 to 7, characterized in that, When the first device displays the first image, the first device recognizes the user's first action, including: When the first device displays the first image, the first device acquires the touch signal on the display screen; The first device recognizes the user's first action based on the touch signal.

12. The method according to any one of claims 1 to 7, characterized in that, When the first device displays the first image, the first device recognizes the user's first action, including: When the first device displays the first image, the first device acquires any two or three of the user's image information, the user's voice information, and the touch signal on the display screen. The first device identifies the user's first action based on any two or three of the acquired information.

13. The method according to any one of claims 1 to 12, characterized in that, The action of the subject in the image frame associated with the first action is the same as the first action, or The action of the subject in the image frame associated with the first action satisfies a preset correspondence with the first action.

14. The method according to any one of claims 1 to 13, characterized in that, The first action is the action of blowing air, and the image frame associated with the first action is an image frame including an image of the subject's hair being blown by the wind; or The first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes a scene of the subject being blown; or The first action is the action of blowing air, and the image frame associated with the first action is an image frame that includes a subject in motion; or The first action is the action of opening and closing the eyes, and the image frame associated with the first action is an image frame that includes the subject's open and closed eyes; or The first action is the action of frowning, and the image frame associated with the first action is an image frame that includes the subject frowning; or The first action is a smiling or laughing action, and the image frame associated with the first action is an image frame that includes a subject smiling or laughing; or The first action is a nodding or shaking motion, and the image frame associated with the first action is an image frame that includes a nodding or shaking motion of the subject; or The first action is jumping, raising an arm, or looking upwards, and the image frame associated with the first action is an image frame that includes a view of the subject moving upwards; or The first action is a multi-finger pinching action, and the image frame associated with the first action is an image frame that includes a subject changing from large to small.

15. The method according to any one of claims 1 to 14, characterized in that, Before the first device displays the first image, the method further includes: The first device acquires the first image; The first device generates a plurality of generated image frames corresponding to the first image based on the first image, wherein the plurality of generated image frames corresponding to the first image are at least a portion of the plurality of image frames corresponding to the first image.

16. The method according to claim 15, characterized in that: At least one first image frame among the plurality of generated image frames corresponding to the first image is associated with the first action; or at least one first image frame among the plurality of generated image frames corresponding to the first image is associated with the first action, and at least one second image frame among the plurality of generated image frames corresponding to the first image is associated with the second action.

17. The method according to any one of claims 1 to 16, characterized in that, The user's first action is used to control the content in the first image.

18. An electronic device, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 17.

19. A computer-readable storage medium, characterized in that, The storage medium stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 17.

20. A chip, characterized in that, The chip stores instructions that, when executed, implement the method as described in any one of claims 1 to 17.

21. A computer program product, characterized in that, The computer program product stores a program or instructions that, when executed, implement the method as described in any one of claims 1 to 17.