Image understanding method and electronic device

CN122795201APending Publication Date: 2026-09-22HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349248.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,对于视障人士(包括全盲和低视力用户)而言,如何高效、准确地感知图片依然面临着巨大的挑战

Benefits of technology

[0056]上述第三方面至第六方面中的各个方面以及各个方面可能达到的技术效果请参照上述针对第一方面或第一方面中的各种可能方案,或者上述第二方面或第二方面中的各种可能方案可以达到的技术效果说明,这里不再重复赘述。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122795201A_ABST
    Figure CN122795201A_ABST
Patent Text Reader

Abstract

An image understanding method and an electronic device are provided to improve user experience. The method can include: obtaining a first image; and outputting a vibration signal based on a first touch position of a user in the first image, the vibration signal being used to indicate image information at the first touch position in the first image or to indicate a distance between the first touch position and at least one target subject in the first image. Through the above method, a visually impaired user can understand an image from a tactile perspective through a vibration signal, such as guiding the user to touch a target subject in the image through a tactile sense or enabling the user to perceive information in the image through a tactile sense, thereby improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic device technology, and in particular to an image understanding method and an electronic device. Background Technology

[0002] With the rapid development of information technology, smart mobile terminals have become one of the main tools for people to obtain and process information. However, for visually impaired people (including those who are totally blind and have low vision), efficiently and accurately perceiving images still faces enormous challenges. Summary of the Invention

[0003] This application provides an image understanding method and an electronic device to improve user experience.

[0004] Firstly, embodiments of this application provide an image understanding method that can be applied to an electronic device. Here, "electronic device" can refer to the electronic device itself, or to a processor, module, chip, or chip system within the electronic device that implements the method. The method may include: acquiring a first image; and outputting a vibration signal based on a user's first touch position in the first image, wherein the vibration signal is used to indicate image information at the first touch position in the first image or to indicate the distance between the first touch position and at least one target object in the first image.

[0005] The above methods can enable visually impaired users to understand images through touch by using vibration signals. For example, users can be guided to touch the target subject in the image through touch, or users can perceive information in the image through touch, thereby improving the user experience.

[0006] In one possible design, the first touch position is located outside the at least one target subject, and the vibration signal has a first vibration frequency, which is used to indicate the distance between the first touch position in the first image and the at least one target subject. This allows the user to identify the distance to the target subject by sensing the vibration frequency of the first vibration signal, thereby enabling the user to touch a specific target subject.

[0007] In one possible design, the distance between the first touch position and any one of the at least one target subject can include: the distance between the first touch position and a first edge point, where the first edge point is the edge point on any one of the target subjects that is closest to the first touch position. This allows for a clear determination of the distance between the user's touch position and any given target subject.

[0008] In one possible design, when the at least one target subject is a single target subject, the magnitude of the first vibration frequency is negatively correlated with the distance between the first touch position and the single target subject; or, when the at least one target subject is multiple target subjects, the magnitude of the first vibration frequency is correlated with the distance between the first touch position and any one of the multiple target subjects, as well as the respective weights of the multiple target subjects. When there is one target subject in the first image, the user can identify the distance to the target subject by the change in the first vibration frequency. For example, the closer the user is to the target subject, the higher the perceived first vibration frequency; the farther the user is from the target subject, the lower the perceived first vibration frequency, thereby guiding the user to touch the target subject. When there are multiple target subjects in the first image, the user can perceive the distance relationship between the first touch position and multiple target subjects by the first vibration frequency, thereby guiding the user to touch a specific target subject.

[0009] In one possible design, when the first image contains multiple target subjects, the magnitude of the first vibration frequency changes differently when the distance between the first touch position and different target subjects changes the same. This allows the user to identify the presence of multiple target subjects by sensing the change in the first vibration frequency, thereby guiding the user to touch a specific target subject.

[0010] In one possible design, the first vibration frequency can satisfy the following formula:

[0011]

[0012] Where f(x, y) is the first vibration frequency at the first touch position (x, y), f max The preset maximum first vibration frequency, n is the number of target subjects in the first image, and a i For the weight coefficients associated with each target subject i, d i (x, y) represents the closest distance between the first touch position and the i-th target, d max This is the preset maximum distance.

[0013] Using the above formula, the vibration frequency of the vibration signal felt by the user at the first touch position can be accurately output by combining information such as the distance between each target subject and the first touch position and the weight of each target subject, thus realizing the understanding of the image from the perspective of touch.

[0014] In one possible design, the first touch position is located within a first target subject, which is any one of the at least one target subject; the vibration signal has at least one of the following: vibration intensity, a second vibration frequency, or a vibration rhythm; the vibration intensity is used to characterize the area and depth of the first target subject, the second vibration frequency is used to characterize the detail complexity of the area where the first touch position is located, and the vibration rhythm is used to characterize the emotional information of the image corresponding to the first image. Thus, when a user touches the first target subject, the user can understand the image information at the first touch position through touch.

[0015] In one possible design, the vibration intensity can satisfy the following formula:

[0016] A(S, Z) = (A min +(A max -A min )×S / S max )×(1-Z / Z max )

[0017] Where A(S, Z) is the vibration intensity, S is the area of ​​the first target body, Z is the depth of field of the first target body in the first image, and A max For the preset maximum vibration intensity, A min S is the preset minimum vibration intensity. max Z is the preset maximum area. max This is the preset maximum depth of field.

[0018] The vibration intensity output by the above formula allows the user to perceive the area of ​​the first target subject and its depth in the image.

[0019] In one possible design, the second vibration frequency can satisfy the following formula:

[0020] F(D)=F min +(F max -F min )×(D / D max )

[0021] Where F(D) is the second vibration frequency, F max F is the preset maximum second vibration frequency. min Let D be the preset minimum second vibration frequency, and D be the detail complexity of the area where the first touch position is located. max This represents the preset maximum detail complexity.

[0022] The second vibration frequency output by the above formula allows the user to perceive the detail and complexity of the area where the first touch position is located.

[0023] In one possible design, the vibration rhythm is positively correlated with the intensity of the emotion corresponding to the emotional information in the image. This allows the user to perceive which emotion the first image corresponds to through the vibration rhythm.

[0024] In one possible design, the emotional information of the image includes one or more of the following: intense emotion or calm emotion.

[0025] In one possible design, the vibration rhythm is periodic.

[0026] In one possible design, a first message is output, indicating that the first target object has been touched; a second message is received, indicating a switch to the first mode; the vibration signal in the first mode has at least one of the following: vibration intensity, the second vibration frequency, or the vibration rhythm. Based on this method, after a user touches a target object, the user can further perceive image information at the touch location through touch based on the user's trigger.

[0027] Secondly, embodiments of this application provide another image understanding method, which can be applied to electronic devices. Here, "electronic device" can refer to the electronic device itself, or to a processor, module, chip, or chip system within the electronic device that implements the method. The method may include: detecting a first operation on a first image; and, in response to the first operation, outputting first image description information via voice, wherein the first image description information is determined based on user-related information and the first image.

[0028] By using the methods described above, we can provide users with detailed image descriptions by combining relevant user information, thereby improving the user experience.

[0029] In one possible design, the first image description information is generated either after the first operation is detected or before the first operation is detected. This allows for greater flexibility in the generation of the first image description information.

[0030] In one possible design, the first operation on the first image includes: tapping the first image or swiping the first image. This allows the electronic device to output first image description information in a relatively simple way.

[0031] In one possible design, the first image description information is determined based on user-related information and the first image. This can include: the first image description information is determined based on contextual information of the content in the first image and the first image itself, whereby the contextual information of the content in the first image is determined based on the user-related information. This allows for the combination of user-related information to obtain the contextual information of the content in the first image, thereby providing the user with a more personalized image description.

[0032] In one possible design, the contextual information of the content in the first image is determined based on the user-related information, which may include: determining the contextual information of the content in the first image based on a feature database, wherein the feature database includes the user-related information. This allows for accurate determination of the contextual information of the content in the first image.

[0033] In one possible design, the contextual information of the content in the first image includes information about the first subject. This allows for relevant descriptions of the first subject, enabling the user to understand the first image.

[0034] In one possible design, the first subject in the first image is identified; information about the first subject is determined based on the user-related information. This allows for accurate determination of the information about the first subject.

[0035] In one possible design, the first subject in the first image can be determined by performing facial recognition on the first image to identify the first subject. This way, when the first subject is a person, the first subject in the first image can be accurately identified.

[0036] In one possible design, the first image description information is determined based on the contextual information of the content in the first image and the first image itself. This may include: the first image description information being determined based on information about the first subject, the shooting time and location corresponding to the first image, and the content contained in the first image. This allows for the combination of information about the first subject and the data information of the first image itself to obtain more detailed first image description information, enabling the user to fully understand the first image.

[0037] In one possible design, the information of the first subject may include one or more of the following: the identity of the first subject, the birth date information of the first subject, the address information of the first subject, or the tag information of the first subject. This allows the first image description information determined based on the information of the first subject to be more detailed, enabling the user to fully understand the first image.

[0038] In one possible design, the user-related information includes one or more of the following: personal identification information, personal date of birth information, personal address information, personal habit information, interpersonal relationship information, or user-related tag information. This allows for a more detailed description of the first image based on the user-related information, enabling the user to fully understand the first image.

[0039] In one possible design, a second image description is output via voice. This second image description is determined based on descriptions corresponding to multiple images associated with the first image and the description information of the first image. This allows for contextualized descriptions to be provided to the user based on the first image and other associated images, such as sequential contextualized descriptions or comparative descriptions across timelines, thus enhancing the user experience.

[0040] In one possible design, multiple images associated with the first image are determined based on the first image description information. This allows for accurate identification of the multiple images associated with the first image, resulting in more accurate second image description information.

[0041] In one possible design, multiple images associated with the first image are determined based on the first image description information. This can be achieved by determining these multiple images based on the first image description information and a feature database, where the feature database includes user-related information. This allows for accurate determination of the multiple images associated with the first image, resulting in more accurate second image description information.

[0042] In one possible design, the plurality of images have at least one of the following relationships with the first image:

[0043] The multiple images and the first image are images from the same activity scene;

[0044] The plurality of images and the first image are images taken of the same subject;

[0045] The multiple images and the first image are images taken at the same location at different times;

[0046] The multiple images and the first image were taken within the same time period.

[0047] The above method allows for a relatively flexible determination of multiple images associated with the first image.

[0048] In one possible design, the second image description information is output via voice. This can be achieved by: responding to the receipt of third information by outputting the second image description information via voice. Based on this method, a contextualized description of the first image and multiple images associated with it can be further provided to the user based on user triggers.

[0049] In one possible design, before receiving the third information, a fourth information can be output. This fourth information requests the user to select whether to switch to the second mode, and the third information indicates whether to switch to the second mode. Based on this method, after providing the user with descriptive information corresponding to the first image, a contextualized description of the first image and multiple images associated with it can be further provided to the user based on user triggers.

[0050] In one possible design, the feature database is generated based on one or more of the following: user-entered information, data from the address book, or data from the photo album. This ensures accurate generation of the feature database.

[0051] Thirdly, embodiments of this application provide an electronic device including one or more processors and a memory; the one or more processors are coupled to the memory, and the one or more processors are configured to read a computer program stored in the memory to execute the method provided in the first or second aspect. The computer program code includes computer instructions.

[0052] Fourthly, embodiments of this application provide a chip, the chip including a processor coupled to a memory, the processor being used to call computer program instructions stored in the memory during runtime to implement the methods provided in the first or second aspect above.

[0053] Optionally, the chip may also include components such as a memory, a communication interface, and a power supply module. The memory is used to store computer programs; the communication interface is used to receive and send data; and the power supply unit is used to supply power to the processor.

[0054] Fifthly, embodiments of this application provide a computer storage medium storing computer program instructions that, when executed on an electronic device, cause the computer to perform the methods provided in the first or second aspect described above.

[0055] In a sixth aspect, embodiments of this application provide a computer program product, the computer program product including computer program instructions; when the computer program instructions are executed on a computer, the computer causes the computer to perform the method provided in the first aspect or the second aspect.

[0056] For the various aspects of the third to sixth aspects mentioned above, and the technical effects that each aspect may achieve, please refer to the above description of the technical effects that can be achieved for the first aspect or the various possible solutions in the first aspect, or the second aspect or the various possible solutions in the second aspect, which will not be repeated here. Attached Figure Description

[0057] Figure 1 A schematic diagram of an electronic device provided in this application;

[0058] Figure 2 A schematic diagram of the software structure of an electronic device provided in this application;

[0059] Figure 3 A schematic diagram of the implementation modules involved in the image understanding method provided in this application;

[0060] Figure 4 A schematic diagram of the processing flow of a perception navigation module provided in this application;

[0061] Figure 5 A schematic diagram of the processing flow of a three-dimensional haptic feedback module provided in this application;

[0062] Figure 6 A schematic diagram illustrating the processing flow of a single-image feature description module provided in this application;

[0063] Figure 7 A schematic diagram illustrating the processing flow of a scene-based image presentation module provided in this application;

[0064] Figure 8 A schematic diagram illustrating the changes in user touch position provided in this application;

[0065] Figure 9 A flowchart of an image understanding method provided in this application;

[0066] Figure 10 A flowchart of another image understanding method provided in this application;

[0067] Figure 11 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0068] This application provides an image understanding method and an electronic device to improve user experience. The method and device described in this application are based on the same technical concept. Since the principles by which the method and device solve problems are similar, their implementations can be mutually referenced, and repeated details will not be elaborated further.

[0069] In the description of this application, the terms "first," "second," etc., are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance or order.

[0070] In the description of this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0071] In the description of this application, "and / or" describes the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A and B can be singular or plural. " / " means "or", for example, a / b means a or b.

[0072] To more clearly describe the technical solutions of the embodiments of this application, the image understanding method and electronic device provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0073] This application provides an electronic device that can implement the image understanding method provided in this application. The electronic device is a device or apparatus with data connectivity, data calculation, and processing capabilities. The electronic device may have a display screen to show a user interface for human-computer interaction.

[0074] For example, the electronic device in this application can be a tablet computer, personal computer (PC), laptop computer, computer, netbook, in-vehicle computer, in-vehicle terminal, mobile phone (such as smartphone), smart wearable device (such as smartwatch, smart bracelet, smart glasses, smart helmet, etc.), personal digital assistant (PDA), smart home device (such as smart TV, smart mirror, smart speaker, etc.). Among these, mobile phones can include foldable phones, non-foldable phones, etc. This application does not limit the specific form of the electronic device.

[0075] Electronic devices can perform their functions and provide services to users through an operating system they run. For example, the electronic device may, but is not limited to, running an operating system... Or other operating systems.

[0076] The following is for reference. Figure 1The structure of the electronic device provided in the embodiments of this application will be described.

[0077] like Figure 1 As shown, the electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a USB interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a SIM card interface 195, etc.

[0078] Processor 110 may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0079] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from this memory. By providing memory, the number of times the processor 110 accesses data in the internal memory 121 can be reduced, thus reducing the processor 110's waiting time and improving system efficiency.

[0080] Display screen 194 is used to display various user interfaces. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays screens 194, where N is a positive integer greater than 1. Display screen 194 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces (GUIs). For example, display screen 194 may display windows, images, videos, web pages, or documents, etc.

[0081] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, program code for at least one application, etc. The operating system may include, but is not limited to, […]. The storage program area can store computer programs for implementing the image understanding method provided in the embodiments of this application. When the processor 110 executes these computer programs, the electronic device can execute the image understanding method. The storage data area can store data created during the use of the electronic device 100, such as the feature database involved in the embodiments of this application.

[0082] In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0083] The sensor module 180 may include, but is not limited to, fingerprint sensors, touch sensors, pressure sensors, magnetic sensors, ambient light sensors, barometric pressure sensors, bone conduction sensors, etc.

[0084] A touch sensor, also known as a "touch panel," can be located on the display screen 194. The touch sensor detects touch operations applied to or near it. It transmits the detected touch operation to an application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some embodiments, the touch sensor can be integrated with the display screen 194 to form a touch display. In other embodiments, the touch sensor may be located on the surface of the electronic device 100, in a different position than the display screen 194.

[0085] The wireless communication function of electronic device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 150 can provide wireless communication solutions for electronic device 100, including 2G / 3G / 4G / 5G. Wireless communication module 160 can provide wireless communication solutions for electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies.

[0086] It should be understood that Figure 1 The electronic device 100 shown is merely an example and does not constitute a limitation on the electronic device. In practical applications, the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0087] The software system of the electronic device 100 provided in this application embodiment can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment takes a layered architecture as an example. Figure 2 The software structure of an electronic device is illustrated by example.

[0088] A layered architecture divides the software within an electronic device into several layers, each with a clear role and function. Layers can communicate with each other through software interfaces. For example... Figure 2 As shown, the software architecture of an electronic device, from top to bottom, includes at least the following: application layer, system service layer, and kernel layer. Additionally, an electronic device may include a hardware abstraction layer (HAL), which provides hardware resource support for the electronic device.

[0089] The application layer is the top layer of the operating system, including the operating system's native applications and installed third-party applications, such as camera, gallery, calendar, Bluetooth, music, video, messaging, calculator, chat applications, video applications, music applications, electronic map applications, social applications, image editing applications, etc.

[0090] The system service layer is an important part of the operating system of an electronic device, including various system services provided by the operating system. For example, the system service layer may include: window management service, operation management service, notification management service, layer management service, input management service, etc.

[0091] The window management service manages the display and configuration of the window library. Based on display settings and multiple management functions, it controls the display of windows or interfaces on the screen, adjusting their transparency, position, and size. It also includes features such as screen locking and screenshot capabilities. The window management service can obtain parameters such as screen size and resolution, and can identify display areas in the user interface, such as the status bar.

[0092] The operation and management service is responsible for starting, switching, and scheduling various components in the system, as well as managing and scheduling applications.

[0093] Notification management services allow applications to display notification information in the status bar, as well as to convey informational messages. These messages can disappear automatically after a short pause, without requiring user interaction.

[0094] The input management service is used to collect, process, and distribute user operation information and input events. For example, the input management service can receive operation information reported by touch drivers or other input device drivers (such as keyboard drivers, mouse drivers, etc.), convert it into input events, and distribute the input events to other modules for processing.

[0095] Layer management service is a standalone service provided by the operating system of an electronic device. It is a low-level, resident service of the operating system. During device operation, this layer management service needs to run continuously, and the operating system prioritizes allocating resources to it to ensure that it can perform image compositing at any time, thereby ensuring that the electronic device can display the user interface in real time. The layer management service is mainly responsible for the creation, control, and management of layers. It composites and renders the image frames of the corresponding window in each layer, and finally combines the graphic frames from all layers into the image of the user interface to be displayed on the screen.

[0096] It should be noted that the system service layer can also have other system services, which will not be elaborated here.

[0097] Optionally, between the application layer and the system service layer, the software architecture of the electronic device may also include an application framework layer (FWK) (not shown in the figure). The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer may include some predefined functions. For example, the application framework layer may include a content provider, a view system, a phone manager, a resource manager, etc. The content provider stores and retrieves data, making this data accessible to applications in the application layer. The view system includes visual components, such as components that display text, images, documents, etc.; components in the view system can be used to build the user interface of the application. The phone manager provides the communication functions of the electronic device.

[0098] The kernel layer provides the core system services of the operating system, such as security, memory management, process management, network protocol stack, and driver models, all of which are implemented based on the kernel layer. The kernel layer also serves as an abstraction layer between the hardware and software stacks. This layer contains many drivers related to electronic devices, including: display drivers; keyboard drivers; mouse drivers; Flash drivers for memory-based devices; camera drivers; audio drivers; Bluetooth drivers; WiFi drivers; and touch drivers.

[0099] HAL (Hardware Interface Layer) is the interface layer between the operating system kernel and the hardware circuitry of an electronic device, and its purpose is to abstract the hardware. For example, HAL can include various sensors (such as touch sensors), displays, cameras, hardware composers (HWCs), and other hardware in an electronic device. The HWC is primarily responsible for image compositing and display.

[0100] It should be noted that, Figure 2This is merely an example of the software architecture of an electronic device, simply listing some layers and software modules, and does not constitute any limitation on the software architecture of the electronic device. In practical applications, the operating system of an electronic device may include other layers, or each layer may include other software modules for implementing one or more functions or services. Furthermore, this application does not limit the specific layer where each software module resides. For example, some system services in the system service layer may be deployed in the application framework layer or in the kernel layer.

[0101] Furthermore, this application does not specifically limit the names of the various software modules in the software structure. For example, in In operating systems, the window management service can be WindowManagerServer (WMS), the layer management service can be the layer management engine (SurfaceFlinger), and the input management service can be InputManagerService. InputManagerService can include modules such as input management (InputManager), event hub (EventHub), and input dispatcher (InputDispatcher). operating system In this context, the window management service can be SceneBoard (SCB), the layer management service can be RenderService (RS), and the input management service can be Multimodal Input Service (INPUT).

[0102] Current accessibility technologies can only provide visually impaired users with basic information about images through voice, such as simple descriptions like "This is a picture of a cat," resulting in a poor user experience for visually impaired users. Therefore, this application provides an image understanding method that can present images to users through touch, provide users with a deeper understanding of images through voice, or provide a deeper understanding of images through a combination of touch and voice.

[0103] In this application, "image" can also be replaced with descriptions such as "picture," "photo," or "photograph."

[0104] To facilitate understanding of the image understanding method provided in the embodiments of this application, the following will be explained... Figures 3-10 This application describes the implementation process of the image understanding method provided. For ease of understanding, the implementation modules that may be involved in the image understanding method provided in the embodiments of this application are first introduced below. These modules can exist independently or in combination; this application does not limit their existence. In the following description, an electronic device is used as an example. In practice, the electronic device can also be replaced by a processor, chip, or a module within an electronic device.

[0105] For example, such as Figure 3 As shown, the image understanding method provided in this application may include at least one of the following implementation modules: an accessibility feature library construction module, a perception and navigation module, a stereoscopic haptic feedback module, a single image feature description module, or a scene-based image presentation module. It should be noted that the names of each module are merely examples and may have other names. This application uses this as an example only and is not intended to limit the scope of this application. The following is a detailed description of each module:

[0106] (1) Accessibility Feature Library Construction Module

[0107] This module can acquire user-related information through proactive information input and / or intelligent data mining to build an accessibility feature database (which can be referred to as a feature database, person / object feature database, etc.) for visually impaired users. By integrating multiple data sources and using data filtering and clustering techniques, this module can collect key information about users' personal preferences, social relationships, and daily lives. The accessibility feature database built by this module helps electronic devices generate more accurate and context-sensitive content, providing users with a highly customized accessibility experience.

[0108] One method of active information entry allows users to proactively input relevant information. For example, the electronic device allows users or their agents to actively enter key information, such as people closely related to the user, family relationships, important events (such as birthdays and anniversaries), residential address, and workplace. For instance, users or their agents can manually add the birthdays of relatives and friends and the user's relationship with those relatives or friends. Optionally, users or their agents can actively enter relevant information through preset editing options in the user's address book on the electronic device, or through other preset editing entry points.

[0109] The aforementioned proactive information input method can provide electronic devices with core, personalized background data, ensuring that the device can generate content that aligns with the user's actual life situation when performing image descriptions. For example, when a user proactively adds the birthdays of relatives and friends and their relationship with those relatives or friends, the electronic device can generate a more contextualized description that better reflects the user's actual needs when it recognizes the relevant individuals.

[0110] Intelligent data mining methods may include automatically obtaining contact information from address book data and / or extracting information from smart portrait photo album data sources. Of course, intelligent data mining methods may also include other methods, which are not limited in this application.

[0111] This module automatically retrieves contact information from the user's address book, extracting contact information (such as name, birthday, relationship, address, etc.) and matching it with people in the photo album. Using this data, the electronic device can further infer the identity of the person in the image and their relationship to the user, thus providing a more personalized description. For example, when the electronic device detects that the person in the image is a contact in the address book, it can combine birthday information to generate a contextualized description such as "This is a photo from your friend [name]'s] birthday party."

[0112] For extracting information from the smart portrait album data source, this module can utilize the built-in intelligent portrait classification function of the album to extract and organize the people data in the smart album, clustering and labeling the people in the photos. Specifically, this module can identify the same person in different photos through facial recognition and clustering algorithms, and associate this data with contact information in the user's address book to further enrich the feature database. For example, this module can automatically identify which photos contain the user's friends, family, or colleagues, and can also identify the shooting location information, clustering and storing these photos to form accessibility feature tags based on people.

[0113] In some implementations, when the module obtains data through at least two of the above methods, the module can integrate the data obtained through different methods to obtain user-related information.

[0114] Furthermore, this module can automatically adjust the data clustering rules and the generation method of contextual descriptions based on user feedback and the input of new data, enabling the feature database to have dynamic update and self-learning capabilities. In other words, this module can continuously optimize the feature database content as users use it daily and data accumulates. For example, as users take more images, the module can more accurately identify and classify different scenes, people, and events, ensuring that the feature database remains updated and adapts to changing user needs.

[0115] (2) Perception and Navigation Module

[0116] The sensory navigation module uses guided tactile feedback to help visually impaired users move their fingers towards the main objects in an image, allowing them to intuitively perceive the location of the main objects and thus gain a coarse-grained understanding of the image's composition. The module works by altering the frequency of vibrations the user perceives based on the distance of their touch relative to the objects in the image, thereby guiding their touch towards the main object area.

[0117] Optionally, the subject in the image described herein can be the target subject in the image. In addition to containing at least one target subject, the image may or may not contain other subjects that are not the target subject. Subjects that are not the target subject can be ignored and do not need to be perceived by the user.

[0118] Optionally, the sensing navigation module can output a vibration signal at a user's touch location, and this vibration signal has a vibration frequency. For example, the vibration frequency of the vibration signal at a user's touch location output by the sensing navigation module is denoted as the first vibration frequency.

[0119] The first vibration frequency is based on the change in distance between the touch position and the subject, and the first vibration frequency can satisfy the following formula:

[0120]

[0121] Where f(x, y) is the first vibration frequency at the touch position (x, y), and f(x, y) changes as the user touches the position. max The preset maximum first vibration frequency (typically used to represent the vibration frequency when the user touches a subject very close to it) is a constant, where n is the number of subjects in the image touched by the user, and a i For the weight coefficients associated with each subject i, d i (x, y) represents the nearest distance between the touch position and the i-th subject, d max The preset maximum distance can represent the maximum distance that a user can touch (e.g., typically the length of the image diagonal). This value is used to normalize the distance so that the variation range of the first vibration frequency is within a controllable range.

[0122] Based on the importance of the subject (such as the priority of people or objects), different weight coefficients a can be assigned to different subjects. i If all subjects are equally important, the weighting coefficients can be set to the same value.

[0123] d i (x, y) can be calculated based on Euclidean distance. For example, Among them, (x ij y ij () represents the j-th pixel of the i-th subject.

[0124] The first vibration frequency f(x, y) at the touch position can be understood as follows: the first vibration frequency f(x, y) is based on the touch position and the corresponding d of each subject. i The frequency is adjusted using (x, y). The closer the user's touch position is to a certain object, the higher the initial vibration frequency. When the touch position is close to a certain object, the distance d... iAs (x, y) decreases, its total contribution to the first vibration frequency decreases, leading to an increase in the first vibration frequency. Conversely, the farther the touch position is from the subject, the greater the distance d. i As (x, y) increases, the first vibration frequency decreases.

[0125] In some implementations, the distance between the user's touch position and a certain subject can be the distance between the touch position and a first edge point, where the first edge point can be the edge point in the subject that is closest to the touch position.

[0126] For different subjects, when the distance between the user's touch position and the subject changes the same, the change in the first vibration frequency can be different. For example, the frequency at which the first vibration frequency increases as the distance between the touch position and the first subject becomes closer is different from the frequency at which the first vibration frequency increases as the distance between the touch position and the second subject becomes closer.

[0127] Optionally, the change in the first vibration frequency can be differentiated based on the importance of different subjects when the distance between the touch position and different subjects changes.

[0128] Optionally, for different subjects, when the user's touch position is at the same distance from the subject, the corresponding first vibration frequency is different.

[0129] For different subjects, the user's touch position is located at the same first vibration frequency in different subjects, which is the preset maximum first vibration frequency.

[0130] For the same subject, the first vibration frequency is the same at any point on the subject where the touch position is located.

[0131] For example, the processing flow of this perception navigation module can be as follows: Figure 4 As shown, when a user touches an image on the touchscreen of an electronic device, this module first performs a subject segmentation mask on the image. Then, based on the user's touch position on the image, it performs a non-zero pixel distance mapping. This allows for the generation of guided vibration intensity sequences based on different touch positions, which are then fed back to the user via vibration signals with different first vibration frequencies. In other words, for visually impaired users, when exploring an image on the touchscreen, the first vibration frequency can indicate the distance between the touch position and the subject in the image. When the touch position is closer to the subject, the vibration becomes more frequent, helping the user accurately locate important objects in the image. This vibration feedback method can help users perceive, navigate, and understand multiple subjects and their relative positions in an image even without visual perception.

[0132] In some implementations, to reduce the performance bottleneck of electronic device vibration, after generating a pixel map based on a first vibration frequency, the pixel values ​​can be approximated by isotropic processing to reduce the enumeration number of the first vibration frequency, ultimately generating a guided intensity map. Thus, for the same image, when the user repeatedly touches the image subsequently, the sensing and navigation module can directly output the vibration signal corresponding to the touch position based on the generated guided intensity map. That is, the user can feel the corresponding first vibration frequency, eliminating the need to determine the frequency for the same image and saving power.

[0133] (3) Three-dimensional haptic feedback module

[0134] A three-dimensional haptic feedback module can provide visually impaired users with multi-dimensional haptic feedback on images. When a visually impaired user browses an image, different haptic feedbacks are provided based on the user's touch location, allowing the user to perceive key content in the image through touch. The module works by combining the area of ​​the main subject, its details, and the emotions of the figures in the image to generate a relatively complex vibration feedback pattern. This vibration feedback pattern can reflect at least one of vibration intensity, vibration frequency (referred to as the second vibration frequency in this application), or vibration rhythm. For example, the user can perceive the size of the subject through vibration intensity, the complexity of the subject's details through the second vibration frequency, and the emotional state of the figures in the image through vibration rhythm.

[0135] For example, such as Figure 5 As shown in the flowchart, the 3D haptic feedback module can perform depth estimation on the image viewed by the user, process the depth estimation result to obtain the corresponding vibration intensity mapping image. The 3D haptic feedback module also performs edge detection on the image, and performs noise filtering and dilation operations on the edge detection result to obtain the second vibration frequency mapping image. Then, the 3D haptic feedback module superimposes the vibration intensity mapping image and the second vibration frequency mapping image into one channel to generate a high-order information map of the image. Different positions in this high-order information map can correspond to different vibration intensities and second vibration frequencies. Furthermore, the 3D haptic feedback module can also perform emotion recognition on the image, mapping different vibration rhythms (or different vibration curves) based on the overall emotional information of the image (such as happiness, sadness, etc.).

[0136] Optionally, the 3D haptic feedback module can output a vibration signal at a user's touch location, wherein the vibration signal has at least one of vibration intensity, a second vibration frequency, or a vibration rhythm. Alternatively, the vibration signal can be described as a vibration feedback mode composed of at least one of vibration intensity, a second vibration frequency, or a vibration rhythm.

[0137] In some embodiments, the vibration feedback mode V can be a synthesis of parameters such as vibration intensity A(S, Z), second vibration frequency F(D), and vibration rhythm P(E) in different dimensions. V can be a vibration effect determined by vibration intensity, second vibration frequency, and vibration rhythm, obtained based on the area S of the subject in the image, the depth information Z of the subject in the image, the detail (or detail complexity) D of the area where the touch position is located, and the image's mood E.

[0138] Alternatively, by reasonably adjusting the weights and parameters of each element, users can perceive important information in the image through touch, such as the presence of people, the outlines of objects, and the emotional atmosphere.

[0139] Specifically, the vibration intensity A(S, Z) can satisfy the following formula:

[0140] A(S, Z) = (A min +(A max -A min )×S / S max )×(1-Z / Z max Formula 2.

[0141] Among them, A max For the preset maximum vibration intensity, A min S is the preset minimum vibration intensity. max Z is the preset maximum area. max This is the preset maximum depth of field.

[0142] Vibration intensity can be controlled by the area of ​​the subject in the image and the depth of field. Here, the vibration intensity A(S, Z) is designed as a function of the area S of the subject and the depth Z of the field, where the area determines the basic intensity of the vibration, while the depth of field can adjust the final intensity of the vibration, simulating the perception of the subject's distance in space.

[0143] The second vibration frequency F(D) can satisfy the following formula:

[0144] F(D)=F min +(F max -F min )×(D / D max Formula 3.

[0145] Among them, F max F is the preset maximum second vibration frequency. min D is the preset minimum second vibration frequency. max This represents the preset maximum detail complexity.

[0146] As can be seen from Formula 3, the second vibration frequency is controlled by the complexity of detail. The more complex the detail (e.g., more edges, textures, etc.), the higher the second vibration frequency. A higher second vibration frequency allows the user to perceive the complexity or richness of detail of the subject. For example, if the hair of a person in an image has more texture, meaning the hair has greater detail complexity, and the face has less texture, meaning the face has lower detail complexity, the second vibration frequency when the user touches the hair will be higher than the second vibration frequency when the user touches the face.

[0147] The vibration rhythm can be determined by the emotional information in an image (such as the emotions of the people in the image). The vibration rhythm can be preset and can change according to the intensity of the emotion; that is, the switching frequency (rhythm) of the vibration will change according to the intensity of the emotion. It can also be understood that the vibration rhythm is positively correlated with the intensity of the emotion corresponding to the emotional information in the image. For example, when the emotion in the image is more intense (such as anger or excitement), the vibration rhythm will speed up, producing a faster pulse sensation; when the emotion is more calm (such as relaxation or happiness), the vibration rhythm will slow down, producing a softer, intermittent vibration.

[0148] Optionally, the vibration rhythm can be periodic. For example, when emotions are intense, the vibration rhythm can be continuous, achieving a faster vibration pace. When emotions are calm, the vibration rhythm can be a cycle of vibrating a few times, pausing for a while, and then vibrating a few times again, achieving a gentler vibration. It should be understood that the above are merely examples, and there may be other examples of different vibration rhythms, which this application does not limit.

[0149] (4) Single Image Feature Description Module

[0150] The single-image feature description module obtains the contextual information of the image based on the understanding of the image viewed by the user, the image shooting information and the relevant content of the subject in the image, so that the user can perceive a richer content in the image.

[0151] For example, such as Figure 6 As shown, when a user browses an image, this module can perform image matting to obtain key information (such as the main subject and / or objects). Based on the information obtained from the matting, it performs image understanding to determine if the image contains people and / or objects, as well as behavioral information of the people or objects. This module can also obtain image information such as the time and location when the image was captured. Furthermore, based on the information obtained from the matting, the module can perform face detection, expression and emotion detection, and, combined with a feature database, match the identities and relationships of people. Further, this module can perform text aggregation on the information obtained above to obtain a characteristic image description of the image.

[0152] For example, when a user browses an image, the module can extract relevant information about the people in the image from a feature database, such as their birthday, home address, and related notes. It can then determine the background of the image activity by combining this information with the time and location of the photo. Finally, it can organize the fragmented information and output a high-level single-image feature description, such as "This is a photo of your birthday party."

[0153] Optionally, the module does not distinguish the order of operations for image understanding, image information acquisition, and feature database matching. Any operation can be performed first, or they can be obtained simultaneously. This application does not limit this.

[0154] In some embodiments, the image description obtained by the single-image feature description module can be output to the user via voice.

[0155] (5) Scene-based image presentation module

[0156] The contextualized image presentation module can match multiple images associated with a single image viewed by a user, and generate contextualized descriptions based on these images, such as sequential contextualized descriptions or comparative descriptions across timelines.

[0157] For example, such as Figure 7 As shown, this module analyzes images viewed by the user and combines them with other related images (such as relevant time, location, people, etc.). Based on a feature database, it clusters and matches a series of images taken at the same time and place, and then performs text aggregation and summarization to produce a contextualized image presentation. For example, the module matches all photos taken in the same active scene and provides a contextualized description of the image content information in chronological order. Another example is that, based on the user's habitual shooting habits—taking photos of the same object or person at different times and locations—the module provides users with comparative descriptive suggestions across timelines.

[0158] In this way, the module not only provides basic information about objects and people, but also generates high-level semantic descriptions through techniques such as sentiment analysis and scene reasoning, helping users perceive the emotional tension or story background behind the image. For example, when the image shows a birthday party scene, the module can match relevant event and key person information to scan related images in the image library to create a connected, contextualized story description, such as "This is a birthday party for XXX held at XXXX place with which friends, and you all happily cut the cake and played XXXX games at the party, etc."

[0159] In some embodiments, the image description obtained by the contextualized image presentation module can be output to the user via voice.

[0160] In this application, the electronic device can provide the user with image understanding content based on at least one of the five modules described above. In some embodiments, the electronic device can use a particular module based on the user's selection. Some examples of image understanding are described below.

[0161] In one example a1, when a user is browsing an image, as the user's touch position moves outside the subject, the electronic device can output vibration signals with different first vibration frequencies based on the distance between the user's touch position and the first edge point of a subject. The closer the touch position is to the subject, the higher the first vibration frequency of the output vibration signal, until the first vibration frequency reaches its maximum when the user touches a subject.

[0162] For example, such as Figure 8 As shown, when a user browses an image, as the user's touch position moves from touch position 1, touch position 2, touch position 3 to touch position 4, the first vibration frequency of the vibration signal output by the electronic device continuously increases. That is, the first vibration frequency at touch position 2 is greater than that at touch position 1, the first vibration frequency at touch position 3 is greater than that at touch position 2, the first vibration frequency at touch position 4 is greater than that at touch position 3, and the first vibration frequency is the highest at touch position 4. Once the user touches a person, the first vibration frequency of the vibration signal no longer changes when touching different positions within the person. In this way, the electronic device can guide the user to locate important information positions in the image using different first vibration frequencies.

[0163] Figure 8 Let's take a person as the main subject in the image as an example. The process of the electronic device guiding the user to locate an animal in the image is similar. For both animals and people, the electronic device can output vibration signals with varying frequencies based on the importance of the subject. For instance, when the user's touch position is closer to an animal, the first vibration frequency changes more slowly; when the user's touch position is closer to a person, the first vibration frequency changes more quickly.

[0164] In some implementations, when a user touches a subject, the electronic device can also output a description of which subject it is via voice, so that the user can understand the image jointly through touch and hearing.

[0165] In example a1, the electronic device can be implemented using the aforementioned perception navigation module.

[0166] In another example, a2, when a user browses an image, the electronic device first guides the user to locate important information positions (such as the position of the main figure) in the image by outputting vibration signals of different first vibration frequencies. The implementation process can be found in example a1. Furthermore, the electronic device can output vibration signals with vibration intensity, a second vibration frequency, and a vibration rhythm based on the area of ​​the subject, the depth information of the subject in the image, the detail complexity of the area where the touch point is located, and the emotional information of the image. This allows the user to perceive information in the image in a three-dimensional way through touch. Through the linkage of the perception navigation module and the three-dimensional haptic feedback module, the electronic device provides visually impaired users with a complete haptic experience flow from "locating important information in the image" to "deeply perceiving image details and emotions." Visually impaired users do not need visual input; they can perceive key information and emotional content in the image in a three-dimensional way solely through haptic feedback, thereby significantly improving their understanding of images and their interactive experience.

[0167] In some implementations, when a user touches a subject, the electronic device can also output a description of which subject it is via voice.

[0168] Optionally, when a user touches a subject, the electronic device can provide a voice prompt indicating that the subject has been touched. When the user responds with a voice message to switch to the 3D haptic feedback mode or when the user performs a continuous click operation, the electronic device outputs a vibration signal with vibration intensity, a second vibration frequency, and vibration rhythm.

[0169] The continuous clicking operation can be a series of clicks, such as clicking twice or three times consecutively; this application does not limit this. Optionally, in addition to the above, the user can also trigger the electronic device to output a vibration signal with vibration intensity, a second vibration frequency, and a vibration rhythm through other operations; this application does not limit this as well.

[0170] Optionally, the voice prompt from the electronic device indicating that a subject has been touched can also be understood as asking whether to switch to a 3D haptic feedback mode.

[0171] When the user's voice response does not switch or the user does not provide any feedback, the electronic device will not perform any further operations; that is, it will remain at the point where the user touches the main body.

[0172] It should be noted that the 3D haptic feedback mode is only an example and can be replaced by other modes or descriptions.

[0173] In example a2, the relationship between the first vibration frequency and the second vibration frequency is not limited. The first vibration frequency is the frequency of the vibration signal when the touch position is outside the subject, and the second vibration frequency is the frequency of the vibration signal when the touch position is inside the subject after it has been located. The first and second vibration frequencies are the vibration frequencies of the electronic device in different modes. The first vibration frequency corresponds to the sensory navigation mode, and the second vibration frequency corresponds to the 3D haptic feedback mode. Of course, the sensory navigation mode here is only an example and can be replaced by other modes or other descriptions.

[0174] In another example, a3, when a user browses an image, the electronic device can combine the people, objects, and actions in the image, as well as the time and location when the image was taken, to match the relationships between people and the scene to obtain the image's contextual content, thus providing a detailed description of the image. Furthermore, the electronic device outputs this detailed description of the image to the user via voice, allowing the user to perceive the emotions in the image, the relationships between people / objects, the context of the scene, and other higher-order semantic information. For example, the electronic device might output, "This is a photo of your birthday party."

[0175] In example a3, the electronic device can be implemented using the aforementioned single-image feature description module.

[0176] In another example, a4, when a user browses an image, the electronic device can first determine the descriptive information in the image, and then match it with multiple images taken in the same activity scene associated with that image to create contextualized descriptions of the multiple images. For example, for an image containing multiple people interacting, the electronic device can not only identify the people, but also describe the interaction relationships between them, the theme of the activity, and interpret their facial expressions, body language, and the context in which they are situated.

[0177] Specifically, electronic devices can match all photos taken in the same live scene to a feature database, providing contextualized descriptions of the image content in chronological order. Alternatively, based on a user's habitual shooting patterns—taking photos of the same object or person at different times and locations—the device can offer comparative descriptions across timelines. For example, if a user has a daily habit of watering and photographing their plants, the electronic device can describe the differences between the current photo and previous ones, and provide relevant watering and fertilization suggestions based on the plant's condition. This semantically enhanced image description will significantly improve visually impaired users' understanding and perception of images.

[0178] In the process described above, electronic devices can use a single image feature description module and a contextualized image presentation module to contextualize a single image and multiple related images.

[0179] Optionally, after determining the description of a single image, the electronic device can output a description via voice. This voice-output description can be used to request the user to select whether to switch to a contextualized image presentation mode. When the user responds with a voice message to switch or performs a series of clicks, the electronic device provides contextualized voice descriptions for multiple subsequent related images.

[0180] For details on continuous clicking, please refer to the previous description, which will not be repeated here.

[0181] When the user's voice response does not switch or the user provides no feedback, the electronic device will no longer provide contextualized voice descriptions of multiple related images; it will remain at the single-image description level.

[0182] It should be noted that the scene-based image presentation mode is only an example and can be replaced by other modes or descriptions.

[0183] In another example, a5, when a user browses an image, the electronic device can simultaneously provide vibration signals and voice output a description of the first image. Specifically, example a1 can be implemented in combination with example a3, or example a1 can be implemented in combination with example a4, or example a2 can be implemented in combination with example a3, or example a2 can be implemented in combination with example a4. This example a5 allows the user to fully understand the image through touch and hearing.

[0184] The above implementation methods enable visually impaired users to achieve a detailed understanding of images through various means. In some embodiments, for non-visually impaired users, haptic feedback can also assist sighted users in performing shallow-level image interaction operations. For example, users often encounter interactive scenarios in daily image operations, such as wanting to cut out a specific subject from an image. Current electronic devices often only allow interaction in certain areas of the image (such as the subject area), and only provide feedback when the user operates in that area. In this application, the electronic device can use a perception navigation module to help users quickly find interactive areas, assisting them in performing the next image operation.

[0185] Based on the above, the flow of the image understanding method provided in the embodiments of this application will be described in detail below.

[0186] For example, Figure 9 A flowchart of an image understanding method is shown, which may include the following steps:

[0187] Step 901: The electronic device acquires the first image.

[0188] When a user views the first image on the screen of an electronic device, the electronic device can acquire the first image.

[0189] Step 902: The electronic device outputs a vibration signal based on the user's first touch position in the first image. The vibration signal is used to indicate image information at the first touch position in the first image or to indicate the distance between the first touch position and at least one target subject in the first image.

[0190] In some embodiments, when the first touch position is located outside at least one target subject, the vibration signal may have a first vibration frequency, which is used to indicate the distance between the first touch position in the first image and at least one target subject.

[0191] The distance between the first touch position and any one of the at least one target body can be the distance between the first touch position and the first edge point, where the first edge point is the edge point on the target body that is closest to the first touch position.

[0192] When the first image contains a target subject, the magnitude of the first vibration frequency is negatively correlated with the distance between the first touch position and the target subject. That is, the closer the first touch position is to the target subject, the higher the first vibration frequency; the farther the first touch position is from the target subject, the lower the first vibration frequency.

[0193] In the case where the first image contains multiple target subjects, the magnitude of the first vibration frequency and the first touch position are related to the distance between each of the multiple target subjects and the weights corresponding to the multiple target subjects respectively.

[0194] When the first image contains multiple target objects, the magnitude of the first vibration frequency changes differently when the distance between the first touch position and different target objects changes the same. For example, the frequency at which the first vibration frequency increases as the distance between the first touch position and the first target object becomes closer is different from the frequency at which the first vibration frequency increases as the distance between the first touch position and the second target object becomes closer.

[0195] Optionally, the change in the first vibration frequency can be differentiated based on the importance of different subjects when the distance between the first touch position and different target subjects changes.

[0196] In some embodiments, when the first image contains multiple subjects, the first vibration frequency is different when the first touch position is at the same distance from different target subjects.

[0197] In some embodiments, the first vibration frequency can be referred to the relevant description in Formula 1 above, and will not be repeated here.

[0198] Based on the above method, the user can be guided to touch a target object by the change in the first vibration frequency. When the user's first touch position is within a target object, the first vibration frequency reaches its maximum, and within that target object, the first vibration frequency perceived by the user no longer changes. Optionally, at this time, the electronic device can provide voice prompts to the user to touch a target object.

[0199] Optionally, when the electronic device outputs a vibration signal with a first vibration frequency, it can be implemented based on the aforementioned sensing and navigation module. For details, please refer to the aforementioned description, which will not be described in detail here.

[0200] Specifically, the process by which the electronic device guides the user to touch the target subject in the first image can also be found in the description of example a1 above, and will not be repeated here.

[0201] In some embodiments, when the first touch position is located within the first target subject, i.e., after the user touches the first target subject, the vibration signal may have at least one of the following: vibration intensity, second vibration frequency, or vibration rhythm; the vibration intensity is used to characterize the area and depth of the first target subject, the second vibration frequency is used to characterize the detail complexity of the area where the first touch position is located, and the vibration rhythm is used to characterize the image emotion information corresponding to the first image. The first target subject is any one of at least one target subject in the first image.

[0202] Specifically, a vibration signal possessing at least one of the following characteristics—vibration intensity, a second vibration frequency, or a vibration rhythm—can refer to the aforementioned descriptions. Vibration intensity can refer to the description in Formula 2, the second vibration frequency can refer to the description in Formula 3, and the vibration rhythm can also refer to the aforementioned descriptions, which will not be repeated here.

[0203] In one optional implementation, when a user touches the first target object, the electronic device can output first information to indicate that the first target object has been touched. After receiving second information, the electronic device can output a vibration signal having at least one of the following: vibration intensity, a second vibration frequency, or a vibration rhythm. The second information is used to indicate switching to a first mode, where the vibration signal has at least one of the following: vibration intensity, a second vibration frequency, or a vibration rhythm. The first information may be a voice prompt output by the electronic device, and the second information may be a voice response from the user, information triggered by continuous clicking, or information triggered by other operations.

[0204] The first mode can also be called a three-dimensional haptic feedback mode or other modes, which are not limited in this application.

[0205] The relationship between the first and second vibration frequencies is not limited. The first vibration frequency is the frequency of the vibration signal when the touch position is outside the target body, and the second vibration frequency is the frequency of the vibration signal when the touch position is inside the first target body after the touch position has been located. The first and second vibration frequencies are the vibration frequencies of the electronic device in different modes. The first vibration frequency corresponds to the perception navigation mode, and the second vibration frequency corresponds to the 3D haptic feedback mode. Of course, the perception navigation mode here is only an example and can be replaced by other modes or other descriptions.

[0206] Optionally, before and after switching modes, users may feel a jump in frequency or other tactile sensations, which is not limited in this application.

[0207] Optionally, when the electronic device outputs a vibration signal having at least one of the following: vibration intensity, second vibration frequency, or vibration rhythm, it can be implemented based on a three-dimensional haptic feedback module, as described above, and will not be detailed here.

[0208] Specifically, the process by which an electronic device guides a user to touch a target subject in the first image and provides further feedback on the details of that target subject and the mood of the image can also be found in the description of example a2 above, and will not be repeated here.

[0209] based on Figure 9 The image understanding method shown adds a tactile aspect to the image understanding, allowing users to perceive information in the image through touch, thus improving the user experience.

[0210] For example, Figure 10 A flowchart of another image understanding method is shown, which may include the following steps:

[0211] Step 1001: The electronic device detects the first operation for the first image.

[0212] In some embodiments, a user's first operation on the first image may include clicking or swiping the first image. That is, when a user views the first image, clicking or swiping it, the electronic device can detect the user's action, which constitutes the first operation.

[0213] Step 1002: In response to the first operation, the electronic device outputs first image description information via voice, the first image description information being determined based on user-related information and the first image.

[0214] For example, user-related information may include one or more of the following: personal identification information, personal date of birth information, personal address information, personal habit information, interpersonal relationship information, or user-related tag information, etc.

[0215] In one alternative implementation, the first image description information may be generated by the electronic device after the first operation is detected or before the first operation is detected.

[0216] For example, an electronic device can generate first image description information when a user views the first image, that is, generate the first image description information before detecting the first operation, and output the first image description information after detecting the user's first operation.

[0217] For example, after a user views the first image and performs the first operation, the electronic device can generate first image description information in response to the first operation, that is, generate first image description information after detecting the first operation and output the first image description information.

[0218] Optionally, the first image description information is not limited to being generated by an electronic device, but can also be generated by other devices (such as a server), and the electronic device obtains it from other devices.

[0219] In some embodiments, the determination of the first image description information based on user-related information and the first image can specifically be as follows: the first image description information is determined based on the context information of the content in the first image and the first image, and the context information of the content in the first image is determined based on user-related information.

[0220] Optionally, the contextual information of the content in the first image can be determined based on a feature database, which includes user-related information.

[0221] For example, an electronic device may pre-generate this feature database. For instance, the electronic device may generate the feature database based on one or more of the following information: information entered by the user, data from a contact list, or data from a photo album.

[0222] Optionally, electronic devices can generate a feature database through the aforementioned accessibility feature library construction module, as detailed in the aforementioned descriptions, which will not be repeated here.

[0223] In one optional implementation, the contextual information of the content in the first image may include information about the first subject. Optionally, the electronic device may determine the first subject in the first image and determine the information of the first subject based on user-related information, thereby realizing the determination of the contextual information of the content in the first image based on user-related information.

[0224] For example, the information of the first subject may include one or more of the following: the identity of the first subject, the birth date information of the first subject, the address information of the first subject, or the tag information of the first subject.

[0225] Optionally, when the first subject is a person, the electronic device can perform facial recognition on the first image to determine the first subject in the first image.

[0226] Optionally, when the first subject is an object, the electronic device can use methods such as target detection or feature extraction to determine the first subject in the first image.

[0227] Optionally, the first image description information is determined based on the context information of the content in the first image and the first image itself, and may include: the first image description information is determined based on the information of the first subject, the shooting time and location corresponding to the first image, and the content contained in the first image. For example, the first image description information may be "This is a photo of your friend's birthday party."

[0228] The aforementioned first image description information can be a detailed description of the first image. For example, see example a3 above. The acquisition of the first image description information can be achieved through the aforementioned single-image feature description module, as described above, and will not be repeated here.

[0229] In some embodiments, after the electronic device outputs a description corresponding to the first image, it can further combine multiple images associated with the first image to provide a contextualized description, thereby providing the user with a more vivid description.

[0230] For example, an electronic device can output second image description information via voice; the second image description information is determined based on description information corresponding to multiple images associated with the first image and the first image description information.

[0231] Optionally, the electronic device may determine multiple images associated with the first image based on the first image description information.

[0232] Optionally, the electronic device can determine multiple images associated with the first image based on the first image description information by means of the following method: the electronic device determines multiple images associated with the first image based on the first image description information and a feature database.

[0233] In some embodiments, the electronic device may also determine a plurality of images associated with the first image based on the first image, in a manner similar to determining a plurality of images associated with the first image based on the first image description information, and these methods may be referenced in turn.

[0234] Optionally, the multiple images and the first image have at least one of the following relationships: the multiple images and the first image are images from the same activity scene; the multiple images and the first image are images taken of the same subject; the multiple images and the first image are images taken at the same location at different times; or the multiple images and the first image are images taken during the same time period.

[0235] For example, if the first image is a photo of a user's friend's birthday party, the electronic device can identify all images taken within that birthday party setting and provide a chronological, contextualized, and narrative description of the image content. For instance, the electronic device could output, "This is from [user's name]'s birthday party, taken at [location] with [number] friends. You were happily cutting the cake and playing [game] together." This method not only provides basic information about objects and people but also uses techniques like sentiment analysis and scene reasoning to generate higher-level semantic descriptions, helping users perceive the emotional tension or backstory behind the images.

[0236] For example, see Example a4 above, which describes how an electronic device generates a contextualized description of a first image.

[0237] Optionally, the electronic device can generate a contextualized description for the first image through a contextualized image presentation module, as described above, and will not be repeated here.

[0238] In one alternative implementation, the electronic device may, in response to receiving third information, output second image description information via voice.

[0239] The third piece of information can be the user's voice information, information triggered by the user performing continuous clicks, or information executed by the user performing other operations.

[0240] Optionally, before receiving the third information, the electronic device may output a fourth information, which is used to request the user to select whether to switch to the second mode; the third information received by the electronic device is used to indicate switching to the second mode. The fourth information may be a voice prompt output by the electronic device, and the second mode may also be called a contextualized image presentation mode or other modes; this application does not limit this.

[0241] In some implementations, the output of the first image description information by the electronic device can be regarded as the output of the fourth information. Of course, the electronic device can also output the fourth information in other ways, which is not limited in this application.

[0242] Through the above, the user can actively trigger the electronic device to output the second image description information by voice, or trigger the electronic device to output the second image description information by voice based on the fourth information output by the electronic device. This application does not limit this.

[0243] based on Figure 10 The image understanding method shown can provide users with detailed image descriptions by combining relevant user information, thereby improving the user experience.

[0244] In some embodiments, it is also possible to Figure 9 The illustrated embodiments and Figure 10 The methods in the illustrated embodiments combine to enable users to fully perceive the detailed content in the image through touch and hearing, thereby enhancing the user experience.

[0245] It should be noted that in the above embodiments, the concepts and explanations of the same terms can be referenced to each other, and the same or similar steps can also be referenced to each other.

[0246] It should also be noted that each step in the above embodiments can be executed by the corresponding device, or by components such as chips, processors, or chip systems within that device. The embodiments of this application do not limit their execution. The above embodiments are merely illustrative examples of execution by the corresponding device. Furthermore, the specific implementation methods or examples in the above embodiments do not limit the solutions provided by the embodiments of this application.

[0247] It is understood that, in order to achieve the functions described in the above embodiments, each device involved in the above embodiments includes a hardware structure and / or software module corresponding to perform each function. Those skilled in the art should readily recognize that, based on the units and method steps of the various examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0248] It is understood that the architecture and application scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this invention. As those skilled in the art will know, with the evolution of network architecture and the emergence of new services, the technical solutions provided by the embodiments of this invention are also applicable to similar technical problems.

[0249] It should be noted that the "steps" in the embodiments of this application are merely illustrative and are intended to better understand one method of presentation used in the embodiments. They do not constitute a substantial limitation on the execution of the solution of this application. For example, the "step" can also be understood as a "feature". Furthermore, the steps do not constitute any limitation on the execution order of the solution of this application. Any changes to the order of steps, or the merging or splitting of steps made on this basis without affecting the overall solution implementation, resulting in a new technical solution, are also within the scope of disclosure of this application.

[0250] Based on the same concept, embodiments of this application also provide an electronic device. For example... Figure 11 As shown, the electronic device 1100 includes a display screen 1101, one or more processors 1102, and a memory 1103. Optionally, the electronic device 1100 may also include a transceiver 1104. For example, the above-mentioned components can be connected via one or more communication buses. The one or more computer programs are stored in the memory 1103 and configured to be executed by the one or more processors 1102. The one or more computer programs include instructions that can be used to cause the electronic device 1100 to perform various steps of the methods in the above embodiments.

[0251] For example, the above-mentioned one or more processors 1102 may specifically be Figure 1 The processor 110; the memory 1103 mentioned above can specifically be... Figure 1 The internal memory 121 is included. The transceiver 1104 can be... Figure 1 The mobile communication module 150 or wireless communication module 160 is included; the display screen 1101 can be... Figure 1 The display screen 194 in this embodiment is not limited in any way in this application.

[0252] Based on the same concept, embodiments of this application also provide an electronic device, which includes units for performing the various steps of the methods provided in the above embodiments.

[0253] For details on the specific functions of each module, please refer to the above embodiments; they will not be repeated here.

[0254] Based on the above embodiments, this application also provides a computer program product, which includes instructions; when the computer program is run on a computer, it causes the computer to execute the method provided in the above embodiments.

[0255] Based on the above embodiments, this application also provides a computer-readable storage medium storing computer program instructions, which, when executed by a computer, cause the computer to perform the methods provided in the above embodiments.

[0256] Optionally, the aforementioned computer may include, but is not limited to, control devices.

[0257] The storage medium can be any available medium that a computer can access. For example, but not limited to, a computer-readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer.

[0258] Based on the above embodiments, this application also provides a chip for reading a computer program stored in a memory to implement the method provided in the above embodiments. Optionally, the chip may include a processor and a memory, wherein the processor is coupled to the memory and is used to read the computer program stored in the memory to implement the method provided in the above embodiments.

[0259] Based on the above embodiments, this application provides a chip system including a processor for supporting a computer device in implementing the functions involved in the control device in the above embodiments. In one possible design, the chip system further includes a memory for storing necessary programs and data of the computer device. This chip system may be composed of chips or may include chips and other discrete components.

[0260] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0261] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0262] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0263] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

Claims

1. An image understanding method applied to electronic devices, characterized in that, include: Get the first image; Based on the user's first touch position in the first image, a vibration signal is output, which is used to indicate image information at the first touch position in the first image or to indicate the distance between the first touch position and at least one target subject in the first image.

2. The method as described in claim 1, characterized in that, The first touch position is located outside the at least one target subject, and the vibration signal has a first vibration frequency, which is used to indicate the distance between the first touch position and the at least one target subject in the first image.

3. The method as described in claim 2, characterized in that, The distance between the first touch position and any one of the at least one target body is the distance between the first touch position and the first edge point, where the first edge point is the edge point on any one of the target bodies that is closest to the first touch position.

4. The method as described in claim 2 or 3, characterized in that, When the at least one target subject is a single target subject, the magnitude of the first vibration frequency and the distance between the first touch position and the single target subject are negatively correlated. or When the at least one target subject is multiple target subjects, the magnitude of the first vibration frequency and the distance between the first touch position and any one of the multiple target subjects, as well as the weights corresponding to the multiple target subjects respectively, are related.

5. The method according to any one of claims 2-4, characterized in that, When the first image contains multiple target subjects, and the distance between the first touch position and different target subjects changes in the same way, the corresponding changes in the magnitude of the first vibration frequency are different.

6. The method as described in claim 1, characterized in that, The first touch position is located within a first target subject, which is any one of the at least one target subject; the vibration signal has at least one of the following: vibration intensity, second vibration frequency, or vibration rhythm; the vibration intensity is used to characterize the area and depth of the first target subject, the second vibration frequency is used to characterize the detail complexity of the area where the first touch position is located, and the vibration rhythm is used to characterize the image emotion information corresponding to the first image.

7. The method as described in claim 6, characterized in that, The vibration rhythm is positively correlated with the intensity of the emotion corresponding to the emotional information in the image.

8. The method as described in claim 6 or 7, characterized in that, The emotional information in the image includes one or more of the following: intense emotion or calm emotion.

9. The method according to any one of claims 6-8, characterized in that, The vibration rhythm is periodic.

10. The method according to any one of claims 6-9, characterized in that, The method further includes: Output first information, which is used to indicate that the first target body has been touched; Upon receiving a second message, the second message is used to indicate switching to a first mode; the vibration signal in the first mode has at least one of the following: the vibration intensity, the second vibration frequency, or the vibration rhythm.

11. An image understanding method applied to electronic devices, characterized in that, include: A first operation targeting the first image was detected. In response to the first operation, first image description information is output via voice, the first image description information being determined based on user-related information and the first image.

12. The method as described in claim 11, characterized in that, The first image description information is generated after the first operation is detected or before the first operation is detected.

13. The method as described in claim 11 or 12, characterized in that, The first operation on the first image includes: clicking the first image or swiping the first image.

14. The method according to any one of claims 11-13, characterized in that, The first image description information is determined based on user-related information and the first image, including: The first image description information is determined based on the context information of the content in the first image and the first image itself. The context information of the content in the first image is determined based on the user-related information.

15. The method as described in claim 14, characterized in that, The contextual information of the content in the first image is determined based on the user-related information, including: The contextual information of the content in the first image is determined based on a feature database, which includes user-related information.

16. The method as described in claim 14 or 15, characterized in that, The contextual information of the content in the first image includes information about the first subject.

17. The method as described in claim 16, characterized in that, The method further includes: Identify the first subject in the first image; The information of the first subject is determined based on the user-related information.

18. The method as described in claim 16 or 17, characterized in that, The first image description information is determined based on the context information of the content in the first image and the first image itself, and includes: The first image description information is determined based on the information of the first subject, the shooting time and location corresponding to the first image, and the content contained in the first image.

19. The method according to any one of claims 16-18, characterized in that, The information of the first subject includes one or more of the following: the identity of the first subject, the birth date information of the first subject, the address information of the first subject, or the tag information of the first subject.

20. The method according to any one of claims 11-19, characterized in that, The user-related information includes one or more of the following: personal identification information, personal birthday information, personal address information, personal habit information, interpersonal relationship information, or user-related tag information.

21. The method according to any one of claims 11-20, characterized in that, The method further includes: The second image description information is output via voice; the second image description information is determined based on the description information corresponding to multiple images associated with the first image and the first image description information.

22. The method as described in claim 21, characterized in that, The method further includes: Based on the first image description information and a feature database, multiple images associated with the first image are determined, wherein the feature database includes user-related information.

23. The method as described in claim 21 or 22, characterized in that, The plurality of images have at least one of the following relationships with the first image: The multiple images and the first image are images from the same activity scene; The plurality of images and the first image are images taken of the same subject; The multiple images and the first image are images taken at the same location at different times; The multiple images and the first image were taken within the same time period.

24. The method according to any one of claims 21-23, characterized in that, The second image description information is output via voice, including: In response to receiving the third information, the second image description information is output via voice.

25. The method as described in claim 15 or 22, characterized in that, The method further includes: A feature database is generated based on one or more of the following information: Information entered by the user, data obtained from the address book, or data obtained from the photo album.

26. An electronic device, characterized in that, The electronic device includes: Memory is used to store computer program instructions; A processor for executing the computer program instructions to support the electronic device in implementing the method as described in any one of claims 1-10, or in implementing the method as described in any one of claims 11-25.