Image recognition method and related device

CN116152814BActive Publication Date: 2026-09-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211640349.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-20
Publication Date
2026-09-11
Estimated Expiration
2042-12-20

AI Technical Summary

Technical Problem

[0004]然而,用户需要在4自由度下移动(3个位移自由度,1个转动自由度),如“向前移动1英尺”“向左移动1英尺”,“旋转到五点钟方向”在移动时容易偏移目标,出错率较高,对于盲人用户来说,无法精准量化自己移动的距离和旋转的角度,不能做出引导语中的精确动作,有时会造成目标偏离程度反而增大

Benefits of technology

[0056] This application embodiment prompts the user to establish a positional association between the assistive part and the object to be identified. Since visually impaired users can perceive the positional relationship between the assistive part and the object to be identified, as well as the positional relationship between the assistive part and the terminal device, through proprioception, the spatial alignment between the terminal and the object to be identified in three degrees of freedom can be maintained. Only the position of the terminal in the vertical direction needs to be adjusted, which reduces the user's action cost and improves the recognition efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152814B_ABST
    Figure CN116152814B_ABST
Patent Text Reader

Abstract

An image recognition method comprises: outputting a first reminder; the first reminder instructs a user to associate an auxiliary part with a to-be-recognized object in position, and controls a terminal to capture the auxiliary part; in a case where the auxiliary part exists in a captured first image, and a target object in the first image has a positional relationship with the auxiliary part that satisfies a first preset condition, an identification result of the target object is obtained according to a captured second image; the first image and the second image are images in a video stream captured by the terminal after the first reminder is output, and the second image is captured after the first image. By prompting the user to associate the auxiliary part with the to-be-recognized object in position, the cost of user action is reduced, and the efficiency of recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more particularly to an image recognition method and related equipment. Background Technology

[0002] In daily life, visually impaired individuals have numerous needs for recognizing textual information in near-field environments, such as recipient information on express delivery slips and names, usage instructions, and dosages on medicine leaflets. Currently, optical character recognition (OCR) and text-to-speech (TTS) technologies enable visually impaired individuals to obtain near-field textual information through terminal devices. However, when using information recognition software equipped with OCR and TTS technologies, visually impaired individuals still encounter problems such as not being able to capture the text, capturing only a small portion, or capturing it unclearly due to the lack of visual feedback.

[0003] Therefore, existing technologies are beginning to explore how to help visually impaired individuals accurately and completely read text information in the areas they want to recognize using image capture devices. One existing implementation calculates the direction and distance the user should move their phone by monitoring the integrity of the document in the current frame in real time, and then uses voice guidance to assist the user.

[0004] However, users need to move in 4 degrees of freedom (3 displacement degrees of freedom and 1 rotation degree of freedom), such as "move forward 1 foot", "move left 1 foot", "rotate to the five o'clock position". When moving, it is easy to deviate from the target and the error rate is high. For blind users, it is impossible to accurately quantify the distance they move and the angle of rotation, and they cannot make the precise actions in the guidance. Sometimes, the deviation from the target will actually increase. Summary of the Invention

[0005] In a first aspect, this application provides an image recognition method, the method comprising: outputting a first reminder; the first reminder instructing a user to establish a positional association between an auxiliary part and an object to be recognized, and controlling a terminal to capture the auxiliary part; if the auxiliary part exists in the captured first image, and a target object exists in the first image whose positional relationship with the auxiliary part satisfies a first preset condition, obtaining a recognition result of the target object based on a captured second image; wherein the first image and the second image are images in a video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

[0006] This application prompts the user to establish a positional association between the assistive device and the object to be identified. Since visually impaired users can perceive the positional relationship between the assistive device and the object to be identified, as well as the positional relationship between the assistive device and the terminal device, through proprioception, the spatial alignment between the terminal and the object to be identified in three degrees of freedom can be maintained. Only the position of the terminal in the vertical direction needs to be adjusted, which reduces the user's action cost and improves the recognition efficiency.

[0007] Furthermore, by using the auxiliary parts as anchor points, the auxiliary parts are identified using computer vision, and the areas with spatial relationships with the auxiliary parts are defined as regions of interest. By utilizing the habitual interactive actions of visually impaired users in recognizing text in daily life, and through the proprioception of visually impaired users, they can quickly use handheld devices to locate the areas that need to be recognized. In addition, this application also significantly improves recognition efficiency in multi-target scenes and cluttered background scenes.

[0008] In one possible implementation, the auxiliary part is the hand.

[0009] In one possible implementation, the first preset condition includes at least one of the following: there is an overlap between the target object and the auxiliary part; the target object is in the direction indicated by the auxiliary part; the target object is the object closest to the auxiliary part among the plurality of objects included in the first image.

[0010] In one possible implementation, the video stream further includes a third image acquired before the first image; the method further includes: when no target object satisfying the first preset condition is found in the third image, outputting a second reminder, the second reminder instructing the user to unassociate the auxiliary part with the position of the object to be identified, or to move the auxiliary part toward the edge of the object to be identified; the second image is acquired after the output of the second reminder.

[0011] In one possible implementation, the method further includes: when the image of the target object in the first image is incomplete or unclear, outputting a third reminder, the third reminder instructing the user control terminal to move away from or closer to the object to be identified; the acquisition time of the second image is after the output of the third reminder.

[0012] In one possible implementation, the method further includes: based on the fact that the pose difference of the terminal when moving away from or nearing the object to be identified is greater than a threshold, outputting a fourth reminder according to the pose difference, wherein the fourth reminder instructs the user to control the terminal to adjust the pose, and the adjustment amount of the pose adjustment is related to the pose difference.

[0013] When photographing an object, there exists a spatial range defined by the relative position and angle of the camera and the object to be photographed. Within this spatial range, the information in the photograph taken by the camera can be well identified. As mentioned above, when guiding the user to move the shooting device to photograph the entire object, due to individual operating habits or the lack of a stable shooting device during movement, the terminal posture may deviate significantly from the initial terminal posture. The shooting device may no longer be able to reach the target position by moving up and down. Therefore, it is necessary to guide the user to restore the terminal posture.

[0014] During the correction process, if the detected change in the terminal's posture exceeds a certain angle, the user is prompted to recalibrate. Timely reminders when the user makes an incorrect action during adjustment reduce the probability of user errors. Furthermore, it allows for timely correction and restarting when errors are significant, avoiding endless corrections.

[0015] In one possible implementation, the object to be identified is a planar object, and the first reminder specifically instructs the user to cover the object to be identified with the auxiliary part; or, the object to be identified is a three-dimensional object, and the first reminder specifically instructs the user to pick up the object to be identified using the auxiliary part or to cover one face of the three-dimensional object with the auxiliary part.

[0016] In one possible implementation, the method further includes: when the auxiliary part exists in the captured first image and a target object exists in the first image whose positional relationship with the auxiliary part satisfies a first preset condition, outputting a fifth reminder, the fifth reminder instructing the user to release the positional association between the auxiliary part and the object to be identified; the second image is acquired after the output of the fifth reminder.

[0017] In one possible implementation, the target object is a screen, and the terminal includes a touch component; the recognition result is the text content corresponding to the target control on the screen; the method further includes: outputting the text content and receiving the user's selection for the target control; outputting a sixth reminder based on the relative position between the touch component and the target control, the sixth reminder instructing the user to control the terminal to adjust the position until the touch component touches the target control, and the adjustment amount of the position adjustment is related to the relative position.

[0018] In one possible implementation, the touch component is a bracket attached to the back of the terminal or a corner point on the terminal.

[0019] Secondly, this application provides an image recognition device, the device comprising:

[0020] The output module is used to output a first reminder; the first reminder instructs the user to establish a location association between the auxiliary part and the object to be identified, and to control the terminal to take a picture of the auxiliary part;

[0021] The recognition module is used to obtain the recognition result of the target object based on the acquired second image when the auxiliary part exists in the captured first image and a target object exists in the first image whose positional relationship with the auxiliary part meets a first preset condition.

[0022] Wherein, the first image and the second image are images from the video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

[0023] In one possible implementation, the auxiliary part is the hand.

[0024] In one possible implementation, the first preset condition includes at least one of the following:

[0025] There is an overlap between the target object and the auxiliary part;

[0026] The target object is in the direction indicated by the auxiliary part;

[0027] The target object is the object closest to the auxiliary part among the multiple objects included in the first image.

[0028] In one possible implementation, the video stream further includes a third image captured before the first image; the output module is further configured to:

[0029] When there is no target object in the third image that meets the first preset condition, a second reminder is output. The second reminder instructs the user to unlink the auxiliary part from the position of the object to be identified, or to move the auxiliary part toward the edge of the object to be identified.

[0030] The second image was acquired after the second reminder was output.

[0031] In one possible implementation, the output module is further configured to:

[0032] When the image of the target object in the first image is incomplete or unclear, a third reminder is output, which instructs the user control terminal to move away from or closer to the object to be identified.

[0033] The second image was acquired after the third reminder was output.

[0034] In one possible implementation, the output module is further configured to:

[0035] Based on the fact that when the terminal moves away from or nears the object to be identified, the difference in posture compared to before moving away from or nearing the object to be identified is greater than a threshold, a fourth reminder is output according to the posture difference. The fourth reminder instructs the user to control the terminal to adjust the posture, and the adjustment amount of the posture adjustment is related to the posture difference.

[0036] In one possible implementation,

[0037] The object to be identified is a planar object, and the first reminder specifically instructs the user to cover the object to be identified with the auxiliary part; or...

[0038] The object to be identified is a three-dimensional object, and the first reminder specifically instructs the user to pick up the object to be identified through the auxiliary part or to cover one surface of the three-dimensional object with the auxiliary part.

[0039] In one possible implementation, the output module is further configured to:

[0040] If the auxiliary part exists in the first image and there is a target object in the first image whose positional relationship with the auxiliary part meets the first preset condition, a fifth reminder is output, which instructs the user to release the positional association between the auxiliary part and the object to be identified.

[0041] The second image was acquired after the fifth reminder in the output.

[0042] In one possible implementation, the target object is a screen, and the terminal includes a touch component; the recognition result is the text content corresponding to the target control on the screen; the output module is further configured to:

[0043] Output the text content;

[0044] The device further includes: a receiving module, configured to receive a user's selection of the target control;

[0045] The output module is also used for:

[0046] Based on the relative position between the touch component and the target control, a sixth reminder is output, instructing the user control terminal to adjust the position until the touch component contacts the target control, and the adjustment amount is related to the relative position; or,

[0047] According to the target control.

[0048] In one possible implementation, the touch component is a bracket attached to the back of the terminal or a corner point on the terminal.

[0049] Thirdly, this application provides an image recognition device, including: a processor, a memory, a camera, and a bus, wherein: the processor, the memory, and the camera are connected via the bus;

[0050] The camera is used to capture video in real time;

[0051] The memory is used to store computer programs or instructions;

[0052] The processor is used to call or execute programs or instructions stored in the memory, and is also used to call the camera to implement the steps described in the first aspect and any possible implementation of the first aspect.

[0053] Fourthly, this application provides a computer storage medium including computer instructions that, when executed on an electronic device or server, perform the steps described in the first aspect and any possible implementation thereof.

[0054] Fifthly, this application provides a computer program product that, when run on an electronic device or server, performs the steps described in the first aspect and any possible implementation thereof.

[0055] Sixthly, this application provides a chip system including a processor for supporting an execution device or training device in implementing the functions involved in the foregoing aspects, such as transmitting or processing data involved in the foregoing methods; or, information. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the execution device or training device. This chip system may be composed of chips or may include chips and other discrete devices.

[0056] This application embodiment prompts the user to establish a positional association between the assistive part and the object to be identified. Since visually impaired users can perceive the positional relationship between the assistive part and the object to be identified, as well as the positional relationship between the assistive part and the terminal device, through proprioception, the spatial alignment between the terminal and the object to be identified in three degrees of freedom can be maintained. Only the position of the terminal in the vertical direction needs to be adjusted, which reduces the user's action cost and improves the recognition efficiency.

[0057] Furthermore, by using the auxiliary parts as anchor points, the auxiliary parts are identified using computer vision, and the areas with spatial relationships with the auxiliary parts are defined as regions of interest. By utilizing the habitual interactive actions of visually impaired users in recognizing text in daily life, and through the proprioception of visually impaired users, they can quickly use handheld devices to locate the areas that need to be recognized. In addition, this application also significantly improves recognition efficiency in multi-target scenes and cluttered background scenes. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application;

[0059] Figure 2 This is a software structure block diagram of a terminal device according to an embodiment of this application;

[0060] Figure 3 A schematic diagram illustrating an embodiment of an image recognition method provided in this application;

[0061] Figure 4 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0062] Figure 5 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0063] Figure 6 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0064] Figure 7 This is an illustration of one scenario in an embodiment of this application;

[0065] Figure 8 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0066] Figure 9 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0067] Figure 10 This is an illustration of one scenario in an embodiment of this application;

[0068] Figure 11 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0069] Figure 12 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0070] Figure 13 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0071] Figure 14This is a schematic diagram of a terminal interface in an embodiment of this application;

[0072] Figure 15 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0073] Figure 16 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0074] Figure 17 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0075] Figure 18 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0076] Figure 19 This is a schematic diagram of an image recognition process in an embodiment of this application;

[0077] Figure 20 This is a schematic diagram of an image recognition embodiment in this application;

[0078] Figure 21 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0079] Figure 22 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0080] Figure 23 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0081] Figure 24 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0082] Figure 25 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0083] Figure 26 This is a schematic diagram of a terminal interface in an embodiment of this application;

[0084] Figure 27 This application provides a schematic diagram of the structure of an image recognition device according to an embodiment of the present application.

[0085] Figure 28 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0086] The embodiments of the present invention will now be described with reference to the accompanying drawings. The terminology used in the embodiments section is for illustrative purposes only and is not intended to limit the scope of the invention.

[0087] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0088] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0089] For ease of understanding, the structure of the terminal 100 provided in the embodiments of this application will be illustrated below. See also Figure 1 , Figure 1 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application.

[0090] like Figure 1 As shown, terminal 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0091] It is understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the terminal 100. In other embodiments of this application, the terminal 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0092] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0093] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0094] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0095] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0096] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the terminal 100.

[0097] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0098] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0099] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music output through Bluetooth headphones.

[0100] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the shooting function of the terminal 100. The processor 110 and the display screen 194 communicate via the DSI interface to enable the display function of the terminal 100.

[0101] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0102] Specifically, the video captured by the camera 193 (including image frame sequences, such as the first image, second image, third image, etc. in this application) can be transmitted to the processor 110 through, but is not limited to, the interface (such as the CSI interface or GPIO interface) described above for connecting the camera 193 and the processor 110.

[0103] The processor 110 can retrieve instructions from the memory and perform video processing (such as image recognition in this application) on the video captured by the camera 193 based on the retrieved instructions to obtain the processed image (such as the recognition result).

[0104] The processor 110 can, but is not limited to, transmit the processed image to the display screen 194 through the interface (e.g., DSI interface or GPIO interface) described above for connecting the display screen 194 and the processor 110, so that the display screen 194 can display video.

[0105] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, or USB Type-C port. USB port 130 can be used to connect a charger to charge terminal 100, and can also be used for data transfer between terminal 100 and peripheral devices. It can also be used to connect headphones for audio output. This interface can also be used to connect other electronic devices, such as AR devices.

[0106] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the terminal 100. In other embodiments of this application, the terminal 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0107] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0108] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0109] The wireless communication function of terminal 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0110] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0111] The mobile communication module 150 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the terminal 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via the antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to the modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0112] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0113] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0114] In some embodiments, antenna 1 of terminal 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling terminal 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0115] Terminal 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information. Specifically, one or more GPUs in processor 110 can perform image rendering tasks (such as the rendering tasks related to the image to be displayed in this application), and pass the rendering results to the application processor or other display driver, which triggers the display screen 194 to display video.

[0116] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, terminal 100 may include one or N display screens 194, where N is a positive integer greater than 1. The display screen 194 can display the target video in the embodiments of this application. In one implementation, terminal 100 can run a camera-related application. When the camera-related application is opened on the terminal, display screen 194 can display a shooting interface, which may include a viewfinder, within which video can be displayed.

[0117] Terminal 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0118] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0119] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, terminal 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0120] The DSP converts the digital image signal into a standard RGB, YUV format image signal to obtain the original image (e.g., the first image, second image, third image, etc. in this embodiment). The processor 110 can perform further image processing on the original image. Image processing includes, but is not limited to, image stabilization, perspective distortion correction, optical distortion correction, and cropping to fit the size of the display screen 194. The processed image can be displayed in the viewfinder of the shooting interface displayed on the display screen 194.

[0121] In this embodiment, the terminal 100 may have at least two cameras 193. For example, with two cameras, one is a front-facing camera and the other is a rear-facing camera; with three cameras, one is a front-facing camera and the other two are rear-facing cameras; with four cameras, one is a front-facing camera and the other three are rear-facing cameras. It should be noted that the camera 193 may be one or more of a wide-angle camera, a main camera, or a telephoto camera.

[0122] For example, taking two cameras, the front camera can be a wide-angle camera, and the rear camera can be the main camera. In this case, the image captured by the rear camera has a larger field of view and richer image information.

[0123] For example, taking three cameras as an example, the front camera can be a wide-angle camera, and the rear camera can be a wide-angle camera and a main camera.

[0124] For example, taking four cameras as an example, the front camera can be a wide-angle camera, and the rear camera can be a wide-angle camera, a main camera, and a telephoto camera.

[0125] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal 100 selects a frequency point, the DSP can perform Fourier transforms on the frequency energy.

[0126] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. Thus, terminal 100 can output or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0127] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in terminals, such as image recognition, facial recognition, speech recognition, and text understanding.

[0128] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0129] Internal memory 121 can be used to store computer executable program code, which includes instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound output function, image output function, etc.), etc. The data storage area may store data created during the use of terminal 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of terminal 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located in the processor.

[0130] Terminal 100 can implement audio functions such as music output and recording through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0131] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0132] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The terminal 100 can listen to music or make hands-free calls through the speaker 170A.

[0133] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal 100 receives a phone call or voice message, the receiver 170B can be brought close to the listener's ear to hear the voice.

[0134] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Terminal 100 may have at least one microphone 170C. In some embodiments, terminal 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal 100 may have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0135] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0136] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Terminal 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, terminal 100 detects the intensity of the touch operation based on pressure sensor 180A. Terminal 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example: when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0137] The shooting interface displayed on the screen 194 may include a first control and a second control. The first control is used to enable or disable the image stabilization function, and the second control is used to enable or disable the perspective distortion correction function. For example, a user can enable the image stabilization function by clicking the first control. The terminal 100 can determine the location of the first control based on the detection signal from the pressure sensor 180A, and then generate an operation command to enable the image stabilization function. Similarly, a user can enable the perspective distortion correction function by clicking the second control. The terminal 100 can determine the location of the second control based on the detection signal from the pressure sensor 180A, and then generate an operation command to enable the perspective distortion correction function.

[0138] The gyroscope sensor 180B can be used to determine the motion attitude of the terminal 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the terminal 100 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the terminal 100's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the terminal 100 through reverse movement, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.

[0139] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0140] The magnetic sensor 180D includes a Hall sensor. The terminal 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the terminal 100 is a flip phone, the terminal 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.

[0141] The 180E accelerometer can detect the magnitude of acceleration of terminal 100 in various directions (typically three axes). When terminal 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices, and is applied to applications such as screen orientation switching and pedometers.

[0142] A distance sensor 180F is used to measure distance. The terminal 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, the terminal 100 can utilize the distance sensor 180F to measure distance for rapid focusing.

[0143] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The terminal 100 emits infrared light outward through the LED. The terminal 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 100. When insufficient reflected light is detected, the terminal 100 can determine that there is no object near the terminal 100. The terminal 100 may use the proximity sensor 180G to detect when a user holds the terminal 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and screen locking.

[0144] The ambient light sensor 180L is used to sense the ambient light intensity. The terminal 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light intensity. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the terminal 100 is in a pocket to prevent accidental touches.

[0145] The fingerprint sensor 180H is used to collect fingerprints. The terminal 100 can use the characteristics of the collected fingerprints to unlock the device, access application locks, take photos with fingerprints, and answer calls with fingerprints.

[0146] Temperature sensor 180J is used to detect temperature. In some embodiments, terminal 100 uses the temperature detected by temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, terminal 100 reduces the performance of the processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is below another threshold, terminal 100 heats battery 142 to prevent abnormal shutdown of terminal 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, terminal 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.

[0147] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of terminal 100, in a different position than display screen 194.

[0148] The bone conduction sensor 180M can acquire vibration signals. In some embodiments, the bone conduction sensor 180M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 180M can also be incorporated into headphones to form bone conduction headphones. The audio module 170 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 180M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 180M to realize heart rate detection functionality.

[0149] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Terminal 100 can receive button input and generate key signal inputs related to user settings and function control of terminal 100.

[0150] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, audio output, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0151] The output device in this embodiment can be a speaker 170A, a headphone jack 170D, a motor 191, etc. Audio prompts can be achieved through the speaker 170A and the headphone jack 170D, and vibration prompts can be achieved through the motor 191.

[0152] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0153] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and separate from the terminal 100. The terminal 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The terminal 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the terminal 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 100 and cannot be separated from the terminal 100.

[0154] The software system of terminal 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to illustrate the software structure of terminal 100.

[0155] Figure 2 This is a software structure block diagram of terminal 100 according to an embodiment of this disclosure.

[0156] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0157] The application layer can include a series of application packages.

[0158] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0159] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0160] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0161] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0162] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0163] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0164] The phone manager is used to provide communication functions for terminal 100. For example, it manages call status (including connection, hang-up, etc.).

[0165] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0166] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0167] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0168] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0169] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0170] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0171] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0172] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0173] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0174] A 2D graphics engine is a graphics engine for 2D drawing.

[0175] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0176] The following example, using a photography scenario, illustrates the workflow of the terminal 100's software and hardware.

[0177] When touch sensor 180K receives a touch operation, a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, etc.). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. Taking a touch click operation as an example, where the control corresponding to the click operation is the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then calls the kernel layer to start the camera driver, capturing static images or videos through camera 193. The captured video may include the first image, second image, third image, etc., in this embodiment of the application.

[0178] In daily life, visually impaired individuals have numerous needs for recognizing textual information in near-field environments, such as recipient information on express delivery slips and names, usage instructions, and dosages on medicine leaflets. Currently, optical character recognition (OCR) and text-to-speech (TTS) technologies enable visually impaired individuals to obtain near-field textual information through terminal devices. However, when using information recognition software equipped with OCR and TTS technologies, visually impaired individuals still encounter problems such as not being able to capture the text, capturing only a small portion, or capturing it unclearly due to the lack of visual feedback.

[0179] Therefore, existing technologies are beginning to explore how to help visually impaired individuals accurately and completely read text information in the areas they want to recognize using image capture devices. One existing implementation calculates the direction and distance the user should move their phone by monitoring the integrity of the document in the current frame in real time, and then uses voice guidance to assist the user.

[0180] However, users need to move in 4 degrees of freedom (3 displacement degrees of freedom and 1 rotation degree of freedom), such as "move forward 1 foot", "move left 1 foot", "rotate to the five o'clock position". When moving, it is easy to deviate from the target and the error rate is high. For blind users, it is impossible to accurately quantify the distance they move and the angle of rotation, and they cannot make the precise actions in the guidance. Sometimes, the deviation from the target will actually increase.

[0181] To address the aforementioned technical problems, this application provides an image recognition method.

[0182] To facilitate understanding, an image recognition method provided in this application embodiment will be specifically described in conjunction with the accompanying drawings and application scenarios.

[0183] Reference Figure 3 , Figure 3This is a schematic diagram of an embodiment of an image recognition method provided in this application, as shown below. Figure 3 As shown, the image recognition method provided in this application includes:

[0184] 301. Output the first reminder; the first reminder instructs the user to establish a location association between the auxiliary part and the object to be identified, and to control the terminal to take a picture of the auxiliary part.

[0185] In step 301, the executing entity can be a terminal device.

[0186] Users (such as blind users) can open applications with image recognition capabilities on the terminal. Users can send requests for information recognition (such as image recognition) through interactive events (such as voice, touch, button, etc.), thereby triggering the recognition process.

[0187] In one possible implementation, the terminal can turn on its own camera and display the camera's captured image. The user can point the camera's field of view at the object to be identified, and the terminal can then identify the content of the captured image (or send it to other computing devices for identification, such as a server).

[0188] As mentioned above, it is difficult for users, especially the blind, to aim the camera at the object to be identified. This difficulty primarily stems from the inability to guarantee the parallelism between the device and the object, and to ensure the object falls completely within the camera's effective capture range. However, blind people possess the ability to use proprioception to guide their hand direction and quickly correct it during movement. Visually impaired individuals tend to use their hands (or hand-operated assistive tools) to locate objects. By using one hand to locate the object, the other hand can sense the distance and orientation between them, allowing for rapid adjustment of the appropriate recognition distance. Visually impaired individuals can use proprioception to capture their hands and maintain a roughly parallel alignment between the phone and the object being photographed.

[0189] In other words, visually impaired individuals can use the terminal to locate their hand (or other assistive tool) and place it roughly parallel to the target area, which is equivalent to the terminal locating the object. The terminal device can prompt the user to place their hand (or other assistive tool) on the area to be recognized, which can be a flat surface or an object, and prompt the user to use the device to photograph their hand (or other assistive tool). After the prompt, the device will determine whether the hand (or other assistive tool) has been recognized. If the hand (or other assistive tool) is recognized, the area in spatial relationship with the hand (or other assistive tool) is defined as the region of interest, i.e., the area to be recognized. If the terminal device indicates that it has not recognized the hand, it will continue to prompt the user to photograph their hand.

[0190] In this embodiment, the hand is used as an example of an assistive tool.

[0191] In one possible implementation, the terminal device may output a first reminder; the first reminder instructs the user to establish a location association between the auxiliary part and the object to be identified, and to control the terminal to take a picture of the auxiliary part.

[0192] In one possible implementation, the auxiliary part is the hand. The first prompt instructs the user to establish a positional association between the hand and the object to be identified.

[0193] In one possible implementation, the object to be identified is a planar object, and the first reminder specifically instructs the user to cover the object with their hand. For example, the object to be identified can be text on a screen, paper, or other flat surface.

[0194] In one possible implementation, the object to be identified is a three-dimensional object. The first reminder specifically instructs the user to pick up the object using the auxiliary part or to cover one face of the three-dimensional object with their hand. The three-dimensional object can be a cylinder or a polyhedron. For some non-planar items, the information is distributed across the entire cylinder or multiple faces of the polyhedron. Before outputting the first reminder, the size characteristics of the object to be identified can be used to determine whether it is necessary to pick it up. The size characteristics are mainly achieved by comparing the size of the auxiliary part with the size of the object. For example, the first reminder could be playing "Please pick up the item and turn your palm towards yourself." For items that are difficult to pick up, such as large-capacity beverages, the user is reminded to place the auxiliary part on the target face. Generally, a cylinder is a cylinder, and a polyhedron is any face.

[0195] It should be understood that when the object to be identified is a three-dimensional object, after outputting the first prompt, the terminal device can detect whether there is target information in the captured image (belonging to the captured video). The target information can be text, images, etc. Optionally, the target information can be pre-specified by the user (for example, it can be input by the user in advance on the terminal). The terminal device can prompt the user to rotate the object (cylinder) or flip its faces (polyhedron).

[0196] In one possible implementation, for a cylinder, the system prompts to stop when the target information is detected, or prompts to stop after one full rotation; for a polyhedron, the system stops when the target information is detected, or stops after all faces have been traversed.

[0197] In one possible implementation, the auxiliary part is an assistive tool for user hand operation, such as a component that may include a plate-like structure, which the user can place over the object to be identified.

[0198] This application embodiment prompts the user to establish a positional association between the assistive part and the object to be identified. Since visually impaired users can perceive the positional relationship between the assistive part and the object to be identified, as well as the positional relationship between the assistive part and the terminal device, through proprioception, the spatial alignment between the terminal and the object to be identified in three degrees of freedom can be maintained. Only the position of the terminal in the vertical direction needs to be adjusted, which reduces the user's action cost and improves the recognition efficiency.

[0199] Furthermore, by using the auxiliary parts as anchor points, the auxiliary parts are identified using computer vision, and the areas with spatial relationships with the auxiliary parts are defined as regions of interest. By utilizing the habitual interactive actions of visually impaired users in recognizing text in daily life, and through the proprioception of visually impaired users, they can quickly use handheld devices to locate the areas that need to be recognized. In addition, this application also significantly improves recognition efficiency in multi-target scenes and cluttered background scenes.

[0200] 302. If the auxiliary part exists in the captured first image, and a target object exists in the first image whose positional relationship with the auxiliary part satisfies a first preset condition, the recognition result of the target object is obtained based on the captured second image.

[0201] Wherein, the first image and the second image are images from the video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

[0202] In one possible implementation, the terminal can acquire a video stream and perform image analysis on the video stream. When it is determined that an image that meets the image recognition conditions has been acquired (for example, an image that meets the above image recognition conditions can be the first image), image recognition can be performed.

[0203] In one possible implementation, the first preset condition includes at least one of the following: there is an overlap between the target object and the auxiliary part; the target object is in the direction indicated by the auxiliary part; the target object is the object closest to the auxiliary part among the plurality of objects included in the first image.

[0204] Optional, such as Figure 4 As shown, the hand recognition model and object recognition model can be invoked to determine whether a hand or an object is detected in the image. If a hand or object is detected, it is determined whether the recognition boxes of the hand and the object overlap and exceed a threshold. If so, it is determined whether the object is in the direction of the hand being held. If so, it is determined whether the object is closest to the center point and exceeds a threshold. If so, the object / area is identified as the region of interest. If any condition is not met, the system returns to the judgment screen to continue recognizing the hand and object until the region of interest is found by moving the shooting equipment.

[0205] In one possible implementation, the video stream further includes a third image captured before the first image; if no target object satisfying the first preset condition is found in the third image, a second reminder may be output, which instructs the user to unassociate the auxiliary part with the position of the object to be identified, or to move the auxiliary part toward the edge of the object to be identified; wherein the second image is captured after the output of the second reminder.

[0206] In one possible implementation, such as Figure 5 As shown, in some scenarios, the hand may obstruct the view, preventing the identification of a suitable object after the hand is detected. If the user fails to identify a suitable object multiple times, the hand information is recorded, including the position of the recognition frame, the direction of hand movement, and the center position of the hand. Features around the hand, such as information about other objects, can also be recorded to assist in locating the target content. The user is then prompted to remove their hand or place it near the edge of an object. Based on the stored information and the current image, the system calculates whether a suitable object exists. Figure 5 If an object and a hand are detected, determine if the bounding boxes of the hand and the object overlap and exceed a threshold. If so, determine if the object is in the direction the hand is pointing. If so, determine if the object is closest to the center point and exceeds a threshold. If so, the object / region is identified as a region of interest.

[0207] Before an image that meets the requirements for image recognition can be obtained, the user can be reminded to correct the posture of the terminal based on the image captured by the camera (such as the second reminder, third reminder, fourth reminder, etc. introduced in subsequent embodiments) so that a clear image containing the complete object to be identified can be captured.

[0208] In one possible implementation, when the image of the target object in the first image is incomplete or unclear, a third reminder may be output, which instructs the user control terminal to move away from or closer to the object to be identified; the second image is acquired after the third reminder is output.

[0209] In one possible implementation, if hands and objects are not captured in the frame, it might be due to the shooting distance being too close. In this case, the shooting device cannot focus, resulting in a persistently blurry image. Figure 6 As shown, the system can determine whether the image remains blurry (current technology mainly relies on image features such as contrast and sharpness). If the image is blurry, it is determined that the user is too close and prompts the user to move the camera away.

[0210] A diagram showing the entire image of an item captured by the terminal can be shown as follows: Figure 7 As shown, there exists a recognizable space. Taking a document as an example, the object to be recognized is... Figure 7 To capture the entire document, the terminal's recognizable space should be within a cone-shaped area, with the bottom being the closest position. The resulting image is shown in the lower right corner of image 7, where the document is captured completely and occupies the entire frame. Figure 7 The shooting effect is as follows: (The image is located at the upper right front position.) Figure 7 The document can be captured in its entirety and is located in the upper left corner of the image.

[0211] As can be seen from the above embodiments, visually impaired individuals can use the camera to locate their hand and position it roughly parallel to the camera. However, after identifying the object and area of ​​interest, the area may not be complete, thus requiring guidance and correction to adjust the phone's position. In this embodiment, if the object is already within the frame, and the camera is parallel to both the hand and the object, then after the user locates the hand and identifies the object, it is only necessary to guide the user to maintain the phone's orientation and move it in the direction normal to the phone's touchscreen surface.

[0212] In one possible implementation, such as Figure 8As shown, the system checks whether the edges of the current recognition area can be identified. If so, it continues to determine whether the information in the current recognition area can be identified. If it can be identified, the text in the area is recognized and announced via voice. If not, it indicates that the object to be identified in the image is incomplete, and the user is prompted to move the camera up ("away from the object"), and it is determined whether the user has captured the entire object / area. The prompt to move the camera up / down can be conveyed through voice prompts such as "Please keep your phone in position and move it slowly up or down," or through a distinctive sound effect or vibration to indicate moving away from / approaching the object, with continuous feedback. Upon reaching the target location, the system provides confirmation of arrival through voice, sound effects, and vibration.

[0213] Feedback can also be provided to the user based on the distance between the target point and the current location. This requires estimating the target point's position on the camera and comparing it to the current position of the camera. As the user moves up or down, feedback is given based on the distance between the target point and the current location. Feedback can take the form of discrete or continuous changes, such as playing short beeps with varying frequencies based on distance; playing continuously changing sound cues with varying pitch and vibration intensity based on distance. The target point's position can be estimated by recognizing the object and determining its actual location within the image.

[0214] In one possible implementation, the terminal may output a fourth reminder based on the pose difference between when it moves away from or near the object to be identified and before moving away from or nearing the object, which is greater than a threshold. The fourth reminder instructs the user to control the terminal to adjust the pose, and the amount of the pose adjustment is related to the pose difference.

[0215] When photographing an object, there exists a spatial range defined by the relative position and angle of the camera and the object to be photographed. Within this spatial range, the information in the photograph taken by the camera can be well identified. As mentioned above, when guiding the user to move the shooting device to photograph the entire object, due to individual operating habits or the lack of a stable shooting device during movement, the terminal posture may deviate significantly from the initial terminal posture. The shooting device may no longer be able to reach the target position by moving up and down. Therefore, it is necessary to guide the user to restore the terminal posture.

[0216] In one possible implementation, such as Figure 9 First, after initiating the correction prompt, the camera's orientation is recorded as the initial orientation, and then the correction prompt begins. During the correction process, the current device orientation is acquired and the current deviation is calculated. The deviation is the difference between the camera's initial orientation and the current orientation when entering the correction process. The camera's orientation can be characterized by the motion sensor in the camera, estimated from the position of a fixed object in the captured image, or a combination of both methods.

[0217] Taking a motion sensor as an example, after recording the initial posture of the motion sensor (such as a quaternion), the difference between the posture data during the subsequent movement and adjustment process and the initial posture data is recorded as the deviation. Taking the estimation of the position of a fixed object as an example, the initial posture is calculated by the posture of the hand in the current image (such as a normal vector), and the difference between the posture data during the subsequent movement and adjustment process and the initial posture data is recorded as the deviation.

[0218] Then, it checks whether the deviation exceeds threshold 1. If the deviation exceeds threshold 1 but not threshold 2, a voice prompt will remind the user to stabilize and adjust the shooting state. If the deviation exceeds threshold 2, the user will be prompted to start again. Threshold 1 covers the range where, at that angle, the shooting image may shift, but the shift is small and will not affect the recognition of object information. Exceeding threshold 2 means that at that angle, or after correction at that angle, the information in the terminal image cannot be correctly recognized.

[0219] During the correction process, if the detected change in the terminal's posture exceeds a certain angle, the user is prompted to recalibrate. Timely reminders when the user makes an incorrect action during adjustment reduce the probability of user errors. Furthermore, it allows for timely correction and restarting when errors are significant, avoiding endless corrections.

[0220] In one possible implementation, if the auxiliary part exists in the first captured image, and a target object exists in the first image whose positional relationship with the auxiliary part meets a first preset condition, a fifth reminder is output. This fifth reminder instructs the user to remove the positional association between the auxiliary part and the object to be identified. The second image is captured after the fifth reminder is output. In other words, after capturing a first image that meets the image recognition requirements, the user can be reminded to remove the auxiliary tool from the object to be identified so that the camera can capture the complete object (the second image).

[0221] This application provides an image recognition method, the method comprising: outputting a first reminder; the first reminder instructing a user to establish a positional association between an auxiliary part and an object to be recognized, and controlling a terminal to capture the auxiliary part; if the auxiliary part exists in the captured first image, and a target object exists in the first image whose positional relationship with the auxiliary part satisfies a first preset condition, obtaining a recognition result of the target object based on a captured second image; wherein the first image and the second image are images in a video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

[0222] This application prompts the user to establish a positional association between the assistive device and the object to be identified. Since visually impaired users can perceive the positional relationship between the assistive device and the object to be identified, as well as the positional relationship between the assistive device and the terminal device, through proprioception, the spatial alignment between the terminal and the object to be identified in three degrees of freedom can be maintained. Only the position of the terminal in the vertical direction needs to be adjusted, which reduces the user's action cost and improves the recognition efficiency.

[0223] Next, using a hand as an auxiliary part and a document as the object to be recognized as an example, we will introduce the image recognition method of this application with a specific example:

[0224] In one possible implementation, user input can be acquired to initiate the recognition process. Specifically, an information recognition request can be sent via interactive events (such as voice, touch, button presses, etc.) to trigger the recognition process. The spatial relationship between the hand and an object can be used to identify the object or area of ​​interest. Specifically, the terminal device can prompt the user to place their hand on the area to be recognized, which can be a flat surface or an object, and prompt the user to use the device to capture a picture of their hand. Figure 10 As shown:

[0225] After the prompt, the device will determine whether it has recognized the hand and the document. If it has recognized the hand and the document, it will define the document, which has a spatial relationship with the hand, as the area of ​​interest, i.e. the area to be recognized. If the device prompts that it has not recognized the hand, it will continue to prompt the user to take a picture of their hand.

[0226] like Figure 11 As shown, after the system recognizes that the user has slapped their hand and identifies the document below as the region of interest, it alerts the user with a sound effect such as "ding" and prompts the user to move their hand away via voice prompts such as "Please move your hand away." In the previous step of identifying the region of interest, if the document is severely obscured by the hand, it may not be able to identify documents that have a spatial relationship with the hand. In this case, the spatial relationship of the hand is recorded, and the user is prompted to move their hand away. After the hand is moved away, the document is identified, and documents that have a spatial relationship with the hand are identified as the region of interest, and the subsequent operations described above are performed.

[0227] like Figure 12 As shown, after the system detects that the user has moved their hand away, it reminds the user that it has detected a sound effect such as "ding". It also detects that the document edge cannot be captured and reminds the user to maintain the phone's position and move it upward. During the upward movement, it will prompt the user to continue moving upward with continuous sound effects such as "ding...ding...ding...".

[0228] like Figure 13As shown, the system recognizes the edge of an object, reminds the user that it has detected a sound effect such as a "ding" or vibration, then prompts the user to steady their hands, automatically focuses and takes a picture, then recognizes the information in the document, such as through OCR, and reads the information content aloud.

[0229] Next, using a hand as an auxiliary part and a three-dimensional object (medicine bottle) as an example, we will introduce the image recognition method of this application with a specific example:

[0230] Sometimes, the information to be read can be categorized into three types based on the object: flat (e.g., documents), cylindrical (e.g., beverages), and polyhedral (e.g., medicine boxes). Different object types require different guidance strategies to obtain the necessary information. For cylindrical, polyhedral, or other irregularly shaped objects, the size and shape can affect the recognition interaction process. Whether the object needs to be picked up and placed in the hand, or whether it needs to be rotated / switched, will all affect the entire interaction process. The user sends an information recognition request through interactive events (voice, touch, button, etc.), triggering the recognition process. The device prompts the user to place their hand on the area to be recognized, which can be a flat surface or an object, and prompts the user to use the device to photograph their hand. Figure 14 As shown.

[0231] After the prompt, the device will determine whether it has recognized the hand and the object, such as the medicine bottle in the picture. If it recognizes the hand and the medicine bottle, and recognizes that the medicine bottle and the hand have a spatial relationship, then it will define this object (medicine bottle) as the area of ​​interest, that is, the area to be recognized. If the device prompts that it has not recognized, the device will continue to prompt you to take a picture of your hand.

[0232] like Figure 15 After identifying the region of interest, the shape features of the object can be identified and classified into three categories: plane, cylinder, and polyhedron.

[0233] For some non-planar objects, the information is distributed across the entire cylindrical surface or multiple faces of a polyhedron. Therefore, it is necessary to determine whether we need to pick up the object based on its size characteristics. Size characteristics are primarily determined by comparing the size of our hand with that of the object. For example... Figure 14 For items like medicine bottles that can be picked up, the system will remind the user to "pick up the item and turn your palm towards yourself." For items that cannot be picked up, such as large-capacity beverages, the system will remind the user to place their hand on the target surface. Generally, cylinders are represented by cylinders, while polyhedra are represented by any facet.

[0234] Once the user's operation is detected, the system first takes a full picture of the object according to the steps described above. Then, depending on whether there is target information on the current display surface, the system prompts the user to rotate the object (cylinder) or flip the surface (polyhedron).

[0235] For cylinders, the system should prompt to stop upon detecting the target information, or to stop after detecting a full rotation. For polyhedra, the system should stop upon detecting the target information, or after all faces have been traversed. The target information can be preset information such as shelf life or ingredients. Recognizing a cylinder's full rotation or traversing multiple faces of a polyhedron requires recording the spatial information of each face and the user's actions. The acquired information and actions can then guide the user on how to rotate the object.

[0236] After recognizing the information, the information content can be read aloud via voice.

[0237] In some scenarios, near-field information also includes touchscreens on self-service terminals, such as hospital registration machines, bank ATMs, and parcel lockers. These devices generally lack accessibility features, such as voice-over or screen reading. In these scenarios, in addition to acquiring the aforementioned information, it's necessary to guide users to click on a target button to complete the operation. For example, with parcel lockers, users need to be guided to click to select "retrieve parcel," and then scan a QR code to open the locker.

[0238] In one possible implementation, the target object is a screen, and the terminal includes a touch component; the recognition result is the text content corresponding to the target control on the screen; the text content can also be output, and the user's selection of the target control can be received; based on the relative position between the touch component and the target control, a sixth reminder is output, the sixth reminder instructing the user to control the terminal to adjust the position until the touch component touches the target control, and the adjustment amount of the position adjustment is related to the relative position.

[0239] In one possible implementation, the relative position can be determined based on the positional relationship between the image area of ​​the target control and the corresponding image area of ​​the touch component. During the video stream acquisition process, the user can be guided to move the terminal device closer to the target control. This guidance process aims to ensure that the area of ​​the target control in the image and the area corresponding to the touch component maintain a match (e.g., overlap), thereby enabling the touch component to successfully contact the target control. The area corresponding to the touch component can be an image area in the video stream display, and this area is related to the fixed position of the touch component on the terminal device.

[0240] In one possible implementation, the touch component is a bracket attached to the back of the terminal or a corner point on the terminal.

[0241] like Figure 16 The following is an example of the interaction flow in this embodiment:

[0242] It identifies information on the screen and determines the target button the user wants to click. There are two ways to lock the button:

[0243] The system pre-defines target buttons for the current scenario and automatically identifies and confirms them. For example, if the current scenario is a parcel locker, the system pre-defines the target button as "a button with parcel pickup function / that can jump to the parcel pickup page." When recognizing the parcel locker touchscreen, the system automatically identifies this target button on the screen and locks it as the target button. Button recognition is generally performed based on the text or icons on the interface using OCR or image template matching.

[0244] In some scenarios, where a large number of functions need to be provided to the user, it is necessary to recognize the information on the control screen and convert it into an accessible information menu on the mobile phone. After generation, a notification should be given to the user indicating that recognition is complete, such as a voice prompt: "Recognition complete. Swipe left or right to browse between items, double-tap to lock an item." The user can then swipe left or right on the phone to browse between menu items, hear the current button name aloud, and double-tap the screen to lock a specific button item. Figure 17 As shown, it can identify interface information and browse through all identified information, or it can directly identify buttons on the interface and browse through those buttons.

[0245] like Figure 18 The image shown is a schematic diagram of the buttons on the parcel locker's identification interface.

[0246] Once the target button is locked, the device tracks the target button on the touchscreen and guides the user to approach it, ultimately allowing the user to tap the button via a touch point on the phone. The tracking of the target button is typically achieved using computer vision technology.

[0247] In guiding a user to approach and click the target button, it's necessary to guide the user to gradually move closer to the touchscreen while simultaneously guiding them to "align" the button with it to ensure they can actually touch it. The method for guiding the user to align with the target button is to ensure that the target button remains within the "target area" both inside and outside the camera's field of view.

[0248] The preset target range is related to the size and position of the phone's camera, touch points, and target buttons. Depending on the touchscreen interface, the phone, camera position, whether it's a wide-angle lens, and the touch point position and method, the target range will vary. Therefore, the target range setting will differ for different application scenarios.

[0249] After determining the target scope, such as Figure 19 As shown, you can follow these steps to guide the user to adjust their phone's position so they can click the target button:

[0250] First, determine if the target button (or the key point of the target button) is within the target area. If it is within the target area, such as... Figure 20As shown in 'a', the user can be guided to move closer to the target until the target is touched, or until the final touch condition is met, and then the touch is triggered. Otherwise, the user is guided to adjust the phone so that the center of the target area is on the target (see Figure 'c' below).

[0251] For the defined range of target buttons, such as Figure 21 As shown, under some technical solutions, the actual button's range can be obtained; under other technical solutions, the range of icon content within the actual button, such as the text or icon on the button, can be obtained.

[0252] One possible way to determine whether a tap was successful is as follows: After a successful tap, the interface will usually switch to another page. Therefore, you can check if there is a sudden change in the screen by taking a picture of the screen with your phone. If there is a sudden change, combined with the data from the phone's motion sensor, i.e., the phone does not shake significantly (causing a drastic change in the screen), then the tap can be determined to be successful.

[0253] In some scenarios, it can be determined whether the final touch condition is met, and then the user can be guided to perform a single touch operation to complete the touch. The final touch condition here is generally the distance between the phone and the touch screen. The distance can be calculated by using a depth camera or by using computer vision to calculate the size and proportion of the image content.

[0254] By employing the above methods, the touchscreen interface of external devices is identified and transferred to the mobile phone to form an accessible menu, allowing users to find the required function entry points on the terminal. This transforms the unknown scope, unknown objects, and easily mis-touched interface into a user-friendly interface familiar to visually impaired users, conforming to their daily interaction habits and reducing learning costs. It not only allows visually impaired users to "see" the interface but also enables them to easily "find" different functions, achieving zero-contact interaction with the external device's touchscreen while accurately obtaining interface information.

[0255] In one possible implementation, the touch component is a stand attached to the back of the terminal or a corner point on the terminal. Using an external component (such as a phone stand) fixed relative to the camera or the phone's own structure (such as a corner point) in conjunction with computer vision technology, the user can interact with target buttons on the screen instead of using a finger. This fixed component or structure guides the user to the target by recognizing and clicking it. The user only needs to move the phone according to voice prompts; the operation is simple and intuitive, requiring no learning or memorization of any special operations, and no contact with the touchscreen, greatly reducing the cognitive burden on the user.

[0256] The following example uses a touch component attached to the back of the terminal as a specific illustration:

[0257] like Figure 22 A fixed folding stand is attached to the back of the phone. When the stand is extended, it extends within the camera's field of view. Generally, the extended length of the stand is greater than the protruding thickness when the phone is held in the hand, typically greater than 4.5cm. When this function is enabled on the phone, the system detects whether the stand is present in the frame or whether it is in the target position to determine if the stand is extended. If not extended, the user is prompted to extend the stand. Once the system detects that the user has extended the stand, the guided process described in the above embodiment is executed.

[0258] After identifying the target button, guide the user to click it. The preset target area is related to the phone's camera, the touch point, and the size and position of the target button. Depending on the touchscreen interface, the phone, the camera position, whether it's a wide-angle lens, and the touch point position and method, the target area will vary. Therefore, the target area setting will differ for different application scenarios.

[0259] In this embodiment, the phone holder is positioned less than 2cm directly below the camera, in an application scenario of a parcel locker, where the minimum effective touch area of ​​the button is approximately 2*5cm. The actual size of the wide-angle lens's shooting plane within the distance of the phone holder is approximately 5*10cm. In this embodiment, as... Figure 23 As shown, the target area can be defined as the rectangular area formed by the top edge of the phone holder, the center line of the screen, the left edge of the screen, and the right edge of the screen. The target area may differ depending on other basic conditions.

[0260] The following example uses a corner point on the terminal as an example:

[0261] like Figure 24 As shown, touch can be performed using the corner of the phone. The user is guided to hold the phone close to the target button until the conditions for a final touch are met, and then guided to tap the screen with the corner of the phone to click the target button.

[0262] In this embodiment, the key point of the target button is defined as the lower left corner of the target button, such as... Figure 25 The illustration is on the left. In this embodiment, the upper right corner of the phone is located at the upper right corner of the camera lens, 3cm horizontally and 2.5cm vertically. The application scenario is a parcel locker, and the minimum effective contact area of ​​the button is approximately 2*5cm. Based on the experimental results of the corner contact distribution for this phone, the following is obtained: Figure 25 The diagram on the right shows the target area, indicated by the darker color. If the user holds the phone 5-6 centimeters from the screen and the key point of the button falls within this area, they can tap the screen with the upper right corner of the phone to click the target button. In this embodiment, the target area may differ under other conditions.

[0263] During the guidance process, to avoid repeated corrections during approach, the target area is expanded four times before reaching the touch distance, as follows: Figure 26 Steps 1-4: When the phone moves to the target distance (approximately 6cm), shrink the target area to the standard size and guide the user to readjust the position of the target button's key point within the camera frame. After adjustment, prompt the user to tap the screen using the upper right corner of the phone.

[0264] This application also provides an image recognition device, which can be a terminal device, see reference. Figure 27 , Figure 27 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application, such as... Figure 27 As shown in the figure, the image recognition device 2700 includes:

[0265] Output module 2701 is used to output a first reminder; the first reminder instructs the user to establish a location association between the auxiliary part and the object to be identified, and to control the terminal to take a picture of the auxiliary part;

[0266] The specific details of the output module 2701 can be found in the description of step 301 in the above embodiments, and will not be repeated here.

[0267] The recognition module 2702 is used to obtain the recognition result of the target object based on the acquired second image when the auxiliary part exists in the captured first image and a target object exists in the first image whose positional relationship with the auxiliary part meets a first preset condition.

[0268] Wherein, the first image and the second image are images from the video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

[0269] The specific details of the identification module 2702 can be found in the description of step 302 in the above embodiments, and will not be repeated here.

[0270] In one possible implementation, the auxiliary part is the hand.

[0271] In one possible implementation, the first preset condition includes at least one of the following:

[0272] There is an overlap between the target object and the auxiliary part;

[0273] The target object is in the direction indicated by the auxiliary part;

[0274] The target object is the object closest to the auxiliary part among the multiple objects included in the first image.

[0275] In one possible implementation, the video stream further includes a third image captured before the first image; the output module is further configured to:

[0276] When there is no target object in the third image that meets the first preset condition, a second reminder is output. The second reminder instructs the user to unlink the auxiliary part from the position of the object to be identified, or to move the auxiliary part toward the edge of the object to be identified.

[0277] The second image was acquired after the second reminder was output.

[0278] In one possible implementation, the output module is further configured to:

[0279] When the image of the target object in the first image is incomplete or unclear, a third reminder is output, which instructs the user control terminal to move away from or closer to the object to be identified.

[0280] The second image was acquired after the third reminder was output.

[0281] In one possible implementation, the output module is further configured to:

[0282] Based on the fact that when the terminal moves away from or nears the object to be identified, the difference in posture compared to before moving away from or nearing the object to be identified is greater than a threshold, a fourth reminder is output according to the posture difference. The fourth reminder instructs the user to control the terminal to adjust the posture, and the adjustment amount of the posture adjustment is related to the posture difference.

[0283] In one possible implementation,

[0284] The object to be identified is a planar object, and the first reminder specifically instructs the user to cover the object to be identified with the auxiliary part; or...

[0285] The object to be identified is a three-dimensional object, and the first reminder specifically instructs the user to pick up the object to be identified through the auxiliary part or to cover one surface of the three-dimensional object with the auxiliary part.

[0286] In one possible implementation, the output module is further configured to:

[0287] If the auxiliary part exists in the first image and there is a target object in the first image whose positional relationship with the auxiliary part meets the first preset condition, a fifth reminder is output, which instructs the user to release the positional association between the auxiliary part and the object to be identified.

[0288] The second image was acquired after the fifth reminder in the output.

[0289] In one possible implementation, the target object is a screen, and the terminal includes a touch component; the recognition result is the text content corresponding to the target control on the screen; the output module is further configured to:

[0290] Output the text content;

[0291] The device further includes: a receiving module, configured to receive a user's selection of the target control;

[0292] The output module is also used for:

[0293] Based on the relative position between the touch component and the target control, a sixth reminder is output, instructing the user control terminal to adjust the position until the touch component contacts the target control, and the adjustment amount is related to the relative position; or,

[0294] According to the target control.

[0295] In one possible implementation, the touch component is a bracket attached to the back of the terminal or a corner point on the terminal.

[0296] The following describes a terminal device provided in an embodiment of this application. The terminal device can be... Figure 27 For image recognition devices, please refer to Figure 28 , Figure 28 This is a schematic diagram of a terminal device provided in an embodiment of this application. The terminal device 2800 can specifically be a virtual reality (VR) device, a mobile phone, a tablet, a laptop computer, a smart wearable device, etc., and is not limited thereto. Specifically, the terminal device 2800 includes: a receiver 2801, a transmitter 2802, a processor 2803, and a memory 2804 (wherein the terminal device 2800 may have one or more processors 2803). Figure 28 (Taking a processor as an example), the processor 2803 may include an application processor 28031 and a communication processor 28032. In some embodiments of this application, the receiver 2801, transmitter 2802, processor 2803, and memory 2804 may be connected via a bus or other means.

[0297] Memory 2804 may include read-only memory and random access memory, and provides instructions and data to processor 2803. A portion of memory 2804 may also include non-volatile random access memory (NVRAM). Memory 2804 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0298] The processor 2803 controls the operation of the terminal device. In specific applications, the various components of the terminal device are coupled together through a bus system. This bus system includes not only the data bus but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.

[0299] The methods disclosed in the embodiments of this application can be applied to or implemented by processor 2803. Processor 2803 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 2803 or by instructions in software form. Processor 2803 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 2803 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 2804. Processor 2803 reads the information in memory 2804 and, in conjunction with its hardware, completes the steps of the above method. Specifically, processor 2803 can read the information in memory 2804 and, in conjunction with its hardware, complete the data processing-related steps 301 to 302 in the above embodiments.

[0300] Receiver 2801 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the terminal device. Transmitter 2802 can be used to output digital or character information through the first interface; transmitter 2802 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 2802 may also include a display device such as a display screen.

[0301] This application also provides a computer program product that, when run on a computer, causes the computer to perform the functions described in the above embodiments. Figure 3 The steps of the image recognition method described in the corresponding embodiments.

[0302] This application also provides a computer-readable storage medium storing a program for signal processing, which, when run on a computer, causes the computer to perform the steps of the image recognition method as described in the foregoing embodiments.

[0303] The image recognition device provided in this application embodiment can specifically be a chip, which includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the data processing method described in the above embodiments, or to cause the chip in the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit can be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit can be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).

[0304] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0305] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0306] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0307] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. An image recognition method, characterized in that, The method includes: Output a first reminder; the first reminder instructs the user to establish a positional association between the auxiliary part and the object to be identified, and to control the terminal to take a picture of the auxiliary part; if the object to be identified is a planar object, the first reminder specifically instructs the user to cover the object to be identified with the auxiliary part; or if the object to be identified is a three-dimensional object, the first reminder specifically instructs the user to pick up the object to be identified through the auxiliary part or to cover one surface of the three-dimensional object with the auxiliary part; If the auxiliary part exists in the first image and a target object exists in the first image whose positional relationship with the auxiliary part meets the first preset condition, the recognition result of the target object is obtained based on the acquired second image. Wherein, the first image and the second image are images from the video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

2. The method according to claim 1, characterized in that, The auxiliary part is the hand.

3. The method according to claim 1 or 2, characterized in that, The first preset condition includes at least one of the following: There is an overlap between the target object and the auxiliary part; The target object is in the direction indicated by the auxiliary part; The target object is the object closest to the auxiliary part among the multiple objects included in the first image.

4. The method according to any one of claims 1 to 2, characterized in that, The video stream also includes a third image captured before the first image; the method further includes: When there is no target object in the third image that meets the first preset condition, a second reminder is output. The second reminder instructs the user to unlink the auxiliary part from the position of the object to be identified, or to move the auxiliary part toward the edge of the object to be identified. The second image was acquired after the second reminder was output.

5. The method according to any one of claims 1 to 2, characterized in that, The method further includes: When the image of the target object in the first image is incomplete or unclear, a third reminder is output, which instructs the user control terminal to move away from or closer to the object to be identified. The second image was acquired after the third reminder was output.

6. The method according to claim 5, characterized in that, The method further includes: Based on the fact that when the terminal moves away from or nears the object to be identified, the difference in posture compared to before moving away from or nearing the object to be identified is greater than a threshold, a fourth reminder is output according to the posture difference. The fourth reminder instructs the user to control the terminal to adjust the posture, and the adjustment amount of the posture adjustment is related to the posture difference.

7. The method according to any one of claims 1 to 2, characterized in that, The method further includes: If the auxiliary part exists in the first image and there is a target object in the first image whose positional relationship with the auxiliary part meets the first preset condition, a fifth reminder is output, which instructs the user to release the positional association between the auxiliary part and the object to be identified. The second image was acquired after the fifth reminder in the output.

8. The method according to any one of claims 1 to 2, characterized in that, The target object is a screen, and the terminal includes a touch component; the recognition result is the text content corresponding to the target control on the screen; the method further includes: Output the text content and receive the user's selection for the target control; Based on the relative position between the touch component and the target control, a sixth reminder is output, which instructs the user control terminal to adjust the position until the touch component touches the target control, and the adjustment amount is related to the relative position.

9. The method according to claim 8, characterized in that, The touch component is a bracket attached to the back of the terminal or a corner point on the terminal.

10. An image recognition device, characterized in that, The device includes: The output module is used to output a first reminder; the first reminder instructs the user to establish a positional association between the auxiliary part and the object to be identified, and to control the terminal to take a picture of the auxiliary part; the object to be identified is a planar object, and the first reminder specifically instructs the user to cover the object to be identified with the auxiliary part; or, the object to be identified is a three-dimensional object, and the first reminder specifically instructs the user to pick up the object to be identified through the auxiliary part or to cover one surface of the three-dimensional object with the auxiliary part; The recognition module is used to obtain the recognition result of the target object based on the acquired second image when the auxiliary part exists in the captured first image and a target object exists in the first image whose positional relationship with the auxiliary part meets a first preset condition. Wherein, the first image and the second image are images from the video stream captured by the user-controlled terminal after the first reminder is output, and the second image is captured after the first image.

11. The apparatus according to claim 10, characterized in that, The auxiliary part is the hand.

12. The apparatus according to claim 10 or 11, characterized in that, The first preset condition includes at least one of the following: There is an overlap between the target object and the auxiliary part; The target object is in the direction indicated by the auxiliary part; The target object is the object closest to the auxiliary part among the multiple objects included in the first image.

13. The apparatus according to any one of claims 10 to 11, characterized in that, The video stream also includes a third image captured before the first image; the output module is further configured to: When there is no target object in the third image that meets the first preset condition, a second reminder is output. The second reminder instructs the user to unlink the auxiliary part from the position of the object to be identified, or to move the auxiliary part toward the edge of the object to be identified. The second image was acquired after the second reminder was output.

14. The apparatus according to any one of claims 10 to 11, characterized in that, The output module is also used for: When the image of the target object in the first image is incomplete or unclear, a third reminder is output, which instructs the user control terminal to move away from or closer to the object to be identified. The second image was acquired after the third reminder was output.

15. The apparatus according to claim 14, characterized in that, The output module is also used for: Based on the fact that when the terminal moves away from or nears the object to be identified, the difference in posture compared to before moving away from or nearing the object to be identified is greater than a threshold, a fourth reminder is output according to the posture difference. The fourth reminder instructs the user to control the terminal to adjust the posture, and the adjustment amount of the posture adjustment is related to the posture difference.

16. The apparatus according to any one of claims 10 to 11, characterized in that, The output module is also used for: If the auxiliary part exists in the first image and there is a target object in the first image whose positional relationship with the auxiliary part meets the first preset condition, a fifth reminder is output, which instructs the user to release the positional association between the auxiliary part and the object to be identified. The second image was acquired after the fifth reminder in the output.

17. The apparatus according to any one of claims 10 to 11, characterized in that, The target object is a screen, and the terminal includes a touch component; the recognition result is the text content corresponding to the target control on the screen; the output module is further used for: Output the text content; The device further includes: a receiving module, configured to receive a user's selection of the target control; The output module is also used for: Based on the relative position between the touch component and the target control, a sixth reminder is output, instructing the user control terminal to adjust the position until the touch component contacts the target control, and the adjustment amount is related to the relative position; or, According to the target control.

18. The apparatus according to claim 17, characterized in that, The touch component is a bracket attached to the back of the terminal or a corner point on the terminal.

19. An image recognition device, characterized in that, The device includes a processor, memory, camera, output device, and bus, wherein: The processor, the memory, and the camera are connected via the bus; The camera is used to capture video in real time; The memory is used to store computer programs or instructions; The processor is used to call or execute programs or instructions stored in the memory, and is also used to call the camera and the output device to implement the method steps of any one of claims 1-9.

20. A computer-readable storage medium comprising a program, which, when run on a computer, causes the computer to perform the method as claimed in any one of claims 1 to 9.

21. A computer program product containing instructions, characterized in that, When the computer program product is run on a terminal, the terminal causes the terminal to perform the method described in any one of claims 1-9.