A scene recognition method and device

CN122841932APending Publication Date: 2026-09-29HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510398101.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,由于目标检测模型输出的图像检测结果数据量过于庞大,这使得NPU与GPU之间的数据传输面临巨大压力,导致GPU对图像检测结果的处理速度大幅降低,影响了终端设备的场景识别效率

Benefits of technology

[0021]第五方面,本申请实施例提供了一种计算机程序产品,包括计算机执行指令,当计算机执行指令在计算机上运行时,使得计算机执行第一方面提供的任意一种方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841932A_ABST
    Figure CN122841932A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a scene recognition method and device, which are applied to the technical field of image processing. The method comprises: detecting a to-be-detected image to obtain a plurality of first bounding boxes corresponding to a target in the to-be-detected image and candidate class information, the candidate class information indicating a plurality of candidate classes to which the target belongs; determining an actual class of the target according to the candidate class information; and transmitting an image detection result to a second processor, the image detection result comprising coordinates of the plurality of first bounding boxes and the actual class of the target. The scheme can reduce the amount of data that needs to be transmitted from the first processor to the second processor, and also reduces the amount of data that needs to be processed by the second processor, thereby ultimately improving the scene recognition efficiency of the terminal device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a scene recognition method and device. Background Technology

[0002] In today's digital age, terminal devices have become deeply integrated into all aspects of people's lives, from daily communication and entertainment to work and study, with a wide variety of usage scenarios. To meet the diverse needs of users in different scenarios and improve user experience, it is crucial to optimize terminals for different usage scenarios.

[0003] In determining the usage scenario of a terminal device, the current common approach is to leverage the collaborative work of the device's network processing unit (NPU) and graphics processing unit (GPU). Specifically, the NPU uses an object detection model to detect the current display screen, generating image detection results. The GPU then uses these results to determine the current usage scenario. However, the sheer volume of image detection data output by the object detection model puts immense pressure on data transmission between the NPU and GPU, significantly reducing the GPU's processing speed and impacting the device's scene recognition efficiency.

[0004] Therefore, there is an urgent need for a fast and efficient scene recognition solution to meet user needs. Summary of the Invention

[0005] This application provides a scene recognition method and device. The method can detect an image to be detected, obtain candidate category information corresponding to the target, determine the actual category of the target based on the candidate category information, and then transmit the actual category of the target to a second processor. The first processor does not need to transmit the large amount of candidate category information to the second processor, which reduces the amount of image detection results transmitted to the second processor and also reduces the amount of data that the second processor needs to process, ultimately improving the scene recognition efficiency of the terminal device.

[0006] In a first aspect, embodiments of this application provide a scene recognition method applied to a first processor of a terminal device. The terminal device includes a first processor and a second processor. The method includes: firstly detecting an image to be detected to obtain multiple first bounding boxes and candidate category information corresponding to a target in the image to be detected, wherein the candidate category information indicates multiple candidate categories to which the target belongs; then determining the actual category of the target based on the candidate category information; and then transmitting an image detection result to the second processor, wherein the image detection result includes the coordinates of the multiple first bounding boxes and the actual category of the target.

[0007] Based on the above technical solution, the first processor can determine the actual category of the target in the image to be detected based on the candidate category information, and then transmit the actual category of the target to the second processor. The first processor does not need to transmit multiple candidate categories to the target to the second processor, but only needs to transmit the actual category of the target to the second processor. This can reduce the amount of image detection results transmitted to the second processor, thereby reducing the data transmission pressure between the first and second processors inside the terminal device, and also reducing the amount of data that the second processor needs to process, ultimately improving the scene recognition efficiency of the terminal device.

[0008] In one possible implementation, determining the actual category of the target based on the candidate category information includes: determining the credibility of each category among multiple candidate categories to which the target belongs based on the candidate category information; and determining the category with the highest credibility as the actual category of the target based on the credibility of each category among multiple candidate categories.

[0009] One possible implementation provides a scheme for determining the actual category of a target based on the credibility of each category, which improves the feasibility of the embodiments of this application.

[0010] In one possible implementation, determining the actual category of the target based on candidate category information includes: determining multiple candidate categories that are related among the multiple candidate categories to which the target belongs; determining the credibility of each of the multiple related candidate categories based on the candidate category information; and determining the category with the highest credibility as the actual category of the target based on the credibility of each of the multiple related candidate categories.

[0011] In this possible implementation, the actual category of the target is determined based on the credibility of each of the multiple related categories. Since the related categories usually belong to the same scenario, the actual category of the target is less likely to be a false detection, thus improving the accuracy of the actual category of the target.

[0012] In one possible implementation, before transmitting the image detection results to the second processor, the method further includes: determining multiple second bounding boxes within multiple first bounding boxes, wherein the scale of the feature maps corresponding to the second bounding boxes is a preset scale; determining the actual category of the target corresponding to the multiple second bounding boxes based on candidate category information; correspondingly, transmitting the image detection results to the second processor includes: transmitting the image detection results to the second processor, wherein the image detection results include the coordinates of the multiple second bounding boxes and the actual category of the target.

[0013] In this possible implementation, the first processor can determine multiple second bounding boxes within multiple first bounding boxes, and then transmit these multiple second bounding boxes to the second processor. The first processor does not need to transmit the coordinates of multiple first bounding boxes with large data volume to the second processor; it only needs to transmit the coordinates of multiple second bounding boxes with smaller data volume to the second processor. This can reduce the data transmission pressure inside the terminal device and also reduce the amount of data that the second processor needs to process, ultimately improving the scene recognition efficiency of the terminal device.

[0014] In one possible implementation, the method further includes: determining the current application of the terminal device; and determining a preset scale of the feature map based on the current application of the terminal device.

[0015] In one possible implementation, the method further includes: in a click scenario, determining at least one third bounding box based on image detection results, wherein the actual category of the target corresponding to the third bounding box is the click target; determining the center coordinates of the at least one third bounding box; and transmitting the center coordinates of the at least one third bounding box to a second processor.

[0016] In this possible implementation, in a click scenario, the first processor of the terminal device can directly transmit the center coordinates of the third bounding box, without needing to transmit the two coordinates of the third bounding box, further reducing the amount of data that needs to be transmitted and lowering the data transmission pressure inside the terminal device.

[0017] In one possible implementation, the first processor is a neural network processor (NPU), and the second processor is a graphics processing unit (GPU).

[0018] Secondly, embodiments of this application provide a terminal device, including a processor and a memory. The processor is coupled to the memory; the memory stores computer instructions, which are loaded and executed by the processor to enable the terminal device to implement any of the methods provided in the first aspect.

[0019] Thirdly, embodiments of this application provide a chip, which includes: a processor and an interface circuit; the interface circuit is used to receive code instructions and transmit them to the processor; the processor is used to run the code instructions to execute any of the methods provided in the first aspect.

[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing at least one computer program instruction, which is loaded and executed by a processor to implement any of the methods provided in the first aspect above.

[0021] Fifthly, embodiments of this application provide a computer program product, including computer execution instructions, which, when executed on a computer, cause the computer to perform any of the methods provided in the first aspect.

[0022] The possible implementations of aspects two through five have similar effects to those of aspect one and the possible designs of aspect one, and will not be elaborated upon here. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the hardware system structure of the terminal device provided in the embodiments of this application;

[0024] Figure 2 This is a schematic diagram of the software system structure of the terminal device provided in the embodiments of this application;

[0025] Figure 3 A flowchart illustrating a scene recognition method provided in an embodiment of this application;

[0026] Figure 4 A scene diagram illustrating a scene recognition method provided in an embodiment of this application;

[0027] Figure 5a A scene diagram illustrating another scene recognition method provided in this application embodiment;

[0028] Figure 5b A scene diagram illustrating another scene recognition method provided in this application embodiment;

[0029] Figure 5c A scene diagram illustrating another scene recognition method provided in this application embodiment;

[0030] Figure 6 A flowchart illustrating another scene recognition method provided in an embodiment of this application;

[0031] Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0032] Figure 8 This is a schematic diagram of another terminal device provided in an embodiment of this application. Detailed Implementation

[0033] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0034] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is merely a description of the relationship between the related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. A and B can be singular or plural.

[0035] In the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0036] Furthermore, to facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0037] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.

[0038] It is understood that the term "embodiment" used throughout the specification means that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, various embodiments throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0039] It is understood that some optional features in the embodiments of this application can be implemented independently in certain scenarios without relying on other features, such as the current solution on which they are based, to solve the corresponding technical problems and achieve the corresponding effects. Alternatively, they can be combined with other features as needed in certain scenarios. Accordingly, the apparatus given in the embodiments of this application can also implement these features or functions, which will not be elaborated here.

[0040] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referred to each other. Unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments are consistent and can be mutually referenced. Different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following embodiments of this application do not constitute a limitation on the scope of protection of this application.

[0041] In today's digital age, terminal devices have become deeply integrated into all aspects of people's lives, from daily communication and entertainment to work and study, with a wide variety of usage scenarios. To meet the diverse needs of users in different scenarios and improve user experience, it is crucial to optimize terminals for different usage scenarios.

[0042] In determining the usage scenario of a terminal device, the current common approach is to leverage the collaborative work of the device's network processing unit (NPU) and graphics processing unit (GPU). Specifically, the NPU uses an object detection model to detect the current display screen, generating image detection results. The GPU then uses these results to determine the current usage scenario. However, the sheer volume of image detection data output by the object detection model puts immense pressure on data transmission between the NPU and GPU, significantly reducing the GPU's processing speed and impacting the device's scene recognition efficiency.

[0043] This application provides a scene recognition method applied to a first processor of a terminal device. The terminal device includes a first processor and a second processor. The method includes: detecting an image to be detected to obtain multiple first bounding boxes and candidate category information corresponding to a target in the image to be detected, wherein the candidate category information indicates multiple candidate categories to which the target belongs; determining the actual category of the target based on the candidate category information; and transmitting an image detection result to the second processor, wherein the image detection result includes the coordinates of the multiple first bounding boxes and the actual category of the target.

[0044] Based on the above technical solution, the first processor can determine the actual category of the target in the image to be detected based on the candidate category information, and then transmit the actual category of the target to the second processor. The first processor does not need to transmit multiple candidate categories to the target to the second processor, but only needs to transmit the actual category of the target to the second processor. This can reduce the amount of image detection results transmitted to the second processor, thereby reducing the data transmission pressure between the first and second processors inside the terminal device, and also reducing the amount of data that the second processor needs to process, ultimately improving the scene recognition efficiency of the terminal device.

[0045] The scene recognition method provided in this application can be applied to terminal devices. The terminal device (UE) in this application can also be referred to as user equipment, terminal, mobile station (MS), mobile terminal (MT), etc. The terminal can be a mobile phone, tablet computer, laptop computer, PDA, mobile internet device (MID), wearable device, virtual reality (VR) device, augmented reality (AR) device, wireless terminal in industrial control, wireless terminal in self-driving, wireless terminal in remote medical surgery, wireless terminal in smart grid, wireless terminal in transportation safety, wireless terminal in smart city, or wireless terminal in smart home, etc. This application does not limit the specific technology or device form used in the terminal.

[0046] The technical solutions provided in this application can be applied to various communication systems, such as: Long Term Evolution (LTE) systems, LTE Frequency Division Duplex (FDD) systems, LTE Time Division Duplex (TDD) systems, Universal Mobile Telecommunication System (UMTS), Worldwide Interoperability for Microwave Access (WiMAX) systems, 5th Generation (5G) mobile communication systems, and New Radio (NR). The 5G mobile communication systems in this application include non-standalone (NSA) 5G mobile communication systems and standalone (SA) 5G mobile communication systems.

[0047] The technical solutions provided in this application can also be applied to future communication systems, such as the sixth-generation mobile communication system, and this application does not limit them.

[0048] To better understand the embodiments of this application, the structure of the terminal device of the embodiments of this application is described below.

[0049] Figure 1 A schematic diagram of the terminal device 100 is shown. The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a multispectral sensor 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0050] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0051] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors.

[0052] In this embodiment, the Neural Processing Unit (NPU) can acquire an image to be detected from the Graphics Processing Unit (GPU). After acquiring the image, the NPU performs detection on it to obtain candidate category information corresponding to the target in the image. This candidate category information indicates multiple candidate categories to which the target belongs. Based on the candidate category information, the actual category of the target is determined. Then, the NPU transmits the image detection result to the GPU, which includes the actual category of the target. The GPU determines the current scene of the terminal device based on the image detection result.

[0053] It is understandable that the GPU also has the ability to detect the image to be detected, and the NPU also has the ability to determine the current scene of the terminal device based on the detection results of the image. Therefore, the work of the NPU and the GPU can be interchanged in the above method.

[0054] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0055] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0056] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal device 100. While charging the battery 142, the charging management module 140 can also supply power to the terminal device via the power management module 141.

[0057] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0058] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0059] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 150 can provide solutions for wireless communication applications, including 2G / 3G / 4G / 5G, on the terminal device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.

[0060] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0061] In some embodiments, the antenna 1 of the terminal device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the terminal device 100 can communicate with the network and other devices through wireless communication technology.

[0062] Terminal device 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0063] The display screen 194 is used to display images, display videos, and receive swipe operations, etc. The display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include one or more display screens 194.

[0064] Terminal device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0065] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's image sensor. The light signal is converted into an electrical signal, and the camera's image sensor transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0066] Camera 193 is used to capture still images or videos. An object is projected onto an image sensor through the lens, generating an optical image. The image sensor can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The image sensor converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device 100 may include one or more cameras 193.

[0067] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0068] Internal memory 121 can be used to store executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the terminal device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the terminal device 100 by running instructions stored in internal memory 121 and / or instructions stored in memory located within the processor.

[0069] Terminal device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0070] In some embodiments, the multispectral sensor 180 can be used to acquire first reflection spectral data and first light source spectral data corresponding to multiple light sources.

[0071] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Terminal device 100 can receive button input and generate key signal inputs related to user settings and function control of terminal device 100.

[0072] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. Indicator 192 can be an indicator light, used to indicate charging status, battery level changes, or to indicate messages, missed calls, notifications, etc.

[0073] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the terminal device 100.

[0074] The software system of terminal device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture, etc. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of terminal device 100.

[0075] Figure 2 This is a software structure block diagram of the terminal device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, the hardware abstraction layer, and the kernel layer.

[0076] The application layer can include a series of application packages. For example... Figure 2 As shown, the application package can include applications such as camera, settings, and calendar.

[0077] The camera app is an application with shooting and video recording functions. The terminal device can respond to the user's action of opening the camera app to take photos or record videos. It is understandable that the photo and video recording functions of the camera app can also be invoked by other applications.

[0078] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes a set of predefined functions.

[0079] Among them, such as Figure 2 As shown, the application framework layer can also include a camera service, which can be called by camera applications to enable functions such as taking photos or recording videos.

[0080] In addition, such as Figure 2 As shown, the application framework layer may also include a window manager, content provider, resource manager, and view system, etc.

[0081] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0082] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.

[0083] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0084] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0085] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.

[0086] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0087] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0088] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0089] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0090] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG2, H.262, MP3, AAC, AMR, JPG, and PNG.

[0091] 3D graphics processing libraries are used to implement 3D graphics drawing, image rendering, compositing, and layer processing. 2D graphics engines are drawing engines for 2D graphics.

[0092] The Hardware Abstraction Layer (HAL) is a layer of abstraction that sits between the kernel layer and the Android runtime. The HAL can be a wrapper around the kernel's hardware drivers, providing a calling interface for the application framework layer.

[0093] In this embodiment, the hardware abstraction layer may include an image detection module and a scene recognition module. The image detection module can detect the image to be detected, obtain candidate category information corresponding to the target in the image, which indicates multiple candidate categories to which the target belongs; and determine the actual category of the target based on the candidate category information. Then, the NPU transmits the image detection result to the GPU, which includes the actual category of the target. The scene recognition module is used to determine the current scene of the terminal device based on the image detection result.

[0094] The kernel layer is the layer between hardware and software. The kernel layer includes at least camera drivers, sensor drivers, and display drivers. In some embodiments, the camera driver controls the operation of the camera, the sensor driver controls the operation of the multispectral sensor, and the display driver controls the display screen to show images.

[0095] The hardware can be a camera, a multispectral sensor, and a display screen, etc. In the embodiments of this application, the camera can be a front-facing camera or a rear-facing camera.

[0096] It should be noted that although the embodiments of this application are described using the Android system, the principle of the image processing method is also applicable to terminal devices with operating systems such as iOS or Windows.

[0097] It is understood that in the embodiments of this application, the executing entity may perform some or all of the steps in the embodiments of this application. These steps or operations are merely examples, and the embodiments of this application may also perform other operations or variations thereof. Furthermore, the various steps may be executed in different orders as presented in the embodiments of this application, and it is not necessarily necessary to execute all the operations in the embodiments of this application.

[0098] It should be noted that the message names between devices or the names of parameters in the messages in the embodiments of this application are just examples. In specific implementations, other names may also be used. This application does not specifically limit this.

[0099] The following is combined Figures 3 to 6 The technical solutions of this application will be described in detail with specific method embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0100] For example, Figure 3 This is a flowchart illustrating a scene recognition method provided in an embodiment of this application. (Refer to...) Figure 3 As shown, the scene recognition method may specifically include the following steps:

[0101] 301. The first processor performs detection on the image to be detected, and obtains multiple first bounding boxes and candidate category information.

[0102] The first processor of the terminal device detects the image to be detected using a target detection model, and obtains multiple first bounding boxes and candidate category information corresponding to the target in the image to be detected; the candidate category information indicates the multiple candidate categories to which the target belongs.

[0103] In this embodiment, when detecting an image, the object detection model first extracts image features through a backbone network to obtain multiple feature maps. In this embodiment, the backbone network of the object detection model typically consists of a series of convolutional layers, pooling layers, and activation functions, and its main function is to extract features from the original image. Convolutional layers slide different convolutional kernels across the image to extract local features such as edges and textures. Pooling layers are used to reduce the size of the feature maps, reducing computation while retaining important feature information. After multi-layer processing by the backbone network, feature maps with different scales and semantic information can be obtained.

[0104] The object detection model then uses a feature pyramid network (FPN) or similar structure to generate multi-scale feature maps. For example, for a 640*640 pixel image to be detected, the object detection model can determine feature maps of 80*80 pixels, 40*40 pixels, and 20*20 pixels. Feature maps of different scales contain different levels of semantic and spatial information. Large-scale feature maps (such as 20×20) retain more detailed information and are suitable for detecting small objects; small-scale feature maps (such as 80×80) have richer semantic information and are suitable for detecting large objects.

[0105] After generating feature maps, the object detection model can divide each feature map into a grid. For example, an 80*80 feature map can be divided into 80*80 grids, and the intersection or center point of each grid is a grid point. The object detection model assumes that if the center point of an object falls within a certain grid, then that grid is responsible for detecting that object. For each grid point, the model predicts multiple bounding boxes (i.e., the multiple first bounding boxes mentioned above), and predicts the corresponding class probability, confidence score, and bounding box offset for each bounding box. In this embodiment, the bounding box is also referred to as a bounding box.

[0106] In this embodiment of the application, an image to be detected may include multiple targets. For a target, the target detection model can determine at least one bounding box. It can be understood that each bounding box is the bounding box of a target detected by the target detection model, and therefore each bounding box has corresponding candidate category information, that is, multiple candidate categories to which the target within the bounding box belongs.

[0107] In this embodiment, the multiple categories in the candidate category information can be preset in the target detection model. For a target, the target detection model can determine the confidence level of each of these preset categories. Alternatively, these multiple categories can also be determined in other ways. For example, the application currently running on the terminal device can be obtained first, and then multiple candidate categories can be determined based on the running application. These candidate categories are the possible categories of targets displayed in the application's screen. Therefore, when it is determined that the terminal device is running the application, these candidate categories can be determined.

[0108] For example Figure 4 As shown, when the object detection model detects an image, it can obtain multiple bounding boxes (i.e., multiple first bounding boxes). Figure 4 The multiple bounding boxes include targets 1 (door), 2 (window), 3 (chair), 4 (computer), and 5 (table). At the same time, candidate category information corresponding to these multiple bounding boxes is obtained. The candidate category information indicates the multiple candidate categories to which the targets within the multiple bounding boxes belong.

[0109] In this embodiment, the candidate category information includes the credibility of the target with each category. For example, for target 4 (computer), the target detection model can determine the credibility of target 4 with multiple categories. For instance, it can include a credibility of 0.9 for the target belonging to the "computer" category, a credibility of 0.7 for the target belonging to the "television" category, and a credibility of 0.4 for the target belonging to the "television" category. In addition, it can also include the credibility of the target with other categories, which will not be elaborated here.

[0110] It is understood that in the embodiments of this application, the candidate category information describes multiple candidate categories to which the target in the image to be detected belongs. However, the number of candidate category information corresponds to the bounding boxes; that is, each bounding box has its corresponding candidate category information, which is used to indicate the category to which the target in the bounding box may belong. For example... Figure 5aAs shown, the total output data volume after detecting an image to be detected is 1*84*8400, of which the output category data volume is 1*80*8400. Here, 1 indicates the number of images to be detected, 80 represents the confidence level of each of the 80 categories preset by the object detection model, and 8400 represents the sum of multiple bounding boxes at different scales. Specifically, the object detection model generates three feature maps at three different scales based on the image to be detected. The grid divisions corresponding to the three different feature maps are 80*80, 40*40, and 20*20, respectively. The 80*80 feature map has 80*80 grid points, and each grid point corresponds to one bounding box. Therefore, the 80*80 feature map has 80*80 grid points and 80*80 bounding boxes. Similarly, a 40*40 feature map corresponds to 40*40 grid points and 40*40 bounding boxes; a 20*20 feature map corresponds to 20*20 grid points and 20*20 bounding boxes. In summary, after the object detection model detects the image, it obtains a total of 80*80 + 40*40 + 20*20 = 8400 bounding boxes. The object detection model has 80 preset categories, so the category data size is 80*8400. Correspondingly, the bounding box coordinates require two diagonal coordinates (x1, y1) and (x2, y2), so the bounding box data size is 4*8400.

[0111] 302. The first processor determines the actual category of the target based on the candidate category information.

[0112] After obtaining candidate category information corresponding to multiple first bounding boxes, the first processor of the terminal device can determine the category corresponding to the multiple first bounding boxes based on the candidate category information, that is, the actual category of the multiple targets corresponding to the multiple first bounding boxes.

[0113] Specifically, the candidate category information includes multiple candidate categories to which the target corresponding to multiple first bounding boxes belongs, as well as the confidence level of each category to which the target may belong. Then, based on the multiple candidate categories to which the target belongs and the confidence level of each category, the category corresponding to the detection is determined. For example, for a target first bounding box, its corresponding candidate category information includes the target corresponding to the bounding box and the confidence level of each of the multiple categories. Then, the category with the highest confidence level is determined as the category corresponding to the bounding box, and this category indicates the actual category of the target within the bounding box.

[0114] In one possible implementation, after obtaining candidate category information, the NPU of the terminal device can determine the category with the highest confidence level as the actual category of the target based on the confidence level of each category to which the target may belong in the candidate category information.

[0115] For example, for a target bounding box, the target detection model may detect the possible categories to which it belongs and the confidence level of each category. For example, the confidence level of the target belonging to the category of computer is 0.9, the confidence level of the target belonging to the category of television is 0.7, and the confidence level of the target belonging to the category of monitor is 0.4. Based on this, the first processor of the terminal device can determine that the category corresponding to the target is the category with the highest confidence level, "computer".

[0116] In another possible implementation, after obtaining candidate category information, the first processor of the terminal device can determine the actual category of the target based on the credibility of each category the target might belong to in the candidate category information, and the relationship between each category. Specifically, the first processor can first determine multiple related categories among the multiple candidate categories to which the target belongs. These related categories include those related to other categories among the multiple candidate categories to which the target belongs, and those related to categories to which other targets belong. Then, among these related categories, the category with the highest credibility is determined as the actual category of the target.

[0117] Understandably, since the first processor may encounter false detections when detecting the actual category of the target, it can first select multiple related categories. These related categories usually belong to a unified scene, making it less likely that the detected category is a false detection. Therefore, the category with the highest confidence among these related categories can be determined as the actual category of the target. For example, among the multiple candidate categories for the first target, there are chairs, tables, and horses; among the multiple candidate categories for the second target, there are computers and televisions. It is understandable that chairs, tables, computers, and televisions belong to indoor scenes and can be considered related categories, while horses belong to outdoor scenes and are not related categories. When determining the actual category of the first target and the actual category of the second target, the first processor can prioritize the determination from the aforementioned related categories to improve the accuracy of the determined actual category of the target.

[0118] In this embodiment of the application, in the category data output by the target detection model, each category may correspond to a score value rather than a probability value. For example, the score for the category "computer" is 90, the score for the category "television" is 70, and the confidence level for the category "monitor" is 40. At this time, the first processor of the terminal device can use logistic regression to map these scores to the interval [0,1]. For example, after performing logistic regression, the probability of the category "computer" can be determined to be 0.9, the probability of the category "television" is 0.7, and the probability of the category "monitor" is 0.4.

[0119] Specifically, when processing categorical data, the terminal device can add a sigmoid layer to the data processing plugin to achieve the above logistic regression processing, and add a max layer to achieve the above maximum selection probability (confidence) processing. The specifics are not limited here.

[0120] In this embodiment of the application, by determining the actual category of the target corresponding to the bounding box based on candidate category information, the amount of category data that needs to be transmitted from the first processor to the second processor is reduced. For example... Figure 5b As shown, when the terminal device detects the image to be detected using the object detection model, the total output data volume is 1*84*8400, of which the output category data volume is 1*80*8400, where 80 represents the confidence level of each of the 80 categories, and 8400 represents the sum of the bounding boxes. Then, the terminal device can determine the actual category of the target based on the confidence level of each category. After processing the above category data, the final output category data volume should be 1*1*8400, which is 1 / 80 of the original 1*80*8400 category data volume.

[0121] 303. The first processor determines multiple second bounding boxes within multiple first bounding boxes.

[0122] The terminal device determines multiple second bounding boxes within multiple first bounding boxes. These multiple second bounding boxes are within the multiple first bounding boxes, and the scale of the corresponding feature map is a bounding box that meets preset conditions.

[0123] Specifically, when a terminal device detects an image using an object detection model, the image to be detected can be input into the model. The convolutional network of the object detection model generates multiple feature maps at different scales. Then, the object detection model determines multiple bounding boxes based on each feature map. In other words, the terminal device obtains multiple bounding boxes corresponding to each feature map based on these different scale feature maps; these multiple bounding boxes are the multiple first bounding boxes. Therefore, it can be understood that each feature map corresponds to multiple bounding boxes, and the terminal device determines multiple second bounding boxes from these first bounding boxes. Essentially, this involves selecting bounding boxes corresponding to a subset of the feature maps from among the bounding boxes corresponding to those feature maps.

[0124] Understandably, feature maps of different scales have their own advantages in detecting large, medium, and small targets. Large-scale feature maps (e.g., 80*80 scale feature maps) have small receptive fields, making them suitable for detecting small targets, allowing the terminal device to capture more details through the target detection model. Small-scale feature maps (e.g., 20*20 scale feature maps) have large receptive fields, making them suitable for detecting large targets and capturing overall target information. Medium-scale feature maps (e.g., 40*40 scale feature maps) perform well in detecting medium-sized targets. Therefore, in different scenarios, terminal devices can use feature maps of different scales for image detection.

[0125] Therefore, the terminal device detects images using an object detection model, obtaining multiple first bounding boxes. These first bounding boxes include multiple bounding boxes corresponding to feature maps of different scales. Then, multiple second bounding boxes are determined from these first bounding boxes, and the size of the feature maps corresponding to the bounding boxes in the multiple second bounding boxes meets a preset requirement. That is, among the multiple first bounding boxes, the bounding boxes corresponding to the feature maps of preset sizes are selected as the multiple second bounding boxes.

[0126] For example Figure 5c As shown, when the terminal device detects an image using the object detection model, the output bounding box data size is 1*4*8400. These bounding boxes are multiple first bounding boxes, where 1 indicates the number of images being detected, 4 represents the two vertex positions of the bounding boxes (x1, y1) and (x2, y2), and 8400 represents the total number of bounding boxes. This includes multiple bounding boxes corresponding to three feature maps of 80*80, 40*40, and 20*20. In some scenarios, such as when the terminal device is running a game application, the target size to be detected is usually moderate, so a medium-sized 40*40 feature map can be used. Therefore, the terminal device can determine multiple bounding boxes corresponding to the 40*40 feature map from the multiple first bounding boxes as multiple second bounding boxes. The 40x40 feature map corresponds to a grid of 40*40 points, and each grid point corresponds to a bounding box. Therefore, the number of bounding boxes corresponding to the 40x40 feature map is also 40*40. Thus, the final output data size for the multiple bounding boxes corresponding to the 40x40 feature map is 1*4*1600. Compared to the initial bounding box data size of 1*4*8400, the processed bounding box data size is 1*4*1600, representing a significant reduction in data size.

[0127] In this embodiment, by filtering out multiple second bounding boxes that are more suitable for the required scenario from multiple first bounding boxes, the amount of data that needs to be transmitted from the first processor to the second processor is reduced, the data transmission speed is improved, and the internal transmission resources of the terminal device are saved.

[0128] In this embodiment, by determining the actual category of the target corresponding to the bounding box in step 302, the amount of data that needs to be transmitted from the first processor to the second processor can be reduced. Similarly, by determining multiple second bounding boxes among multiple first bounding boxes in step 303, only the data of multiple second bounding boxes needs to be transmitted to the second processor, which can also reduce the amount of data that needs to be transmitted from the first processor to the second processor. Therefore, in this embodiment, if only step 302 or only step 303 is executed, the amount of data that needs to be transmitted from the first processor to the second processor can be reduced. However, preferably, both steps 302 and 303 are executed to minimize the amount of data that needs to be transmitted from the first processor to the second processor.

[0129] In this embodiment of the application, the execution order of steps 302 and 303 is not limited. That is, step 302 can be executed first and then step 303 can be executed; or step 303 can be executed first and then step 302. The specific order is not limited here.

[0130] In this embodiment, the terminal device can execute both step 302 and step 303. Specifically, regarding the category data in step 302, after determining the actual categories of multiple targets corresponding to multiple first bounding boxes based on the candidate category information, the actual categories of multiple targets corresponding to multiple second bounding boxes can also be determined based on the actual categories of the multiple targets corresponding to the multiple first bounding boxes. Then, the actual categories of the multiple targets corresponding to the multiple second bounding boxes are transmitted to the second processor, further reducing the amount of data that needs to be transmitted from the first processor to the second processor.

[0131] In one possible implementation, for example, with the aforementioned 1*80*8400 category data, the first processor can first determine the category data corresponding to the second bounding boxes, that is, the 40*40 second bounding boxes corresponding to the 40*40 feature maps in the aforementioned 8400 bounding boxes. Therefore, the number of candidate category data corresponding to the second bounding box is 1*80*1600. Then, the actual category of the target corresponding to the second bounding box is determined, resulting in multiple actual categories of the target corresponding to the second bounding boxes. The data volume of this category data is 1*1*1600. Compared to the aforementioned 1*1*8400 actual category data of the target corresponding to the first bounding box, this category data is further reduced.

[0132] In another possible implementation, for example, for the aforementioned 1*80*8400 category data, the first processor can determine the actual category data volume of the target corresponding to the first bounding box as 1*1*8400 using the above method, and then determine the actual category data volume of the target corresponding to the second bounding box as 1*1*1600, thereby further reducing the category data.

[0133] In summary, both implementation methods involve determining the actual category of the target and filtering based on the corresponding feature map when processing candidate category data. The order of execution differs, but the results and the achieved effects are the same. The specific implementation steps will not be elaborated here.

[0134] In one possible implementation, when the current display interface of the terminal device is a clickable interface, the scene recognition method in this embodiment may further include step 304:

[0135] 304. The first processor determines the coordinate information of the clicked target based on the image detection results.

[0136] When the current display interface of the terminal device is a click interface, the terminal device determines the coordinates of the click target corresponding to the click interface based on the image detection results. The image detection results include the coordinate information of the multiple first bounding boxes and the actual category of the target corresponding to the multiple first bounding boxes, or the coordinate information of the multiple second bounding boxes and the actual category of the target corresponding to the multiple second bounding boxes.

[0137] Specifically, after acquiring interface information, if the interface information indicates that the current display interface of the terminal device is a clickable interface, the first processor can determine the coordinate information of the clickable target corresponding to the clickable interface based on the image detection results. For example... Figure 6 As shown, the first processor can determine the coordinate information of the clicked target through the following steps:

[0138] S1. The first processor determines the third bounding box based on the image detection results.

[0139] In a click scenario, the first processor determines at least one third bounding box from among the multiple second bounding boxes based on the image detection results. The actual category of the target corresponding to the third bounding box is the click target.

[0140] In one possible implementation, the first processor will also execute step S2, filtering the third bounding box based on its confidence level. Specifically:

[0141] S2. The first processor determines whether the confidence level corresponding to the third bounding box is greater than the threshold.

[0142] After determining the third bounding box, the first processor determines whether the confidence level corresponding to the at least one third bounding box is greater than a preset threshold. If it is greater than the preset threshold, step S3 is executed for the third bounding box with a confidence level greater than the threshold; if it is not greater than the threshold, step S3 is not executed.

[0143] In one possible implementation, the number of clickable targets on the current interface is N. If the number of determined third bounding boxes is greater than N, the first processor determines the top N targets with the highest confidence as third bounding boxes that meet the preset conditions, and performs step S3 on these third bounding boxes that meet the preset conditions.

[0144] S3. The first processor determines the center coordinates of the third bounding box.

[0145] After determining the third bounding box, the first processor determines the center coordinates of the third bounding box.

[0146] Specifically, in this embodiment of the application, the coordinate information of the bounding box includes the coordinates of the two opposite corners of the bounding box (x1, y1) and (x2, y2), and then the center coordinates of the third bounding box are determined as (x1 / 2+x2 / 2, y1 / 2+y2 / 2) based on the coordinates of the two opposite corners (x1, y1) and (x2, y2).

[0147] In one possible implementation, during a click scenario, the first processor of the terminal device can directly transmit the center coordinates of the third bounding box, eliminating the need to transmit the two coordinates of the third bounding box. This further reduces the amount of data that needs to be transmitted and lowers the data transmission pressure within the terminal device. In another possible implementation, during a click scenario, the second processor has already determined that it is a click scenario. The first processor only needs to transmit the center coordinates of the third bounding box to the second processor, without needing to transmit multiple first bounding boxes and the actual categories of the targets corresponding to those first bounding boxes, or multiple second bounding boxes and the actual categories of the targets corresponding to those second bounding boxes.

[0148] 305. The first processor transmits the image detection results to the second processor.

[0149] After the first processor of the terminal device obtains the image detection result, it can transmit the image detection result to the second processor. The image detection result includes the coordinates of multiple first bounding boxes and the actual category of the target corresponding to the multiple first bounding boxes, or the coordinates of multiple second bounding boxes and the actual category of the target corresponding to the multiple second bounding boxes.

[0150] In one possible implementation, this embodiment of the application further includes step 304, whereby the image detection result may also include the coordinates of the clicked target.

[0151] In one possible implementation, the image detection result transmitted from the first processor to the second processor may not include the coordinates of the multiple first bounding boxes and the actual category of the target corresponding to the multiple first bounding boxes, or the multiple second bounding boxes and the actual category of the target corresponding to the multiple second bounding boxes, but only include the center coordinates of the third bounding box.

[0152] In this embodiment of the application, the above-mentioned scene recognition method can be implemented by a software plugin. For example, the software plugin can be combined with a target detection model so that the plugin can obtain the output of the target detection model and perform the above-mentioned processing on the data output by the target detection model to realize the above-mentioned scene recognition method.

[0153] In this embodiment, the first processor may be the NPU of the terminal device, and the second processor may be the GPU of the terminal device; otherwise, the first processor may be other processors, and the second processor may also be other processors, which are not limited here.

[0154] Based on the above technical solution, the first processor can determine the actual category of the target based on the candidate category information, and then transmit the actual category of the target to the second processor; it can also determine multiple second bounding boxes in multiple first bounding boxes, and then transmit the multiple second bounding boxes to the second processor; this reduces the amount of data that the first processor needs to transmit to the second processor, the first processor does not need to transmit a large amount of data to the second processor, reduces the data transmission pressure between the first processor and the second processor inside the terminal device, reduces the amount of data that the second processor needs to process, and improves the scene recognition efficiency of the terminal device.

[0155] The above combination Figures 3 to 6 The image processing method provided in the embodiments of this application has been described. The terminal device that performs the above image processing method provided in the embodiments of this application is described below.

[0156] like Figure 7 As shown, Figure 7 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 7 As shown, the terminal device 700 may include:

[0157] The detection module 701 is used to detect the image to be detected and obtain multiple first bounding boxes and candidate category information corresponding to the target in the image to be detected. The candidate category information indicates the multiple candidate categories to which the target belongs.

[0158] The first determining module 702 is used to determine the actual category of the target based on the candidate category information;

[0159] In one possible implementation, the first determining module 702 is specifically used to: determine the credibility of each of the multiple candidate categories to which the target belongs based on the candidate category information; and determine the category with the highest credibility as the actual category of the target based on the credibility of each of the multiple candidate categories.

[0160] In one possible implementation, the first determining module 702 is specifically used to: determine multiple candidate categories that are related among multiple candidate categories to which the target belongs; determine the credibility of each category among the multiple candidate categories that are related based on the candidate category information; and determine the category with the highest credibility as the actual category of the target based on the credibility of each category among the multiple candidate categories that are related.

[0161] In one possible implementation, the terminal device further includes a second determining module 703, which is used to: determine a plurality of second bounding boxes among a plurality of first bounding boxes, wherein the scale of the feature map corresponding to the second bounding boxes is a preset scale; and determine the actual category of the target corresponding to the plurality of second bounding boxes based on candidate category information.

[0162] In one possible implementation, the second determining module 703 is further configured to: in a click scenario, determine at least one third bounding box based on the image detection result, wherein the actual category of the target corresponding to the third bounding box is the click target; and determine the center coordinates of the at least one third bounding box.

[0163] The transmission module 704 is used to transmit image detection results to the second processor. The image detection results include the coordinates of multiple first bounding boxes and the actual categories of the targets corresponding to the multiple first bounding boxes.

[0164] Figure 8 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Figure 8 As shown, the terminal device 800 includes one or more processors 801, communication lines 802 and communication interfaces 803. Optionally, the terminal device 800 also includes a memory 804.

[0165] In some implementations, memory 804 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.

[0166] The methods described in the embodiments of this application can be applied to or implemented by processor 801. Processor 801 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above methods can be completed by the integrated logic circuit in the hardware of processor 801 or by instructions in software form. The processor 801 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. Processor 801 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0167] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 804, and processor 801 reads information from memory 804 and, in conjunction with its hardware, completes the steps of the above method.

[0168] The processor 801, memory 804 and communication interface 803 can communicate with each other through communication line 802.

[0169] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.

[0170] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the process or function performed by the satellite base station or user equipment according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).

[0171] This application provides a terminal device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the computer program to perform the image processing method described above.

[0172] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program or instructions. When the computer program or instructions are executed by a processor, they implement the method executed by the aforementioned terminal device. The methods described in the above embodiments can be implemented wholly or partially by software, hardware, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted on the computer-readable medium. The computer-readable medium can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0173] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0174] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0175] The above specific embodiments further illustrate the purpose, technical solution and beneficial effects of this application. It should be understood that the above are only specific embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of this application should be included within the scope of protection of this application.

Claims

1. A scene recognition method, characterized in that, A first processor applied to a terminal device, the terminal device including the first processor and a second processor, the method comprising: The image to be detected is detected to obtain multiple first bounding boxes and candidate category information corresponding to the target in the image to be detected, wherein the candidate category information indicates the multiple candidate categories to which the target belongs; The actual category of the target is determined based on the candidate category information; The image detection results are transmitted to the second processor, the image detection results including the coordinates of the plurality of first bounding boxes and the actual category of the target.

2. The method according to claim 1, characterized in that, Determining the actual category of the target based on the candidate category information includes: The credibility of each of the multiple candidate categories to which the target belongs is determined based on the candidate category information; Based on the credibility of each of the multiple candidate categories, the category with the highest credibility is determined as the actual category of the target.

3. The method according to claim 1, characterized in that, Determining the actual category of the target based on the candidate category information includes: Identify multiple candidate categories that are related to the target from among multiple candidate categories; The credibility of each of the multiple candidate categories with correlation is determined based on the candidate category information; Based on the credibility of each of the multiple candidate categories with the aforementioned relationship, the category with the highest credibility is determined as the actual category of the target.

4. The method according to any one of claims 1-3, characterized in that, Before transmitting the image detection results to the second processor, the method further includes: Multiple second bounding boxes are determined within the plurality of first bounding boxes, and the scale of the feature map corresponding to the second bounding box is a preset scale; The actual category of the target corresponding to the plurality of second bounding boxes is determined based on the candidate category information; Correspondingly, the transmission of image detection results to the second processor includes: The image detection results are transmitted to the second processor, the image detection results including the coordinates of the plurality of second bounding boxes and the actual category of the target.

5. The method according to claim 4, characterized in that, The method further includes: Determine the current application of the terminal device; The preset scale of the feature map is determined based on the current application of the terminal device.

6. The method according to claim 5, characterized in that, The method further includes: In a click scenario, at least one third bounding box is determined based on the image detection results, and the actual category of the target corresponding to the third bounding box is the click target; Determine the center coordinates of the at least one third bounding box; The center coordinates of the at least one third bounding box are transmitted to the second processor.

7. The method according to claim 6, characterized in that, The first processor is a neural network processor (NPU), and the second processor is a graphics processing unit (GPU).

8. A terminal device, characterized in that, The terminal device includes: The detection module is used to detect the image to be detected and obtain multiple first bounding boxes and candidate category information corresponding to the target in the image to be detected, wherein the candidate category information indicates the multiple candidate categories to which the target belongs; The determination module is used to determine the actual category of the target based on the candidate category information; A transmission module is used to transmit image detection results to the second processor, the image detection results including the coordinates of the plurality of first bounding boxes and the actual category of the target.

9. A terminal device, characterized in that, It includes a memory and a processor, the memory being used to store a computer program, and the processor being used to invoke the computer program to execute the scene recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions, which, when executed, implement the scene recognition method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, Includes a computer program, which, when run, causes the computer to perform the scene recognition method as described in any one of claims 1 to 7.