Scene recognition method and electronic device
Patent Information
- Application Number
- PCT/CN2025/085568
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025085568_01102026_PF_FP_ABST
Abstract
Description
A scene recognition method and electronic device Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a scene recognition method and electronic device. Background Technology
[0002] Electronic devices often perform scene recognition on the interface images of some applications (such as game applications) in order to adopt different optimization strategies for different scenes.
[0003] In the process of scene recognition of interface images by electronic devices, if the electronic devices need to maintain a high recognition accuracy, they need to detect and recognize the entire area of the interface image for a long time and at a high frequency. As a result, the electronic devices generate a lot of unnecessary scene recognition operations, leading to high power consumption and affecting system performance. Summary of the Invention
[0004] This application provides a scene recognition method and an electronic device that can perform scene recognition on the interface image of the electronic device based on different levels. In this way, while ensuring the accuracy of scene recognition, this application does not need to perform scene recognition on the entire area of the interface image, but only on a part of the interface image, avoiding unnecessary scene recognition operations, reducing the power consumption of the electronic device, and improving the system performance of the electronic device.
[0005] In a first aspect, embodiments of this application provide a scene recognition method applied to an electronic device, comprising: displaying a first interface; determining a first region and a second region in the first interface, wherein the second region includes at least a portion of the first interface excluding the first region, and the area of the second region is greater than or equal to the area of the first region; performing scene recognition on the image corresponding to the second region to obtain a first scene type corresponding to the second region; and, if the first scene type is a key scene, performing scene recognition on the image corresponding to the first region to determine whether the first interface is in a target scene.
[0006] The scene recognition method provided in this application does not require scene recognition for the entire area of the interface image. Instead, it forms different levels by dividing the interface image into regions, and then performs scene recognition on a portion of the interface image of the electronic device based on the different levels. In this way, while ensuring the accuracy of scene recognition, unnecessary scene recognition operations are avoided, the power consumption of the electronic device is reduced, and the system performance of the electronic device is improved.
[0007] In one implementation, after performing scene recognition on the image corresponding to the second region to obtain the first scene type corresponding to the second region, the method further includes: if the first scene type is a non-critical scene, then in both the first and second regions, scene recognition is performed only on the image corresponding to the second region based on a preset first frequency, and scene recognition is not performed on the image corresponding to the first region. This implementation, by performing scene recognition only on the image corresponding to the second region at a lower frequency, can save power consumption of the electronic device.
[0008] In one implementation, when the first scene type is a critical scene, scene recognition is performed on the image corresponding to the first region, including: performing scene recognition on the image corresponding to the first region based on a preset second frequency, wherein the second frequency is greater than the first frequency. By using this implementation, performing scene recognition on the image corresponding to the first region based on a higher frequency can improve the accuracy of scene recognition.
[0009] In one implementation, displaying the first interface includes: responding to a first operation on the game application, running the game application, and displaying the first interface, which is the game running interface. This implementation allows for different levels of scene recognition in game scenarios, thereby reducing the power consumption of electronic devices and improving system performance in high-power scenarios like game scenarios.
[0010] In one implementation, determining a first region and a second region in a first interface includes: dividing a first image displayed on the first interface into multiple first image blocks; determining a first region enclosed by preset first boundary coordinates, wherein the first region is composed of a portion of the multiple first image blocks; and obtaining a difference region between the first image and the first region based on the second boundary coordinates corresponding to the first image, thereby determining the difference region as the second region, wherein the second region is composed of another portion of the multiple first image blocks. Using this implementation, the electronic device can accurately locate the boundaries between the first region and the second region. Thus, even when performing scene recognition on complex images, the electronic device can accurately distinguish different regions, avoiding interference from image content.
[0011] In one implementation, scene recognition is performed on the image corresponding to the second region to obtain a first scene type corresponding to the second region. This includes: dividing a preset image into multiple second image blocks, wherein each second image block corresponds to a position of a first image block; obtaining each first image block in the second region, and a first average distance between the second image blocks corresponding to each first image block; if the minimum value of the first average distance is less than a first average distance threshold, then the first scene type corresponding to the second region is determined to be a critical scene; if the minimum value of the first average distance is greater than or equal to the first average distance threshold, then the first scene type corresponding to the second region is determined to be a non-critical scene. This implementation achieves the recognition of non-critical scenes through a low-consumption, low-cost, and highly tolerant algorithm, reducing the power consumption of electronic devices.
[0012] In one implementation, when the first scene type is a key scene, scene recognition is performed on the image corresponding to the first region to determine whether the first interface is in the target scene. This includes: acquiring a first image feature of the first region image corresponding to the first region, wherein the first region image is a part of a first image, or a part of a second image, the second image being displayed after the first image is displayed on the first interface; acquiring a second image feature of the target region in a preset image, wherein the target region corresponds to the position of the first region; acquiring a first Euclidean distance between the first image feature and the second image feature; if the first Euclidean distance is less than a preset first Euclidean distance threshold, then the scene recognition of the first region is determined to be successful, thereby determining that the first interface is in the target scene. This implementation achieves target scene recognition through a high-precision algorithm, improving the accuracy of the scene recognition process.
[0013] In one implementation, the first image feature includes at least one of Scale Invariant Feature Transform (SIFT) features, Orientation Fast (OFFT) features, and Rotation Invariant (ORB) features; the second image feature includes at least one of Scale Invariant Feature Transform (SIFT) features, Orientation Fast (OFFT) features, and Rotation Invariant (ORB) features; the first image feature and the second image feature have the same feature type. Using this implementation, electronic devices can support scene recognition based on multiple different types of image features, improving the compatibility of electronic devices.
[0014] In one implementation, before displaying the first interface, the method further includes: setting a preset image, and setting first boundary coordinates based on the preset image. This implementation demonstrates a preprocessing method for an electronic device, whereby, based on the preset first boundary coordinates, a first region and a second region are divided to perform scene recognition on the interface image of the electronic device at different levels.
[0015] In one implementation, the first boundary coordinates are preset based on a preset image, including: comparing at least one third image with the preset image using cluster analysis to determine the first boundary coordinates, wherein the third image is used for display in key scenes. This implementation automatically determines the boundary coordinates through cluster analysis, reducing manual intervention and enabling electronic devices to accurately locate the boundary between the first and second regions.
[0016] In one implementation, at least one third image is compared with a preset image using clustering analysis to determine the first boundary coordinates. This includes: dividing each third image into multiple third image blocks and dividing the preset image into multiple second image blocks, wherein each third image block in each third image corresponds to a second image block; obtaining a second average distance between each third image block and each second image block in each third image to determine a target average distance for each third image, wherein the target average distance includes the N largest second average distances, where N is a positive integer greater than 0; obtaining a first number of third image blocks corresponding to the target average distance in each third image; determining the third image corresponding to the first number with the minimum value as the target image; and determining the boundary coordinates of the region containing the third image block corresponding to the target average distance in the target image as the first boundary coordinates. Using this implementation, by calculating the second average distance between the third image blocks and the second image blocks and obtaining the first number corresponding to the target average distance, target images can be effectively filtered out, thereby accurately determining the first boundary coordinates in the target image to achieve scene recognition at different levels.
[0017] In one implementation, in each third image, obtaining a second average distance between each third image block and each second image block includes: in each third image, dividing a plurality of third image blocks into a first cluster of image blocks and a second cluster of image blocks, wherein the first cluster of image blocks includes a portion of the third image blocks, and the second cluster of image blocks includes another portion of the third image blocks; determining a second Euclidean distance between each first pixel and each second pixel based on each first pixel in each third image block in the first cluster of image blocks, and a second pixel in the second image block corresponding to the first pixel; and determining a second Euclidean distance between each first pixel and each second pixel based on each first pixel in each third image block in the second cluster of image blocks. The third Euclidean distance between each third pixel and each fourth pixel is determined using three pixels and a fourth pixel corresponding to the third pixel in the second image block. Based on all the second Euclidean distances corresponding to each third image block in the first cluster of image blocks, a third average distance is determined between each third image block in the first cluster of image blocks and its corresponding second image block. Based on all the third Euclidean distances corresponding to each third image block in the second cluster of image blocks, a fourth average distance is determined between each third image block in the second cluster of image blocks and its corresponding second image block, wherein the second average distance includes both the third and fourth average distances. This implementation demonstrates a specific method for determining the second average distance, which facilitates the subsequent determination of the target average distance from the second average distance, effectively filtering out target images, thereby accurately determining the first boundary coordinates in the target image and achieving scene recognition at different levels.
[0018] In one implementation, within each third image, multiple third image blocks are divided into a first cluster image block and a second cluster image block, including: arbitrarily determining two cluster centers within the multiple third image blocks; determining the target Euclidean distance between each third image block and each cluster center; assigning each third image block to a cluster center with the smallest target Euclidean distance based on the two target Euclidean distances corresponding to each third image block, to obtain a first cluster target image block corresponding to one cluster center and a second cluster target image block corresponding to the other cluster center; determining the first cluster target image block and its corresponding cluster. The method calculates the first average error between the centers and the second average error between the second cluster target image patch and its corresponding cluster center. If the first average error and the second average error are different, the two cluster centers are updated, and the target Euclidean distance between each third image patch and each updated cluster center is determined again, thus obtaining the updated first average error and the updated second average error. If the updated first average error and the updated second average error are the same, the updated first cluster target image patch is determined as the first cluster image patch, and the updated second cluster target image patch is determined as the second cluster image patch. By dividing the third image patch into two clusters, this method can more clearly distinguish regions with different features, facilitating subsequent division of the first and second regions, thereby achieving scene recognition at different levels.
[0019] In one implementation, a second Euclidean distance is determined between each first pixel and each second pixel based on each first pixel in each third image block within a first cluster of image blocks, and the second pixel corresponding to the first pixel in a second image block. This includes: setting a first color value (R1, G1, B1) for each first pixel based on RGB coordinates, and setting a second color value (R2, G2, B2) for each second pixel based on RGB coordinates; obtaining a first difference between a first red component R1 in the first color value and a second red component R2 in the second color value, a second difference between a first green component G1 in the first color value and a second green component G2 in the second color value, and a third difference between a first blue component B1 in the first color value and a second blue component B2 in the second color value; and determining the square root of the sum of the squares of the first, second, and third differences as the second Euclidean distance. This implementation illustrates a specific method for determining the second Euclidean distance, which is then used to subsequently obtain the second average distance. In this way, electronic devices can effectively filter out target images based on the second average distance, thereby accurately determining the first boundary coordinates in the target image and realizing scene recognition at different levels.
[0020] In one implementation, a third Euclidean distance is determined between each third pixel and each fourth pixel in each third image block within a second cluster of image blocks, and a fourth pixel in the second image block corresponding to the third pixel. This includes: setting the third color value corresponding to each third pixel to (R3, G3, B3) based on RGB coordinates, and setting the fourth color value corresponding to each fourth pixel to (R4, G4, B4) based on RGB coordinates; obtaining the fourth difference between the third red component R3 and the fourth red component R4 in the third color value, the fifth difference between the third green component G3 and the fourth green component G4 in the third color value, and the sixth difference between the third blue component B3 and the fourth blue component B4 in the third color value; and determining the square root of the sum of the squares of the fourth, fifth, and sixth differences as the third Euclidean distance. This implementation illustrates a specific method for determining the third Euclidean distance, which is then used to subsequently obtain the second average distance. In this way, electronic devices can effectively filter out target images based on the second average distance, thereby accurately determining the first boundary coordinates in the target image and realizing scene recognition at different levels.
[0021] In one implementation, based on all the second Euclidean distances corresponding to each third image block in the first cluster of image blocks, a third average distance is determined between each third image block in the first cluster of image blocks and its corresponding second image block. This includes: in each third image block in the first cluster of image blocks, dividing the sum of all the second Euclidean distances in each third image block by the total number of first pixels in each first image block to obtain the third average distance between each third image block and its corresponding second image block. This implementation illustrates a specific method for determining the third average distance, which is then used to obtain the second average distance. In this way, the electronic device can effectively filter out target images based on the second average distance, thereby accurately determining the first boundary coordinates in the target image and achieving scene recognition at different levels.
[0022] In one implementation, based on all third Euclidean distances corresponding to each third image block in the second cluster of image blocks, a fourth average distance is determined between each third image block in the second cluster of image blocks and its corresponding second image block. This includes: in each third image block in the second cluster of image blocks, dividing the sum of all third Euclidean distances in each third image block by the total number of third pixels in each third image block to obtain the fourth average distance between each third image block and its corresponding second image block. This implementation illustrates a specific method for determining the fourth average distance, which is then used to obtain the second average distance. In this way, the electronic device can effectively filter out target images based on the second average distance, thereby accurately determining the first boundary coordinates in the target image and achieving scene recognition at different levels.
[0023] In one implementation, determining the target average distance for each third image includes: in each second image, comparing each second average distance with a preset second average distance threshold, and determining N second average distances greater than the second average distance threshold as the target average distance. This implementation demonstrates a specific method for obtaining the target average distance based on threshold comparison. This method simplifies the decision-making process of electronic devices, avoids complex statistical analysis, and can quickly filter out target images, thereby accurately determining the first boundary coordinates in the target image and achieving scene recognition at different levels.
[0024] In one implementation, determining the target average distance for each third image includes: sorting the third and fourth average distances from largest to smallest numerical value, and determining the average distance of the top N largest numerical values as the target average distance. This implementation demonstrates a specific method for obtaining the target average distance based on sorting. This method accurately obtains the target average distance, enabling rapid filtering of target images. This allows electronic devices to accurately determine the first boundary coordinates in the target image, achieving scene recognition at different levels.
[0025] Secondly, embodiments of this application provide a scene recognition device, which includes a display module for displaying a first interface; a determination module for determining a first region and a second region in the first interface, wherein the second region includes at least a portion of the first interface excluding the first region, and the area of the second region is greater than or equal to the area of the first region; a first recognition module for performing scene recognition on the image corresponding to the second region to obtain a first scene type corresponding to the second region; and a second recognition module for performing scene recognition on the image corresponding to the first region when the first scene type is a key scene, to determine whether the first interface is in a target scene.
[0026] The scene recognition device provided in this application does not require scene recognition for the entire area of the interface image. Instead, it forms different levels by dividing the interface image into regions, and performs scene recognition on a portion of the interface image of the electronic device based on the different levels. In this way, while ensuring the accuracy of scene recognition, unnecessary scene recognition operations are avoided, the power consumption of the electronic device is reduced, and the system performance of the electronic device is improved.
[0027] Thirdly, embodiments of this application provide an electronic device, which includes a touch screen, a memory, and one or more processors; the touch screen, the memory, and the processor are coupled; wherein the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the scene recognition method provided by the first aspect and any of its implementations.
[0028] Fourthly, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the scene recognition method provided by the first aspect and any of its implementations.
[0029] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the scene recognition method provided by the first aspect and any of its implementations.
[0030] It is understood that the beneficial effects that the technical solutions provided in the third to fifth aspects above can achieve can be referred to the beneficial effects in the first aspect and any of its implementation methods, and will not be repeated here. Attached Figure Description
[0031] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 is a schematic diagram of the first scene in which the electronic device performs scene recognition on the interface image;
[0033] Figure 2 is a schematic diagram of the second scene of scene recognition of interface images by electronic devices;
[0034] Figure 3 is a schematic diagram of the third scene of scene recognition of interface images by electronic devices;
[0035] Figure 4 is a schematic diagram of the fourth scene in which the electronic device performs scene recognition on the interface image;
[0036] Figure 5 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application;
[0037] Figure 6 is a schematic diagram of the software structure of the electronic device provided in an embodiment of this application;
[0038] Figure 7 is a schematic flowchart of the scene recognition method provided in the embodiment of this application;
[0039] Figure 8 is a first interactive schematic diagram of the scene recognition method provided in the embodiment of this application;
[0040] Figure 9 is a schematic diagram of a preset image provided in an embodiment of this application;
[0041] Figure 10 is a schematic diagram of a first interface image provided in an embodiment of this application for comparison with a preset image;
[0042] Figure 11 is a schematic diagram of the first interface image block and the second interface image block provided in the embodiments of this application;
[0043] Figure 12 is a schematic diagram of the method for determining the first cluster of image blocks and the second cluster of image blocks provided in the embodiments of this application;
[0044] Figure 13 is a schematic diagram of the first cluster of image blocks and the second cluster of image blocks in different first interface images provided in the embodiments of this application;
[0045] Figure 14 is a schematic diagram of a scenario in which the electronic device provided in the embodiment of this application determines the Euclidean distance between the first pixel and the second pixel.
[0046] Figure 15 is a first schematic diagram of the interface area corresponding to the target average distance provided in the embodiments of this application;
[0047] Figure 16 is a second schematic diagram of the interface area corresponding to the target average distance provided in the embodiment of this application;
[0048] Figure 17 is a third schematic diagram of the interface area corresponding to the target average distance provided in the embodiments of this application;
[0049] Figure 18 is a schematic diagram of a scenario in which an electronic device acquires a first image according to an embodiment of this application;
[0050] Figure 19 is a schematic diagram of the interactive process of an electronic device acquiring a first image according to an embodiment of this application;
[0051] Figure 20 is a schematic diagram of a scenario where an electronic device provides an embodiment of this application determines a first region;
[0052] Figure 21 is a schematic diagram of a scenario where an electronic device provides an embodiment of this application determines a second region;
[0053] Figure 22 is a first schematic diagram of scene recognition of a second region by an electronic device provided in an embodiment of this application;
[0054] Figure 23 is a second schematic diagram of scene recognition of the second region by the electronic device provided in the embodiment of this application;
[0055] Figure 24 is a third schematic diagram of the electronic device provided in this application performing scene recognition in the second region;
[0056] Figure 25 is a fourth schematic diagram of scene recognition of the second region by the electronic device provided in the embodiment of this application;
[0057] Figure 26 is a schematic diagram of scene recognition of a first region by an electronic device provided in an embodiment of this application;
[0058] Figure 27 is a second flowchart of the scene recognition method provided in the embodiments of this application;
[0059] Figure 28 is a schematic diagram of the scene recognition device provided in an embodiment of this application;
[0060] Figure 29 is a schematic diagram of the structure of a scene recognition device provided in another embodiment of this application;
[0061] Figure 30 is a schematic diagram of the chip system provided in an embodiment of this application. Detailed Implementation
[0062] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the protection scope of this application.
[0063] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0064] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0065] The application scenarios of the embodiments of this application will be described below first.
[0066] Electronic devices often perform scene recognition on the interface images of some applications (such as game applications) in order to adopt different optimization strategies for different scenes.
[0067] The process of scene recognition of interface images of an action game application using an electronic device is illustrated by way of example.
[0068] Figure 1 is a schematic diagram of the first scene in which an electronic device performs scene recognition on an interface image.
[0069] As shown in Figure 1, the electronic device runs an action game application, displays a first game interface, and displays at least one frame of the first game interface image 1 in the first game interface.
[0070] Users can control a first game character (2) to perform real-time physical actions in an action game application. These actions include walking, running, jumping, fighting, climbing, and throwing. The electronic device can display these real-time physical actions on the first game interface image (1), for example, displaying combat actions performed by the first game character (2).
[0071] Based on the network synchronization mechanism, the electronic device can also display real-time physical actions performed by game characters controlled by other users in the first game interface image 1, such as displaying the running action performed by the second game character 3 and the combat action performed by the third game character 4.
[0072] During the aforementioned process, the electronic device needs to perform long-term, high-frequency scene recognition on the entire interface area of the first game interface image 1 to identify that the first game interface image 1 is in a team battle scene (i.e., a scene where multiple game characters are fighting on the same map). In this way, the electronic device can adopt optimization strategies to increase the frequency of the central processing unit (CPU) and / or the graphics processing unit (GPU) to avoid lag caused by insufficient resource supply.
[0073] However, this long-term, high-frequency scene recognition process is not applicable to all scenes corresponding to the first game interface image 1.
[0074] Figure 2 is a schematic diagram of the second scene in which the electronic device performs scene recognition on the interface image.
[0075] As shown in Figure 2, during scene recognition of the first game interface image 1 by the electronic device, if the real-time physical action of the first game character 2 controlled by the user in an action game application is walking, and the electronic device only displays the first game character 2 in the first game interface image 1 without displaying other game characters, it indicates that the first game interface image 1 is in a non-team battle scene. In this scenario, the electronic device does not actually need to perform scene recognition on the entire interface area of the first game interface image 1 for a long time and at a high frequency to determine when to adopt optimization strategies to increase the CPU and / or GPU frequencies. In other words, the long-term, high-frequency scene recognition operation performed by the electronic device in this scenario is unnecessary. These unnecessary scene recognition operations can easily lead to higher power consumption of the electronic device.
[0076] In other scenarios, electronic devices also perform unnecessary scene recognition operations.
[0077] The process of scene recognition of the interface images of a gacha game application using an electronic device is illustrated by way of example.
[0078] Figure 3 is a schematic diagram of the third scene in which the electronic device performs scene recognition on the interface image.
[0079] As shown in Figure 3, the electronic device runs a card-drawing game application, displays a second game interface, and displays at least one frame of the second game interface image 5 in the second game interface.
[0080] Users can click the card-drawing button 6 in the card-drawing game application. In response to the click of the card-drawing button 6, the electronic device displays a card-drawing animation. After the card-drawing animation ends, the game character or equipment drawn is displayed on the second game interface image 5.
[0081] During the aforementioned process, the electronic device needs to perform long-term, high-frequency scene recognition on the entire interface area of the second game interface image 5 to identify that the second game interface image 5 is in a card-drawing scene. In this way, the electronic device can record the moment when a rare equipment or rare character is drawn, and optimize the card-drawing strategy based on the recorded moments.
[0082] However, this long-term, high-frequency scene recognition process is not applicable to all scenes corresponding to the second game interface image 5.
[0083] Figure 4 is a schematic diagram of the fourth scene in which the electronic device performs scene recognition on the interface image.
[0084] As shown in Figure 4, during the scene recognition process of the electronic device on the second game interface image 5, if the user controls the fourth game character 7 to perform a walking action in the real-time physical action of the gacha game application, it indicates that the second game interface image 5 is in a non-gacha scene. In this scenario, the electronic device does not actually need to perform scene recognition on the entire interface area of the second game interface image 5 for a long time and at a high frequency to record the moment when a rare equipment or rare character is drawn. In other words, the long-term, high-frequency scene recognition operation performed by the electronic device in this scenario is unnecessary. These unnecessary scene recognition operations can easily lead to higher power consumption of the electronic device.
[0085] Since game applications themselves have high loads, such unnecessary scene recognition operations performed by electronic devices on the entire area of the interface image will further increase power consumption and affect system performance.
[0086] To address the aforementioned issues, this application provides a scene recognition method.
[0087] The scene recognition method provided in this application can be applied to electronic devices. These electronic devices include, but are not limited to, mobile phones, tablets, personal computers, personal digital assistants (PDAs), workstations, large-screen devices (e.g., smart screens, smart TVs), wearable devices (e.g., smart bracelets, smartwatches), handheld game consoles, home game consoles, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality devices, and in-vehicle intelligent terminals. This application does not limit the specific technology or form of the electronic device. The electronic devices involved in this application can be equipped with… Harmony This application does not restrict the use of other operating systems.
[0088] Figure 5 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application.
[0089] As shown in Figure 5, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 122, antennas 01 and 02, a mobile communication module 150, a wireless communication module 160, a sensor module 180, a camera 192, and a display screen 193. The sensor module 180 may include a touch sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a geomagnetic sensor 180D, an accelerometer sensor 180E, a proximity sensor 180F, and a proximity light sensor 180G. The gyroscope sensor 180B, barometric pressure sensor 180C, geomagnetic sensor 180D, and accelerometer sensor 180E can all be used to detect the motion state of the electronic device; therefore, they can also be called motion sensors.
[0090] Processor 110 may include one or more processing units, such as a central processing unit (CPU), application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.
[0091] In this embodiment, the GPU can be used to process graphics rendering and image processing tasks. For example, in a game scene, the GPU can render 3D models and scenes to provide an immersive gaming experience. The GPU can also perform graphics pipeline processing to ensure real-time rendering of game visuals. The GPU can also execute shader programs to control the visual effects of graphics rendering. In this embodiment, the GPU can be used to draw a first image and display it.
[0092] The controller can generate operation control signals based on the instruction opcode and timing signals, thereby controlling the process of acquiring and executing instructions.
[0093] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0094] DSPs are used to process digital signals. In addition to processing digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency point, a digital signal processor is used to perform Fourier transforms on the frequency point energy, etc.
[0095] NPU stands for Neural Network (NN) computing processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.
[0096] In some embodiments, processor 110 may include one or more interfaces. Interfaces may include Inter-Integrated Circuit (I2C) interfaces, Inter-Integrated Circuit Sound (I2S) interfaces, Pulse Code Modulation (PCM) interfaces, Universal Asynchronous Receiver / Transmitter (UART) interfaces, Mobile Industry Processor Interface (MIPI) interfaces, General-Purpose Input / Output (GPIO) interfaces, Subscriber Identity Module (SIM) interfaces, and / or Universal Serial Bus (USB) interfaces, etc. These interfaces are used to connect to other components in electronic device 100.
[0097] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0098] The memory may include internal memory 122. Internal memory 122 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0099] Random access memory (RAM) can be directly read and written by the processor 110. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data. Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into RAM for direct read and write by the processor 110.
[0100] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions.
[0101] The wireless communication function of electronic device 100 can be implemented through antenna 01, antenna 02, mobile communication module 150, wireless communication module 160, modem processor, and baseband processor.
[0102] Antennas 01 and 02 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 01 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0103] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 01, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 01. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0104] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device or displays an image or video through the display screen 193. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0105] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including Wireless Local Area Networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. The wireless communication module 160 receives electromagnetic waves via antenna 02, modulates and filters the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, modulate and amplify them, and then convert them into electromagnetic waves for radiation via antenna 02.
[0106] In some embodiments, the antenna 01 of the electronic device 100 is coupled to the mobile communication module 110, and the antenna 02 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology.
[0107] Electronic device 100 implements display functions through a GPU, a display screen 193, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 193 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0108] The display screen 193 is used to display images, videos, etc. The display screen 193 includes a display panel. In some embodiments, the electronic device may include one or N display screens 193, where N is a positive integer greater than 1.
[0109] In this embodiment of the application, the display screen 193 can be used to display a first image.
[0110] Electronic device 100 can perform shooting functions through ISP, camera 192, video codec, GPU, display 193 and application processor.
[0111] The ISP (Image Signal Processor) is used to process data fed back from the camera 192. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set within the camera 192.
[0112] Camera 192 is used to capture still images or videos. An object passes through the lens to generate an optical image that is projected onto a photosensitive element. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP (Internet Service Provider) for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP (Digital Signal Processor) for processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the electronic device 100 may include one or N cameras 192, where N is a positive integer greater than 1.
[0113] Touch sensor 180A, also known as a "touch device," can be located on display screen 193. The touch sensor 180A and display screen 193 together form a touchscreen, also known as a "touchscreen." Touch sensor 180A is used to detect touch operations applied to or near it. The touch sensor can then transmit the detected touch operation to the application processor to determine the type of touch event.
[0114] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100.
[0115] The 180C barometric pressure sensor is used to measure barometric pressure.
[0116] The geomagnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the geomagnetic sensor 180D to detect the opening and closing of the flip cover.
[0117] The 180E accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices and applied to applications such as screen orientation switching and pedometers.
[0118] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.
[0119] The proximity light sensor 180G may include, for example, a light-emitting diode and a light detector, such as a photodiode.
[0120] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This embodiment of the invention uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.
[0121] Figure 6 is a schematic diagram of the software structure of the electronic device provided in an embodiment of this application.
[0122] As shown in Figure 6, the layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0123] The application layer can include a series of application packages.
[0124] As shown in Figure 6, the application package may include applications such as camera, gallery, calendar, call, map, navigation, music, video, and SMS.
[0125] In this embodiment of the application, the application package may include a game application.
[0126] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0127] As shown in Figure 6, the application framework layer may include a window manager, an input manager, a sensor manager, a phone manager, a resource manager, a notification manager, etc.
[0128] In this embodiment, the application framework layer also includes an input service, InputFlinger. The electronic device 100 can obtain user input events through InputFlinger.
[0129] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0130] The input manager can be used to listen for user input events, such as clicks and swipes performed by the user's finger on the display screen 193 of the electronic device 100. By listening to input events, the electronic device 100 can determine whether it is in use.
[0131] The sensor manager is used to monitor data returned by various sensors in the electronic device 100, such as motion sensor data, proximity sensor data, and temperature sensor data. Using the data returned by each sensor, the electronic device 100 can determine whether it is shaking or whether the display screen 193 is obstructed.
[0132] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).
[0133] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0134] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0135] InputFlinger manages events from input devices (such as touchscreens, keyboards, mice, etc.) and distributes these events to applications in the system to ensure the accuracy of user input. Specifically, InputFlinger receives signals from input devices, such as those from a touchscreen, converts these signals into standard input events, and then distributes these input events to currently active applications to ensure that user actions are accurately responded to by the applications.
[0136] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0137] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0138] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0139] System libraries can include multiple functional modules. For example: Surface Manager, Media Libraries, 3D graphics processing libraries (e.g., OpenGL, EGL), 2D graphics engines (e.g., SGL), game engines, scene recognition modules, optimization strategy modules, etc.
[0140] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0141] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0142] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0143] A 2D graphics engine is a graphics engine for 2D drawing.
[0144] In this embodiment of the application, the game engine can be a service for managing and distributing game content updates. Such an engine can be integrated into the game client or run as a standalone service to perform services such as content distribution, incremental updates, version control, compatibility checks, user notifications, and automatic updates.
[0145] In this embodiment, the scene recognition module can identify different elements in the game environment, such as terrain, objects, and characters. The scene recognition module can also detect the occurrence of specific events, such as a game character entering a certain area or collecting a specific item. Thus, the scene recognition module can trigger corresponding game logic based on the occurrence of specific events. The scene recognition module can also recognize user interactions with the game environment to provide a dynamic gaming experience. Furthermore, the scene recognition module can identify specific scenes to trigger visual and animation effects, enhancing the visual expressiveness of the game environment. The scene recognition method shown in this embodiment can be implemented based on the scene recognition module.
[0146] It should be noted that in this embodiment, the scene recognition module can be set as an independent module in the system library, or it can be integrated into the game engine, or it can be set in other software layers according to the actual situation. This embodiment does not limit the specific setting method of the scene recognition module. This embodiment only illustrates the scene recognition module as an independent module in the system library.
[0147] In this embodiment, the optimization strategy module is a module that ensures smooth system performance and efficient resource utilization. This module can implement optimization strategies such as frame rate optimization, asset optimization, material optimization, code optimization, graphics rendering pipeline optimization, and physics calculation optimization. For example, a frame rate optimization strategy could increase the CPU / GPU frequency, and a physics calculation optimization strategy could record the moment when rare equipment or rare characters are drawn.
[0148] It should be noted that in this embodiment, the optimization strategy module can be set as an independent module in the system library, or it can be integrated into the game engine, or it can be set in other software layers according to the actual situation. This embodiment does not limit the specific setting method of the optimization strategy module. This embodiment only illustrates the optimization strategy module as an independent module in the system library.
[0149] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0150] The kernel layer can include the base kernel, extended kernel, hardware abstraction layer (HAL), and driver layer.
[0151] The basic kernel can implement core functions such as memory management, task management, inter-process communication, and interrupt management based on the memory (Mem) interface, task (Task) interface, inter-process communication (IPC) interface, and interrupt interface.
[0152] The Hardware Abstraction Layer (HAL) can abstract the hardware operation interface, encapsulating the underlying driver interface into a unified Application Programming Interface (API). This simplifies the complexity of hardware operations for applications.
[0153] It should be noted that the embodiments of this application are only illustrated by taking the HAL layer as part of the kernel layer. In fact, the HAL layer can also be independent of the kernel layer, and the embodiments of this application do not limit this.
[0154] The driver layer includes at least the display driver, camera driver, audio driver, sensor driver, and graphics driver.
[0155] In this embodiment of the application, the graphics driver can be used to manage and control the GPU in the computer system and the associated graphics functions and display devices.
[0156] In this embodiment, the display driver can manage the output of a display device (such as display screen 193) and transfer image frames rendered in the GPU from the GPU to the display device. The display driver can cooperate with the graphics driver to control the output of image frames.
[0157] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0158] The following embodiments of this application are all illustrated by taking the scene recognition process corresponding to a gacha game application as an example.
[0159] Figure 7 is a schematic diagram of the first process of the scene recognition method provided in the embodiments of this application.
[0160] Figure 8 is the first interactive schematic diagram of the scene recognition method provided in the embodiments of this application.
[0161] As shown in Figures 7 and 8, in some embodiments, the method includes the following steps S11-S17.
[0162] In step S11, the electronic device 100 sets a preset image and presets first boundary coordinates based on the preset image.
[0163] Among them, the preset image is a pre-stored image sample used to define the basic features of the target scene.
[0164] Figure 9 is a schematic diagram of a preset image provided in an embodiment of this application.
[0165] As shown in Figure 9, the preset image 11 can be a frame of a card-drawing image displayed by the electronic device 100 in a card-drawing scenario. The preset image 11 has significant features of the card-drawing scenario. For example, the preset image 11 can display the "orange meteor" that appears when a rare character is drawn. If the electronic device 100 displays this preset image 11, it can be determined that the electronic device 100 is in a card-drawing scenario.
[0166] The electronic device 100 needs to preset the first boundary coordinates based on the preset image 11 in order to distinguish the key area and non-key area corresponding to the card-drawing scene in the subsequent scene recognition process. The key area may be, for example, the area where the "orange meteor" is located.
[0167] The following is a detailed explanation of how the electronic device 100 presets the first boundary coordinates.
[0168] In one implementation, the electronic device 100 automatically presets the first boundary coordinates.
[0169] Specifically, the electronic device 100 compares at least one first interface image with a preset image 11 using a clustering analysis method to determine the first boundary coordinates. This method automatically presets the first boundary coordinates, reduces manual intervention, and enhances the robustness of the electronic device 100.
[0170] The clustering analysis method used in this application includes the Bisecting K-Means clustering algorithm. The Bisecting K-Means clustering algorithm is a clustering algorithm that combines the advantages of hierarchical clustering and the K-Means algorithm. This algorithm is based on the idea of progressively dividing clusters into two groups, achieving efficient and accurate clustering results.
[0171] It should be noted that other types of clustering analysis methods can also be used in the embodiments of this application to determine the first boundary coordinates, such as the traditional K-Means clustering algorithm, K-Means++ clustering algorithm, hierarchical clustering algorithm, etc. The embodiments of this application do not limit the specific type of clustering analysis method.
[0172] Figure 10 is a schematic diagram of a first interface image provided in an embodiment of this application for comparison with a preset image.
[0173] As shown in Figures 9 and 10, the first interface image 12 can be an image displayed in a card-drawing scenario.
[0174] An example is given by comparing the first interface image 12 with the preset image 11 in three frames.
[0175] The first frame, first interface image 12, can be the image corresponding to the start of the card-drawing animation, at which time the "meteor" is displayed before it is launched. The second frame, first interface image 12, can be the image corresponding to the middle process of the card-drawing animation, at which time the "meteor" is launched. The third frame, first interface image 12, can be another image corresponding to the middle process of the card-drawing animation, at which time the color state of the "meteor" is displayed.
[0176] The electronic device 100 can compare these three first interface images 12 with the preset image 11 in sequence to determine the first boundary coordinates.
[0177] It should be noted that in actual applications, the electronic device 100 needs to compare more frames of the first interface image 12 with the preset image 11. This application embodiment only uses the comparison process of three frames of the first interface image 12 with the preset image 11 as an example. For the comparison process of other first interface images 12 with the preset image 11, please refer to this embodiment. This application will not elaborate on this.
[0178] Based on the above scenario, the electronic device 100 determines the first boundary coordinates by including the following steps S111-S116.
[0179] In step S111, the electronic device 100 divides each first interface image 12 into multiple first interface image blocks.
[0180] In step S112, the electronic device 100 divides the preset image 11 into multiple second interface image blocks.
[0181] In each first interface image 12, each first interface image block corresponds to the position of a second interface image block.
[0182] Figure 11 is a schematic diagram of the first interface image block and the second interface image block provided in the embodiments of this application.
[0183] As shown in Figures 9 to 11, based on the aforementioned example, the electronic device 100 can divide the first frame first interface image 12 into a first interface image block 121 with i rows and j columns, where i equals 5 and j equals 8.
[0184] Accordingly, the electronic device 100 can divide the second frame first interface image 12 into first interface image blocks 121 in rows i and columns j, for example, i equals 5 and j equals 8.
[0185] Accordingly, the electronic device 100 can divide the third frame first interface image 12 into a first interface image block 121 with i rows and j columns, for example, i equals 5 and j equals 8.
[0186] Accordingly, the electronic device 100 can divide the preset image 11 into a second interface image block 111 with i rows and j columns. For example, i equals 5 and j equals 8.
[0187] Thus, each first interface image block 121 in each first interface image 12 corresponds to the position of a second interface image block 111.
[0188] In step S113, the electronic device 100 obtains the average interface distance between each first interface image block 121 and each second interface image block 111 in each first interface image 12, so as to determine the target average distance corresponding to each first interface image 12.
[0189] In this embodiment, the average interface distance refers to the average of the Euclidean distances between corresponding pixels in the first interface image block 121 and the second interface image block 111. Euclidean distance is the straight-line distance between two points in a multidimensional space. The grayscale value or color value of each pixel in an image block can be considered as one dimension; thus, different dimensions can be considered as a multidimensional space.
[0190] The process by which electronic device 100 obtains the average distance to the interface is further explained below.
[0191] In one implementation, step S113 includes steps S1131-S1135.
[0192] In step S1131, the electronic device 100 divides multiple first interface image blocks 121 into first cluster image blocks and second cluster image blocks in each first interface image 12.
[0193] The first cluster of image blocks includes a portion of the first interface image block 121, and the second cluster of image blocks includes another portion of the first interface image block 121.
[0194] The process of dividing the electronic device 100 into the first cluster of image blocks and the second cluster of image blocks is explained in detail below.
[0195] Specifically, step S1131 includes steps S1131a-S1131f.
[0196] In step S1131a, the electronic device 100 arbitrarily determines two cluster centers among a plurality of first interface image blocks 121.
[0197] Figure 12 is a schematic diagram of the method for determining the first cluster image block and the second cluster image block provided in the embodiments of this application.
[0198] As shown in Figure 12, based on the aforementioned example, the first interface image 12 of the first frame is used for further illustrative explanation.
[0199] In the first frame of the first interface image 12, the electronic device 100 can determine the first interface image block 121 in region A1 (i.e., row 1, column 1) as the first cluster center, and determine the first interface image block 121 in region A2 (i.e., row 5, column 8) as the second cluster center.
[0200] In step S1131b, the electronic device 100 determines the target Euclidean distance between each first interface image block 121 and each cluster center.
[0201] In one implementation, the electronic device 100 can extract feature values for each first interface image block 121 and each cluster center to determine the target Euclidean distance between each first interface image block 121 and each cluster center based on the feature values.
[0202] The feature values include, for example, at least one of pixel values, texture feature values, and color values.
[0203] A pixel value refers to the brightness or color intensity value of each pixel in an image. In a grayscale image, each pixel corresponds to one brightness value. In a color image, each pixel corresponds to three color intensity values, which correspond to the red, green, and blue (RGB) color channels, respectively. In some cases, each pixel may also correspond to a fourth color intensity value, which corresponds to the alpha channel.
[0204] Texture features are a set of values used to describe the texture patterns in an image. Texture can be regular features (such as a brick wall) or irregular features (such as grass), and these features can be extracted based on methods such as gray-level co-occurrence matrix, local binary pattern, and wavelet transform. Texture features can be statistical measures corresponding to the extracted texture features.
[0205] Color values are numerical values used to describe the color information of pixels in an image. Different color spaces represent color values differently; these color spaces include, for example, RGB, Hue Saturation Value (HSV), and Lab color spaces. In the RGB color space, color values can be represented based on (R, G, B) color coordinates.
[0206] Based on different types of feature values, electronic device 100 can obtain different target Euclidean distances. In this way, different target Euclidean distances form different benchmarks, and different benchmarks can determine the subsequent clustering methods.
[0207] Specifically, if the feature value is only the brightness value among pixel values, the different clusters obtained in the subsequent clustering process can include a cluster with higher brightness and a cluster with lower brightness. If the feature value is only the texture feature value, the different clusters obtained in the subsequent clustering process can include a cluster with regular texture and a cluster with irregular texture. If the feature value is only the color value, the different clusters obtained in the subsequent clustering process can include a cluster with one similar color (such as orange meteors) and a cluster with another similar color (such as blue sky).
[0208] To ensure the accuracy of clustering, this embodiment of the application can select multiple feature values, assign different weights to the feature values, and normalize these feature values to obtain normalized feature values. In this way, the electronic device 100 can determine the target Euclidean distance based on the normalized feature values, making the subsequent clustering results more accurate.
[0209] In one implementation, the electronic device 100 can assign the largest weight to the color value among different types of feature values so that the subsequent clustering results can distinguish between key and non-key regions with color differences.
[0210] For ease of explanation, the following embodiments of this application are all illustrated by illustrative examples where the feature value is only a color value. In fact, the embodiments of this application do not limit the specific method of determining the target Euclidean distance based on the feature value.
[0211] Further, as shown in Figure 12, based on the aforementioned example, the first interface image 12 of the first frame is used for illustrative purposes.
[0212] In the first frame of the first interface image 12, the electronic device 100 first determines the first target Euclidean distance between the first interface image block 121 in the first row and the first column and the first cluster center (i.e., the first interface image block 121 in region A1).
[0213] When the Euclidean distance of the first target is determined based on the color value, the Euclidean distance of the first target can be obtained based on the following formula.
[0214] Formula 1:
[0215] Where, d ij-1 R is the Euclidean distance to the first target corresponding to the first interface image block 121 in the i-th row and j-th column. ij G represents the color value of the first interface image block 121 in the i-th row and j-th column on the red channel. ij B represents the color value of the first interface image block 121 in the i-th row and j-th column on the green channel. ijR represents the color value of the first interface image block 121 in the i-th row and j-th column on the blue channel. p1 G represents the color value of the first cluster center on the red channel. p1 B is the color value corresponding to the first cluster center in the green channel. p1 Let i be the color value of the first cluster center in the blue channel. i is a positive integer greater than 0, and j is a positive integer greater than 0.
[0216] Then, with i=1 and j=1, the first target Euclidean distance d between the first interface image patch 121 in the first row and first column and the first cluster center is... 11-1 It is 0.
[0217] Next, the electronic device 100 determines the second target Euclidean distance between the first interface image block 121 in the first row and first column and the second cluster center (i.e., the first interface image block 121 in region A2).
[0218] When the Euclidean distance of the second target is determined based on the color value, the Euclidean distance of the second target can be obtained based on the following formula.
[0219] Formula 2:
[0220] Where, d ij-2 R is the Euclidean distance to the second target corresponding to the first interface image patch 121 in the i-th row and j-th column. ij G represents the color value of the first interface image block 121 in the i-th row and j-th column on the red channel. ij B represents the color value of the first interface image block 121 in the i-th row and j-th column on the green channel. ij R represents the color value of the first interface image block 121 in the i-th row and j-th column on the blue channel. p2 G represents the color value of the second cluster center on the red channel. p2 B represents the color value of the second cluster center in the green channel. p2 Let i be the color value corresponding to the second cluster center in the blue channel. i is a positive integer greater than 0, and j is a positive integer greater than 0.
[0221] Then, with i=1 and j=1, the second target Euclidean distance d between the first interface image patch 121 in the first row and first column and the second cluster center is... 11-2 For example, it could be 30 (this value is for illustrative purposes only).
[0222] Based on a similar approach, the electronic device 100 can determine the Euclidean distance between two targets corresponding to each first interface image block 121. The specific process by which the electronic device 100 determines the Euclidean distances corresponding to other first interface image blocks 121 will not be elaborated upon in this embodiment.
[0223] In step S1131c, the electronic device 100 assigns each first interface image block 121 to a cluster center with the smallest target Euclidean distance based on the two target Euclidean distances corresponding to each first interface image block 121, so as to obtain a first cluster target image block corresponding to one cluster center and a second cluster target image block corresponding to the other cluster center.
[0224] Further, as shown in Figure 12, based on the aforementioned example, the first interface image 12 of the first frame is used for illustrative purposes.
[0225] In the first frame of the first interface image 12, for the first interface image block 121 in the first row and first column, the electronic device 100 can obtain the first target Euclidean distance d corresponding to the first interface image block 121 in the first row and first column. 11-1 The value is 0, and the corresponding Euclidean distance d to the second target is... 11-2 The value is 30 (this value is for illustrative purposes only). Therefore, the first interface image block 121 in the first row and first column has the minimum target Euclidean distance with the first cluster center, and the electronic device 100 can assign the first interface image block 121 in the first row and first column to the first cluster center.
[0226] Based on a similar approach, the electronic device 100 can complete the allocation of all the first interface image blocks 121.
[0227] The result after allocation is, for example, the first cluster of target image blocks 13 includes: the first interface image block 121 in the first row and first column, the first interface image block 121 in the first row and second column, the first interface image block 121 in the first row and third column, the first interface image block 121 in the first row and fourth column, ..., the first interface image block 121 in the fifth row and fourth column.
[0228] The second cluster of target image blocks 14 includes the first interface image block 121 in the 1st row and 5th column, the first interface image block 121 in the 1st row and 6th column, the first interface image block 121 in the 1st row and 7th column, the first interface image block 121 in the 1st row and 8th column, ..., the first interface image block 121 in the 5th row and 8th column.
[0229] The specific image blocks included in the first cluster of target image blocks 13 and the second cluster of target image blocks 14 are shown in Figure 11, and are not listed in detail in this embodiment.
[0230] In step S1131d, the electronic device 100 determines the first average error between the first cluster target image block 13 and its corresponding cluster center, and determines the second average error between the second cluster target image block 14 and its corresponding cluster center.
[0231] The electronic device 100 can obtain the first sum of all first target Euclidean distances between each first interface image block 121 in the first cluster target image block 13 and the first cluster center, and then divide the first sum by the number of first interface image blocks 121 in the first cluster target image block 13 to obtain the first average error.
[0232] Accordingly, the electronic device 100 can obtain the second sum of all second target Euclidean distances between each first interface image block 121 in the second cluster target image block 14 and the second cluster center, and then divide the second sum by the number of first interface image blocks 121 in the second cluster target image block 14 to obtain the second average error.
[0233] Further, as shown in Figure 12, based on the aforementioned example, the first interface image 12 of the first frame is used for illustrative purposes.
[0234] In the first frame of the first interface image 12, if the electronic device 100 obtains a first target Euclidean distance of 0 corresponding to the first interface image block 121 in the first row and first column of the first cluster of target image blocks 13, and obtains a first target Euclidean distance of 10 corresponding to the first interface image block 121 in the first row and second column (this value is for illustrative purposes only), and obtains a first target Euclidean distance of 20 corresponding to the first interface image block 121 in the first row and third column (this value is for illustrative purposes only), ..., and obtains a first target Euclidean distance of 20 corresponding to the first interface image block 121 in the fifth row and fourth column (this value is for illustrative purposes only), then the first sum can be 0+10+20+...+20. Since the number of first interface image blocks 121 in the first cluster of target image blocks 13 is 16, the first average error is (0+10+20+...+20) / 16, and the first average error is, for example, 25 (this value is for illustrative purposes only).
[0235] Accordingly, if the electronic device 100 acquires a first target Euclidean distance of 10 for the first interface image block 121 in the first row and fifth column of the second cluster of target image blocks 14, and acquires a first target Euclidean distance of 30 for the first interface image block 121 in the first row and sixth column, and acquires a first target Euclidean distance of 20 for the first interface image block 121 in the first row and seventh column, ..., and acquires a second target Euclidean distance of 0 for the first interface image block 121 in the fifth row and eighth column, then the second sum can be 10+30+20+...+0. Since the number of first interface image blocks 121 in the second cluster of target image blocks 14 is 16, the second average error is (10+30+20+...+0) / 16, and the second average error is, for example, 40 (this value is for illustrative purposes only).
[0236] It should be noted that the values shown above are only used to describe the logical relationship between the data and do not represent the actual values. The embodiments of this application do not limit the specific value of each data.
[0237] In step S1131e, if the first average error and the second average error are different, the electronic device 100 updates the two cluster centers and determines the target Euclidean distance between each first interface image block 121 and each updated cluster center again, thereby obtaining the updated first average error and the updated second average error.
[0238] Further, as shown in Figure 12, based on the aforementioned example, when the first average error is 25 and the second average error is 40, the first average error and the second average error are different. Therefore, the electronic device 100 needs to update the first cluster center of region A1 and update the second cluster center of region A2, and again determine the target Euclidean distance between each first interface image block 121 and each updated cluster center in the aforementioned manner to obtain the updated first average error and the updated second average error.
[0239] The electronic device 100 may repeat step S1131e multiple times until the updated first average error and the updated second average error are the same.
[0240] Further as shown in Figure 12, based on the aforementioned example, if the electronic device 100 updates the first interface image block 121 with the first cluster center as region A3, and updates the first interface image block 121 with the second cluster center as region A4, and obtains that the first average error corresponding to the updated first cluster center is 5 (this value is only for illustrative purposes), and the second average error corresponding to the updated second cluster center is 5, the electronic device 100 can determine that the updated first average error and the updated second average error are the same, and execute step S1131f.
[0241] It should be noted that in this embodiment, the first average error and the second average error can be exactly the same value or not exactly the same value; that is, there can be a small difference between the first average error and the second average error. This embodiment only illustrates the case where the first average error and the second average error are exactly the same. If the difference between the first average error and the second average error can be used to characterize the cluster centers corresponding to the first cluster of target image blocks and the second cluster of target image blocks respectively, and these cluster centers no longer change significantly, then the first average error and the second average error can be considered to be approximately the same. In this case, the electronic device 100 can also execute step S1131f. In step S1131f, if the updated first average error and the updated second average error are the same, the electronic device 100 determines the updated first cluster of target image blocks as the first cluster of image blocks 15, and determines the updated second cluster of target image blocks as the second cluster of image blocks 16.
[0242] Further as shown in Figure 12, based on the aforementioned example, the electronic device 100 can determine the first cluster target image block corresponding to the first cluster center (i.e., the first interface image block 121 in region A3) as the first cluster image block 15, and determine the second cluster target image block corresponding to the second cluster center (i.e., the first interface image block 121 in region A4) as the second cluster image block 16.
[0243] In this way, the electronic device 100 can distinguish between key and non-key regions in the first interface image 12 based on the first cluster of image blocks 15 and the second cluster of image blocks 16.
[0244] Figure 13 is a schematic diagram of the first cluster image block and the second cluster image block in different first interface images provided in the embodiments of this application.
[0245] As shown in Figure 13, based on the aforementioned example, in the first frame first interface image 12, the first cluster of image blocks 15 can be used to distinguish the key area corresponding to the "unlaunched meteor", and the second cluster of image blocks 16 can be used to distinguish the non-key area corresponding to the "blue sky".
[0246] In the first interface image 12 of the second frame, the first cluster of image blocks 15 can be used to distinguish the key area corresponding to the "shooting meteor", and the second cluster of image blocks 16 can be used to distinguish the non-key area corresponding to the "blue sky".
[0247] In the first interface image 12 of the third frame, the first cluster of image blocks 15 can be used to distinguish the key area corresponding to the "meteor with color", and the second cluster of image blocks 16 can be used to distinguish the non-key area corresponding to the "blue sky".
[0248] In step S1132, the electronic device 100 determines the first pixel Euclidean distance between each first pixel and each second pixel based on each first pixel in each first interface image block 121 in the first cluster image block 15 and the second pixel corresponding to the first pixel in the second interface image block 111.
[0249] In one implementation, the electronic device 100 sets a first color value (R1, G1, B1) for each first pixel based on RGB coordinates, and sets a second color value (R2, G2, B2) for each second pixel based on RGB coordinates. Further, the electronic device 100 obtains a first difference between a first red component R1 in the first color value and a second red component R2 in the second color value; a second difference between a first green component G1 in the first color value and a second green component G2 in the second color value; and a third difference between a first blue component B1 in the first color value and a second blue component B2 in the second color value. Further, the electronic device 100 determines the square root of the sum of the squares of the first, second, and third differences as the Euclidean distance of the first pixel.
[0250] Figure 14 is a schematic diagram of a scenario in which the electronic device provided in the embodiment of this application determines the Euclidean distance between the first pixel and the second pixel.
[0251] As shown in Figure 14, based on the aforementioned example, an example is given of a scenario in which the electronic device 100 determines the Euclidean distance of the first pixel based on the first frame first interface image 12 and the preset image 11.
[0252] The electronic device 100 can acquire the first color value corresponding to each first pixel in the first interface image block 121 in the second row and third column, the first interface image block 121 in the second row and fourth column, the first interface image block 121 in the third row and third column, and the first interface image block 121 in the third row and fourth column, and acquire the second color value corresponding to each second pixel in the second interface image block 111 in the second row and third column, the second interface image block 111 in the second row and fourth column, the second interface image block 111 in the third row and third column, and the second interface image block 111 in the third row and fourth column, to determine the Euclidean distance of the first pixel.
[0253] In step S1133, the electronic device 100 determines the second pixel Euclidean distance between each third pixel and each fourth pixel based on each third pixel in each first interface image block 121 in the second cluster image block 16 and the fourth pixel corresponding to the third pixel in the second interface image block 111.
[0254] In one implementation, the electronic device 100 sets the third color value corresponding to each third pixel to (R3, G3, B3) based on RGB coordinates, and sets the fourth color value corresponding to each fourth pixel to (R4, G4, B4) based on RGB coordinates. Further, the electronic device 100 obtains a fourth difference between the third red component R3 in the third color value and the fourth red component R4 in the fourth color value, a fifth difference between the third green component G3 in the third color value and the fourth green component G4 in the fourth color value, and a sixth difference between the third blue component B3 in the third color value and the fourth blue component B4 in the fourth color value. Further, the electronic device 100 determines the square root of the sum of the squares of the fourth, fifth, and sixth differences as the Euclidean distance of the second pixel.
[0255] Further, as shown in Figure 14, based on the aforementioned example, an example is given of a scenario in which the electronic device 100 determines the Euclidean distance of the second pixel based on the first frame first interface image 12 and the preset image 11.
[0256] The electronic device 100 can acquire the third color value corresponding to each third pixel in the first interface image block 121 other than the first interface image block 121 in the second row and third column, the first interface image block 121 in the second row and fourth column, the first interface image block 121 in the third row and third column, and the first interface image block 121 in the third row and fourth column, and acquire the fourth color value corresponding to each fourth pixel in the second interface image block 111 other than the second interface image block 111 in the second row and third column, the second interface image block 111 in the second row and fourth column, the second interface image block 111 in the third row and third column, and the second interface image block 111 in the third row and fourth column, in order to determine the Euclidean distance of the second pixel.
[0257] The above steps S1132-S1133 can be calculated based on the following formula:
[0258] Formula 3:
[0259] Where d(I0(i,j), Ik(i,j)) is the pixel Euclidean distance. k is equal to 1 or 2, i is a positive integer greater than 0, and j is a positive integer greater than 0.
[0260] I 0(i,j) I represents the color value corresponding to the pixel in the i-th row and j-th column of the second interface image block 111. 0(i,j) equals (R) 0(i,j) G 0(i,j) B 0(i,j) ).
[0261] I k(i,j) I represents the color value corresponding to the pixel in the i-th row and j-th column of the first interface image block 121. k(i,j) equals (R) k(i,j) G k(i,j) B k(i,j) ).
[0262] When k equals 1, I 1(i,j) This is the first color value corresponding to the first pixel in the i-th row and j-th column of the first interface image block 121 in the first cluster image block 15. Correspondingly, it corresponds to I... 1(i,j) Corresponding I 0(i,j) It is the second color value corresponding to the second pixel in the i-th row and j-th column of the second interface image block 111.
[0263] When k equals 2, I 2(i,j) This refers to the three color values corresponding to the third pixel in the i-th row and j-th column of the first interface image block 121 within the second cluster image block 16. Correspondingly, this is related to I... 2(i,j) Corresponding I 0(i,j) It is the fourth color value corresponding to the fourth pixel in the i-th row and j-th column of the second interface image block 111.
[0264] It should be noted that the first pixel, second pixel, third pixel, and fourth pixel in the embodiments of this application are only used to distinguish pixels corresponding to different clusters, and do not indicate that the pixels have different numbers or types. Similarly, the first color value, second color value, third color value, and fourth color value in the embodiments of this application are only used to distinguish the color values of pixels corresponding to different clusters, and do not indicate that the types of color values are different.
[0265] In other words, d(I0(i,j), I1(i,j)) is the Euclidean distance of the first pixel, and d(I0(i,j), I2(i,j)) is the Euclidean distance of the second pixel.
[0266] In step S1134, the electronic device 100 determines the average distance between each first interface image block 121 in the first cluster image block 15 and its corresponding second interface image block 111 based on the Euclidean distance of all first pixels corresponding to each first interface image block 121 in the first cluster image block 15.
[0267] In one implementation, the electronic device 100 divides the sum of the Euclidean distances of all first pixels in each first interface image block 121 in the first cluster of image blocks 15 by the total number of all first pixels in each first interface image block 121 to obtain the average distance of first pixels between each first interface image block 121 and its corresponding second interface image block 111.
[0268] In step S1135, the electronic device 100 determines the average second pixel distance between each first interface image block 121 in the second cluster image block 16 and its corresponding second interface image block 111 based on all the second pixel Euclidean distances corresponding to each first interface image block 121 in the second cluster image block 16.
[0269] In one implementation, the electronic device 100 divides the sum of the Euclidean distances of all second pixels in each first interface image block 121 in the second cluster of image blocks 16 by the total number of all third pixels in each first interface image block 121 to obtain the average distance of the second pixels between each first interface image block 121 and its corresponding second interface image block 111.
[0270] In this embodiment of the application, the average distance of the interface includes the average distance of the first pixel and the average distance of the second pixel.
[0271] The above steps S1134-S1135 can be calculated based on the following formula:
[0272] Formula 4:
[0273] Among them, D (I0(i,j),Ik(i,j)) d(I0(i,j), Ik(i,j)) is the average distance between the interfaces. d(I0(i,j), Ik(i,j)) is the pixel Euclidean distance. M is the total number of pixels in the first interface image 12. H is the image height of the first interface image 12. W is the image width of the first interface image 12.
[0274] When k equals 1, D (I0(i,j),I1(i,j)) The average distance of the first pixel.
[0275] When k equals 2, D (I0(i,j),I2(i,j)) The average distance of the second pixel.
[0276] In this way, the electronic device 100 can obtain the average interface distance between each first interface image block 121 and each corresponding second interface image block 111 in the first interface image 12.
[0277] Furthermore, the electronic device 100 needs to determine the target average distance as the average distance of the top N largest values of the interface in each first interface image 12. Here, N is a positive integer greater than 0.
[0278] In one implementation, the electronic device 100 compares the average distance of each interface (i.e., the average distance of the first pixel, or the average distance of the second pixel) with a preset first interface average distance threshold in each first interface image 12, so as to determine the average distance of N interfaces that is greater than the first interface average distance threshold as the target average distance.
[0279] It should be noted that the specific value of the average distance threshold on the first interface can be set according to the actual situation, and this application embodiment does not limit it.
[0280] The larger the average interface distance, the greater the difference between the first interface image block 121 and its corresponding second interface image block 111; the smaller the average interface distance, the smaller the difference between the first interface image block 121 and its corresponding second interface image block 111.
[0281] The reason why the electronic device 100 needs to determine the average distance to the target is to determine the interface areas occupied by the first interface image block 121 and the second interface image block 111, which are corresponding in position and have significant differences, in the first interface image 12. In this way, the electronic device 100 can select a suitable interface area as the key area of the target in the subsequent process.
[0282] The following describes the interface area corresponding to the average distance to the target.
[0283] Figure 15 is the first schematic diagram of the interface area corresponding to the target average distance provided in the embodiments of this application.
[0284] As shown in Figure 15, further based on the aforementioned example, in the first frame of the first interface image 12, the electronic device 100 can determine that the first interface image block 121 covered by the first key region 17 corresponding to the "unlaunched meteor" has a large average interface distance between the second interface image blocks 111 corresponding to these first interface image blocks 121. Furthermore, in the preset image 11, the electronic device 100 can determine that the second interface image block 111 covered by the second key region 18 corresponding to the "orange meteor" has a large average interface distance between the first interface image blocks 121 corresponding to these second interface image blocks 111.
[0285] Therefore, in the first frame of the first interface image 12, the first key region 17 and the second key region 18, mapped onto the first interface image 12, can form a first overlapping region. The first overlapping region is the interface region corresponding to the average distance of the target.
[0286] In other words, the electronic device 100 can determine the average distance of the interface corresponding to the first interface image block 121 covered by the first overlapping area as the target average distance.
[0287] In this way, the electronic device 100 can determine the average distance of the interfaces corresponding to the first interface image block 121 in the second row and third column, the first interface image block 121 in the second row and fourth column, the first interface image block 121 in the second row and fifth column, the first interface image block 121 in the third row and third column, the first interface image block 121 in the third row and fourth column, and the first interface image block 121 in the third row and fifth column as the target average distance in the first interface image 12 of the first frame.
[0288] Figure 16 is a second schematic diagram of the interface area corresponding to the target average distance provided in the embodiments of this application.
[0289] As shown in Figure 16, further based on the aforementioned example and in a similar manner, the electronic device 100 can determine the average interface distances corresponding to the first interface image block 121 in the second frame and fourth column, the first interface image block 121 in the second row and fifth column, the first interface image block 121 in the third row and fourth column, the first interface image block 121 in the third row and fifth column, the first interface image block 121 in the third row and sixth column, the first interface image block 121 in the fourth row and fourth column, the first interface image block 121 in the fourth row and fifth column, and the first interface image block 121 in the fourth row and sixth column as the target average distance.
[0290] Figure 17 is a third schematic diagram of the interface area corresponding to the target average distance provided in the embodiments of this application.
[0291] As shown in Figure 17, based on the aforementioned example and in a similar manner, the electronic device 100 can determine the average interface distances corresponding to the first interface image block 121 in the second row and fourth column, the first interface image block 121 in the second row and fifth column, the first interface image block 121 in the third frame and the first interface image 12 as the target average distance.
[0292] Electronic device 100 can also determine the target average distance from the interface average distance using other methods.
[0293] In one implementation, the electronic device 100 sorts the average distances of the first pixel and the second pixel from largest to smallest, and determines the average distance of the top N largest values as the target average distance. This application does not limit the specific method by which the electronic device 100 determines the target average distance.
[0294] In step S114, in each first interface image, the electronic device 100 acquires the first number of first interface image blocks 121 corresponding to the average distance to the target.
[0295] Further as shown in Figures 15 to 17, based on the aforementioned example, the electronic device 100 can obtain a first number of six first interface image blocks 121 corresponding to the average distance to the target in the first frame of the first interface image 12.
[0296] In the second frame of the first interface image 12, the electronic device 100 can obtain a first number of 8 first interface image blocks 121 corresponding to the average distance of the target.
[0297] In the second frame of the first interface image 12, the electronic device 100 can obtain a first number of four first interface image blocks 121 corresponding to the average distance of the target.
[0298] In one implementation, the electronic device 100 may perform steps S111-S114 once for each first interface image 12. Thus, after the electronic device 100 obtains the first quantity corresponding to the current first interface image 12, it can compare the current first quantity with a preset first quantity threshold. If the current first quantity is greater than the preset first quantity threshold, the electronic device 100 can discard the current first interface image. If the current first quantity is less than the preset first quantity threshold, the electronic device 100 can confirm that it has obtained the first quantity with the minimum value, and then proceed to step S115.
[0299] It should be noted that the specific value of the first quantity threshold can be set according to the actual situation, and this application embodiment does not limit it.
[0300] In one implementation, the electronic device 100 may further perform steps S111-S114 for a plurality of first interface images 12 to determine a first quantity with the minimum value from a plurality of first quantities, and then perform step S115.
[0301] The embodiments of this application do not restrict the specific execution order of each sub-step in step S11.
[0302] In step S115, the electronic device 100 determines the first interface image 12 corresponding to the first quantity with the minimum value as the target image.
[0303] Further as shown in Figures 15 to 17, based on the aforementioned embodiments, the electronic device 100 can determine the first interface image 12 (i.e., the third frame first interface image 12) corresponding to the first quantity with the minimum value as the target image.
[0304] In step S116, in the target image, the electronic device 100 determines the boundary coordinates of the area where the first interface image block 121 corresponding to the average distance of the target is located as the first boundary coordinates.
[0305] Further, as shown in Figure 17, based on the aforementioned example, in the third frame of the first interface image 12, the electronic device 100 can determine the boundary coordinates (x1, y1), (x2, y1), (x1, y2), (x2, y2) of the area where the first interface image block 121 corresponding to the average distance of the target is located as the first boundary coordinates.
[0306] In one implementation, the electronic device 100 may preset first boundary coordinates based on user input.
[0307] Users can directly input the values of (x1, y1), (x2, y1), (x1, y2), and (x2, y2) to enable the electronic device 100 to preset the first boundary coordinates.
[0308] Users can also manually adjust the first boundary coordinates automatically preset by the electronic device 100 after the electronic device 100 has automatically preset them.
[0309] The embodiments of this application do not limit the specific method by which the electronic device 100 presets the first boundary coordinates.
[0310] As further shown in Figure 6, the electronic device 100 can execute step S11 within an independent scene recognition module in the system library.
[0311] The electronic device 100 can also execute step S11 within the game engine in the system library, based on the scene recognition module integrated into the game engine.
[0312] The embodiments of this application do not limit the specific manner in which the electronic device 100 performs step S11.
[0313] In step S12, the electronic device 100 displays the first interface.
[0314] The first interface can be the running interface of the first application.
[0315] The first application can be a game application or a non-game application. This application only uses a gacha game application as an example to illustrate the first application, and the specific type of the first application is not limited in the embodiments of this application.
[0316] When the first application is a gacha game application, the first interface can be the gacha game's running interface. Since the gacha game's running interface includes not only multiple frames of game images in gacha scenes but also multiple frames of game images in non-gacha scenes, the electronic device 100 needs to perform scene recognition on such game images and then adopt corresponding optimization strategies.
[0317] It should be noted that although the electronic device 100 can perform the scene recognition method shown in this application on any of the images displayed on its interface, not all images need to be scene recognized. For example, when the electronic device 100 displays the interface image of the main screen interface, the interface image of the main screen interface does not need to be scene recognized.
[0318] Therefore, this application embodiment only illustrates the scene recognition process of the first image in the first interface.
[0319] In one implementation, step S12 includes step S121.
[0320] Step S121: The electronic device acquires at least one frame of the first image in the first interface.
[0321] The first image is, for example, the game image currently displayed on the first interface.
[0322] Figure 18 is a schematic diagram of a scenario in which an electronic device acquires a first image according to an embodiment of this application.
[0323] As shown in Figure 18, the user can perform a first operation on the first icon 21 corresponding to the game application on the main screen interface 20 of the electronic device 100. In response to the first operation on the first icon 21, the electronic device 100 runs the game application, displays the game running interface, and acquires the first image 22 in the game running interface.
[0324] In the above process, the electronic device 100 can determine whether to enter the game application based on whether the package name of the foreground application matches the preset list of game package names, thereby triggering the acquisition of the first image 22.
[0325] The electronic device 100 can also take a screenshot of the display window based on the detection of changes in the display window of the foreground interface, obtain the window feature information corresponding to the screenshot, and match the window feature information with a preset game window feature library to determine whether to enter the game application, thereby triggering the acquisition of the first image 22.
[0326] Electronic device 100 can also obtain detailed information of the foreground application through system-level monitoring tools, and determine whether the foreground application is a game application based on the detailed information, so as to determine whether to enter the game application, and then trigger the acquisition of the first image 22.
[0327] This application embodiment does not limit the specific method by which the electronic device 100 triggers the acquisition of the first image 22.
[0328] Figure 19 is a schematic diagram of the interactive process of an electronic device acquiring a first image according to an embodiment of this application.
[0329] As shown in Figure 19, after the electronic device 100 enters the game application, the game application can convert user input events into game logic instructions. Furthermore, the game application can call the game engine's API to send the game logic instructions to the game engine. Further, the game engine updates the game state according to the game logic instructions and generates rendering instructions corresponding to the updated game state. Further, the game engine calls the graphics API to send the rendering instructions to the GPU, causing the GPU to render the first image 22 and display it.
[0330] The electronic device 100 can acquire the first image 22 from the GPU via a graphics API, either based on a game engine or a standalone scene recognition module.
[0331] The electronic device 100 can also capture the current screen using a system-level screenshot tool and send the current screen image to the scene recognition module so that the scene recognition module can obtain the first image 22.
[0332] This application does not limit the specific method by which the electronic device 100 acquires the first image 22.
[0333] The following embodiments of this application are all illustrated by way of electronic device 100 acquiring first image 22 based on scene recognition module.
[0334] It should also be noted that steps S13-S16 in subsequent embodiments of this application can all be executed by the electronic device 100 based on the scene recognition module, and will not be described in detail in subsequent embodiments of this application.
[0335] In step S13, the electronic device 100 determines the first region and the second region in the first interface.
[0336] Specifically, the electronic device 100 can determine the first region and the second region based on each frame of the first image 22 in the first interface.
[0337] In one implementation, step S13 includes steps S131-S133.
[0338] In step S131, the electronic device 100 divides the first image 22 into multiple first image blocks.
[0339] The way in which the electronic device 100 divides the first image 22 into multiple first image blocks can be referred to the way in which the electronic device 100 divides the first interface image block 121 in the first interface image 12 in the embodiment corresponding to FIG10 above. This application embodiment will not elaborate on this.
[0340] In step S132, the electronic device 100 determines the first region enclosed by the first boundary coordinates based on the preset first boundary coordinates, wherein the first region is composed of a portion of a plurality of first image blocks.
[0341] Figure 20 is a schematic diagram of a scenario where an electronic device, according to an embodiment of this application, determines a first region.
[0342] As shown in Figure 20, further based on the aforementioned example, the preset first boundary coordinates in the electronic device 100 can be (x1, y1), (x2, y1), (x1, y2), and (x2, y2). These first boundary coordinates can enclose the interface areas occupied by the first image block 221 in the second row and fourth column, the first image block 221 in the second row and fifth column, and the first image block 221 in the third row and fourth column in the first image 22 to form the first region 23.
[0343] In step S133, the electronic device 100 obtains the difference region between the first image 22 and the first region 23 based on the second boundary coordinates and the first boundary coordinates corresponding to the first image 22, so as to determine the difference region as the second region, wherein the second region is composed of another part of a plurality of first image blocks 221.
[0344] Figure 21 is a schematic diagram of a scenario where an electronic device, according to an embodiment of this application, determines a second region.
[0345] As shown in Figure 21, further based on the aforementioned example, after obtaining the first boundary coordinates, the electronic device 100 can determine the difference region (i.e., the shaded region in Figure 20) between the first image 22 and the first region 23 based on the second boundary coordinates (x3, y3), (x4, y3), (x3, y4), (x4, y4) and the first boundary coordinates (x1, y1), (x2, y1), (x1, y2), (x2, y2). In this way, the electronic device 100 can determine the difference region as the second region 24.
[0346] The second boundary coordinates can be built into the electronic device 100 and correspond to the display screen specifications or current display mode of the electronic device 100. The values of the second boundary coordinates differ when the electronic device 100 uses different display screen specifications or different current display modes. For example, when the electronic device 100 is a foldable device, the second boundary coordinates correspond to different values in the folded and unfolded states. The second boundary coordinates also correspond to different values when the electronic device 100 is in a 16:9 display mode or a 4:3 display mode. In other words, the electronic device 100 needs to select the corresponding second boundary coordinates based on its current state. This application embodiment does not limit the specific values of the second boundary coordinates used by the electronic device 100.
[0347] In one implementation, after determining the first region 23, the electronic device 100 may also select the second region 24 based on a neural network model to determine the second region. This application embodiment does not limit the specific method by which the electronic device 100 determines the second region.
[0348] In this way, after determining the first region 23 and the second region 24, the electronic device 100 can achieve hierarchical scene recognition for these two different regions.
[0349] In step S14, the electronic device 100 performs scene recognition on the image corresponding to the second region 24 to obtain the first scene type corresponding to the second region 24.
[0350] In one implementation, step S14 includes steps S141-S144.
[0351] In step S141, the electronic device 100 divides the preset image 11 into multiple second image blocks.
[0352] It should be noted that, since the electronic device 100 has divided the preset image 11 into multiple second interface image blocks 111 in the foregoing embodiments, in step S141, the electronic device 100 can re-divide the preset image 11 into multiple second image blocks, or reuse the second interface image blocks 111 as second image blocks. The specific implementation of step S141 is not limited in this application embodiment.
[0353] This application embodiment is only used as an example to illustrate how the electronic device 100 redivides the preset image 11 into multiple second image blocks. The method by which the electronic device 100 redivides the preset image 11 into multiple second image blocks can be referred to in the embodiment corresponding to FIG10 above, where the electronic device 100 divides the preset image 11 into second interface image blocks 111; this application embodiment will not elaborate on that.
[0354] After the electronic device 100 divides the preset image 11 into multiple second image blocks, each second image block corresponds to the position of a first image block 221.
[0355] In step S142, the electronic device 100 acquires each first image block 221 in the second region 24, and the first average distance between the second image blocks corresponding to each first image block 221.
[0356] The specific method by which the electronic device 100 acquires each first image block 221 in the second region 24 and the first average distance between the second image blocks corresponding to each first image block 221 can be referred to in the foregoing embodiments, which describes the method by which the electronic device 100 acquires the average interface distance between each first interface image block 121 in the first interface image 12 and each corresponding second interface image block 111. This application embodiment will not elaborate on this method.
[0357] After obtaining the first average distance corresponding to each first image block 221 in the second region 24, the electronic device 100 can determine the minimum value D(I01(i,j), Ismall(i,j)) among these first average distances.
[0358] Step S143: If the minimum value in the first average distance is greater than or equal to the first average distance threshold, then the electronic device 100 determines that the first scene type corresponding to the second area 24 is a non-critical scene.
[0359] The following embodiments of this application illustrate the scene recognition process corresponding to the first image 22 in a non-card drawing scenario and the first image 22 in a card drawing scenario, respectively.
[0360] Figure 22 is a first schematic diagram of scene recognition of a second region by an electronic device provided in an embodiment of this application.
[0361] As shown in Figure 22, in the first frame of the first image 22 in a non-card-drawing scenario, the electronic device 100 can obtain the first average distance between each first image block 221 in the second region 24 and its corresponding second image block 112, and determine the minimum value among these first average distances. However, since each first image block 221 in the second region 24 of the first frame of the first image 22 has a large difference from its corresponding second image block 112, even if the first average distance corresponding to the first image block 221 is the smallest, it will still be greater than the first average distance threshold.
[0362] For example, in the first frame, the first image 22, the first average distance between the first image block 221 in region B1 and its corresponding second image block 112 is the smallest; however, there are still significant differences between the two image blocks (e.g., completely different colors).
[0363] Therefore, the electronic device 100 can determine that the first scene type corresponding to the second area 24 is a non-critical scene, for example, a non-card drawing scene.
[0364] Step S144: If the minimum value in the first average distance is less than the first average distance threshold, the electronic device 100 determines that the first scene type corresponding to the second area 24 is a critical scene.
[0365] Figure 23 is a second schematic diagram of scene recognition of a second region by an electronic device provided in an embodiment of this application.
[0366] Figure 24 is a third schematic diagram of the electronic device provided in this application performing scene recognition in the second region.
[0367] Figure 25 is a fourth schematic diagram of the electronic device provided in this application performing scene recognition in the second region.
[0368] As shown in Figures 23, 24 and 25, in the second frame first image 22, the third frame first image 22 and the fourth frame first image 22 in the card drawing scenario, the electronic device 100 can obtain the first average distance corresponding to each first image block 221 in the second region 24 for each frame first image 22, and determine the minimum value among these first average distances.
[0369] Since the first image 22 in the second, third, and fourth frames are all in a card-drawing scenario, these first images 22 have a certain scene similarity. Each frame of the first image 22 contains a first image block 221 that has a small difference from the corresponding second image block 112 (such as the image block corresponding to "blue sky"). Therefore, in these first images 22 in the card-drawing scenario, the minimum value of the first average distance obtained by the electronic device 100 can be less than the first average distance threshold.
[0370] For example, in the first image 22 of the second frame, the first image block 221 in region B2 has the smallest first average distance with its corresponding second image block 112, and the two image blocks are similar. Therefore, the electronic device 100 can determine that the first scene type corresponding to the second region 24 in the first image 22 of the second frame is a key scene, for example, a card-drawing scene.
[0371] For example, in the first image 22 of the third frame, the first image block 221 in region B3 has the smallest first average distance with its corresponding second image block 112, and the two image blocks are similar. Therefore, the electronic device 100 can determine that the first scene type corresponding to the second region 24 in the first image 22 of the third frame is a key scene, for example, a card-drawing scene.
[0372] For example, in the first image 22 of the fourth frame, the first image block 221 in region B4 has the smallest first average distance with its corresponding second image block 112, and the two image blocks are similar. Therefore, the electronic device 100 can determine that the first scene type corresponding to the second region 24 in the first image 22 of the fourth frame is a key scene, for example, a card-drawing scene.
[0373] Since the first average distance between any image block between two images with significant differences is relatively large, the range of the first average distance threshold can be relatively large, so that the scene recognition process in the second region has a high tolerance. The specific value of the first average distance threshold can be set according to the actual situation, and this application embodiment does not limit it.
[0374] In step S15, when the first scene type is a non-critical scene, the electronic device 100 performs scene recognition on the image corresponding to the second region 24 based on a preset first frequency, and obtains the first scene type corresponding to the second region 24 again.
[0375] In other words, when the electronic device 100 determines that the current scene is a non-critical scene, it does not need to perform scene recognition on the first region 23, but only needs to continuously perform scene recognition on the second region 24 in the other first images 22.
[0376] Since the scene recognition process of the electronic device 100 in the second region 24 is a long-term and wide-range detection process, and in non-critical scenarios, the electronic device 100 does not need to perform further operations (e.g., record the moment when rare equipment or rare characters are drawn, and perform related operations of corresponding optimization strategies), the electronic device 100 can perform scene recognition in the second region 24 at a lower first frequency to reduce the power consumption of the electronic device 100. The specific value of the first frequency is not limited in this embodiment.
[0377] It should also be noted that the method used by the electronic device 100 to determine the first scene type by threshold comparison based on the first average distance during the scene recognition process of the second region 24 is a low-consumption, low-cost, and high-tolerance algorithm. The use of this algorithm in the second region 24 by the electronic device 100 further reduces the power consumption of the electronic device 100 and improves the system performance of the electronic device 100.
[0378] The electronic device 100 can continuously perform scene recognition on the second region 24 in other first images 22 until the first scene type is determined to be a key scene.
[0379] In step S16, when the first scene type is a key scene, the electronic device 100 performs scene recognition on the image corresponding to the first region 23 based on a preset second frequency to determine whether the first interface is in the target scene.
[0380] The second frequency is greater than the first frequency.
[0381] Specifically, during the continuous scene recognition process of the electronic device 100 on the second region 24, if the electronic device 100 determines that the first scene type corresponding to a certain frame of the first image 22 (such as the second frame of the first image 22) is a key scene, then the electronic device 100 performs scene recognition at the next level, that is, performs scene recognition on the first region 23 of the current first image 22 (such as the second frame of the first image 22). Since this level of scene recognition process has higher accuracy requirements, the electronic device 100 needs to perform scene recognition on the first region 23 based on a higher second frequency.
[0382] It should be noted that, since the electronic device 100 has a high frame rate in the game scene, if the sub-device 100 determines that the first scene type corresponding to the first image 22 of the second frame is a key scene, the next level of scene recognition performed by the electronic device 100 may not be targeted at the first region 23 of the current first image 22, but rather at the first region 23 of the subsequent first image 22 (such as the first image 22 of the third frame).
[0383] This application embodiment only illustrates the process of scene recognition performed by the electronic device 100 on the first region 23 of the current first image 22. Furthermore, the electronic device 100 may also employ more precise algorithms to further improve the accuracy of the scene recognition process.
[0384] It should be noted that although the more accurate algorithm used by the electronic device 100 will generate higher power consumption, the scene recognition process performed in the first region 23 has a small area, which saves the power consumption of the electronic device 100 to a certain extent. In this way, the power saving due to the smaller area offsets the power consumption generated by the complex algorithm, so that the solution shown in this application can save the system performance of the electronic device 100 while ensuring the accuracy of scene recognition.
[0385] The following is an exemplary description of the algorithm used by the electronic device 100 in the first region 23.
[0386] In one implementation, step S16 includes steps S161-S164.
[0387] In step S161, the electronic device 100 acquires the first image feature in the first region image corresponding to the first region 23.
[0388] Specifically, the first region image can be a portion of the currently acquired first image 22, or a portion of the first image 22 displayed subsequently. This application embodiment does not limit the specific type of the first region image.
[0389] It should be noted that, as shown in Figure 18, the time point at which the scene recognition module currently acquires the first image 22 may be later than the time point at which the GPU sends the first image 22 for display. Therefore, when the GPU sends the next frame of the first image 22 for display, the scene recognition module may reacquire the next frame of the first image 22, or it may not acquire the next frame of the first image 22, thus resulting in different images of the first region.
[0390] In one implementation, the first image feature includes at least one of Scale Invariant Feature Transform (SIFT) features, Oriented Fast and Rotated Brief (ORB) features.
[0391] SIFT features are image features obtained by detecting key points in an image and extracting descriptors for those key points.
[0392] The following is an exemplary description of the process by which electronic device 100 acquires SIFT features.
[0393] In one implementation, the electronic device 100 can construct a Gaussian pyramid in the first region 23 to form different scale spaces based on the Gaussian pyramid. Further, the electronic device 100 can search for extreme points in the different scale spaces to obtain the corresponding keypoints in the image. Further, the electronic device 100 can remove edge response points from the keypoints to improve their stability. Further, the electronic device 100 can assign an orientation to each keypoint, making each keypoint rotationally invariant. Further, the electronic device 100 can acquire image data around each keypoint to generate a 128-dimensional descriptor vector, which can be used to describe the local features of the keypoint, i.e., SIFT features.
[0394] ORB features are image features obtained based on an improved feature detection and description method. These features combine the advantages of Fast keypoints and Brief descriptors, and introduce orientation information, exhibiting good rotation invariance.
[0395] The following is an exemplary description of the process by which electronic device 100 acquires ORB features.
[0396] In one implementation, the electronic device 100 can use the Fast algorithm to detect corner points in an image and filter out stable keypoints through non-maximum suppression. Furthermore, the electronic device 100 assigns an orientation to each keypoint, making each keypoint rotationally invariant. Further, the electronic device 100 applies the Brief algorithm to the keypoints to generate a binary descriptor vector, which can be used to describe the grayscale features of the keypoints, i.e., ORB features.
[0397] In this embodiment of the application, the electronic device 100 can acquire any image feature as the first image feature, and the specific type of the first image feature is not limited in this embodiment of the application.
[0398] In step S162, the electronic device 100 acquires the second image feature in the target region 25 in the preset image 11, wherein the target region 25 corresponds to the position of the first region 23.
[0399] In one implementation, the second image feature includes at least one of SIFT features and ORB features.
[0400] In order to facilitate the comparison of image features, the first image feature and the second image feature need to have the same feature type.
[0401] The method by which the electronic device 100 acquires the second image features can be referred to the method of acquiring the first image in the foregoing embodiments, and will not be described in detail here.
[0402] In step S163, the electronic device 100 acquires the first Euclidean distance between the first image feature and the second image feature.
[0403] An example is given where both the first and second image features are ORB features.
[0404] In one implementation, the electronic device 100 can convert both the first image feature and the second image feature into floating-point numbers, and then calculate the first Euclidean distance based on the floating-point numbers.
[0405] The above step S163 can be calculated based on the following formula:
[0406] Formula 4:
[0407] Where, d (f1,f2) f1 is the first Euclidean distance. f2 is the first floating-point number corresponding to the first image feature and f1 is the second floating-point number corresponding to the second image feature.
[0408] It should be noted that, compared with the method of determining Euclidean distance based on color values in the previous embodiments, the method of determining Euclidean distance based on floating-point numbers corresponding to image features used in this embodiment has higher robustness and higher accuracy. Therefore, it is suitable for scene recognition of the first region 23.
[0409] In one implementation, if the first Euclidean distance is greater than the first Euclidean distance threshold, the electronic device 100 determines that the scene in the first region 23 has not been successfully recognized, skips the current first image 22, and performs scene recognition on the first region 23 of the next frame of the first image 22.
[0410] In step S164, if the first Euclidean distance is less than the preset first Euclidean distance threshold, the electronic device 100 determines that the scene recognition of the first region 23 is successful, and then determines that the first interface is in the target scene.
[0411] The following embodiments of this application illustrate the scene recognition process corresponding to multiple frames of the first image 22 in a card-drawing scenario.
[0412] Figure 26 is a schematic diagram of scene recognition of a first region by an electronic device provided in an embodiment of this application.
[0413] As shown in Figures 23 to 26, in the second frame first image 22, the third frame first image 22, and the fourth frame first image 22 in the card drawing scenario, the electronic device 100 can obtain the first image feature in the first region 23 and the first Euclidean distance between the second image feature in the target region 25 for each frame first image 22.
[0414] Since the first Euclidean distance can more accurately measure the difference between the first region 23 and the target region 25, the first region 23 in the second frame first image 22 and the first region 23 in the third frame first image 22 have a large difference relative to the target region 25, while the first region 23 in the fourth frame first image 22 has a small difference relative to the target region 25. Therefore, the electronic device 100 can obtain a first Euclidean distance less than the first Euclidean distance threshold in the fourth frame first image 22.
[0415] In this way, the electronic device 100 can determine that the scene recognition is successful in the first image 22 of the fourth frame, and then determine that the first image 22 of the fourth frame is the target scene of extracting rare equipment or rare characters.
[0416] It should be noted that the value range of the first Euclidean distance threshold can be relatively small so that the scene recognition process of the first region 23 has high accuracy. The specific value of the first Euclidean distance threshold can be set according to the actual situation, and this application embodiment does not limit it.
[0417] In step S17, if it is determined that the first interface is in the target scene, the electronic device 100 executes the optimization strategy corresponding to the target scene.
[0418] In one implementation, after the scene recognition module determines that the first interface is in the target scene, it sends a strategy execution instruction to the optimization strategy module. The optimization strategy module then performs corresponding operations based on the strategy execution instruction. For example, it might perform an operation to record the moment when a rare character or rare equipment is extracted.
[0419] The following explains the optimization strategies executed by the optimization strategy module in team battle scenarios.
[0420] As further illustrated in Figure 19, if the scene recognition module determines that the first interface is in a team battle scene, the scene recognition module can send a target strategy execution instruction to the optimization strategy module. This target strategy execution instruction instructs the GPU to increase its frequency. The optimization strategy module can then parse the target strategy execution instruction and convert it into a format readable by the GPU. Upon receiving the parsed target strategy execution instruction, the GPU increases its own frequency. Based on this process, the electronic device 100 does not need to perform scene recognition on the entire area of the first image 22 in the first interface. Instead, it creates different layers by dividing the first interface into a first region 23 and a second region 24. Scene recognition is then achieved on portions of the first image 22 in the first interface based on these different layers. This ensures the accuracy of scene recognition while avoiding unnecessary scene recognition operations, reducing power consumption, and improving system performance.
[0421] Figure 27 is the second flowchart of the scene recognition method provided in the embodiments of this application.
[0422] As shown in Figure 27, the scene recognition method provided in this application embodiment may include the following steps S21-S24.
[0423] Step S21: The electronic device displays the first interface.
[0424] Step S22: The electronic device determines a first region and a second region in the first interface, wherein the second region includes at least a portion of the region in the first interface other than the first region, and the area of the second region is greater than or equal to the area of the first region.
[0425] Step S23: The electronic device performs scene recognition on the image corresponding to the second region to obtain the first scene type corresponding to the second region.
[0426] In step S24, if the first scene type is a key scene, the electronic device performs scene recognition on the image corresponding to the first area to determine whether the first interface is in the target scene.
[0427] The scene recognition method provided in this application does not require scene recognition for the entire area of the interface image. Instead, it forms different levels by dividing the interface image into regions, and then performs scene recognition on a portion of the interface image of the electronic device based on the different levels. In this way, while ensuring the accuracy of scene recognition, unnecessary scene recognition operations are avoided, the power consumption of the electronic device is reduced, and the system performance of the electronic device is improved.
[0428] The foregoing primarily describes the solutions provided in the embodiments of this application from the perspective of electronic devices. It is understood that, in order to achieve the aforementioned functions, the electronic device includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the scene recognition method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed through hardware or by software-driven hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0429] This application embodiment can divide the above-described electronic device into functional modules or functional units according to the above method examples. For example, each function can be divided into its own functional modules or functional units, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module or functional unit. The module or unit division in this application embodiment is illustrative and represents only one logical functional division; other division methods may be used in actual implementation.
[0430] Figure 28 is a schematic diagram of the scene recognition device provided in an embodiment of this application.
[0431] As shown in Figure 28, the scene recognition device 200 provided in this application embodiment can be applied to the electronic device in the above embodiment. The scene recognition device 200 may include: a display screen 201, a memory 202, a processor 203, and a communication module 204. These devices can be connected via one or more communication buses 205. The display screen 201 may include a display panel 2011 and a touch sensor 2012. The display panel 2011 is used to display images, and the touch sensor 2012 can transmit detected touch operations to the application processor to determine the touch event type. The display panel 2011 provides visual output related to the touch operation. The processor 203 may include one or more processing units, such as an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. Different processing units may be independent devices or integrated into one or more processors. The memory 202 is coupled to the processor 203 and is used to store various software programs and / or computer instructions. The memory 202 may include volatile memory and / or non-volatile memory. When the processor executes computer instructions, the electronic device can perform the various functions or steps performed in the above method embodiments.
[0432] Figure 29 is a schematic diagram of the structure of a scene recognition device provided in another embodiment of this application.
[0433] As shown in Figure 29, in some other embodiments, this application also provides a scene recognition device 300, which can be applied to the electronic device in the above embodiments. The scene recognition device 300 includes a display module 301, a determination module 302, a first recognition module 303, and a second recognition module 304. The display module 301 is used to display a first interface. The determination module 302 is used to determine a first region and a second region in the first interface, wherein the second region includes at least a portion of the first interface excluding the first region, and the area of the second region is greater than or equal to the area of the first region. The first recognition module 303 is used to perform scene recognition on the image corresponding to the second region to obtain a first scene type corresponding to the second region. The second recognition module 304 is used to perform scene recognition on the image corresponding to the first region when the first scene type is a key scene, to determine whether the first interface is in a target scene.
[0434] The scene recognition device provided in this application does not require scene recognition for the entire area of the interface image. Instead, it forms different levels by dividing the interface image into regions, and performs scene recognition on a portion of the interface image of the electronic device based on the different levels. In this way, while ensuring the accuracy of scene recognition, unnecessary scene recognition operations are avoided, the power consumption of the electronic device is reduced, and the system performance of the electronic device is improved.
[0435] This application also provides an electronic device, which includes a processor and a memory coupled together. The memory stores program instructions, and when the program instructions are executed by the processor, the processor performs the various functions or steps as described in the above method embodiments.
[0436] Figure 30 is a schematic diagram of the chip system provided in an embodiment of this application.
[0437] As shown in Figure 30, this application embodiment also provides a chip system 400, such as a SoC, which includes at least one processor 401 and at least one interface circuit 402. The processor 401 and the interface circuit 402 can be interconnected via lines. For example, the interface circuit 402 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 402 can be used to send signals to other devices (e.g., the processor 401 or the touchscreen of an electronic device). Exemplarily, the interface circuit 402 can read instructions stored in the memory and send the instructions to the processor 401. When the instructions are executed by the processor 401, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, which are not specifically limited in this application embodiment.
[0438] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.
[0439] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform various functions or steps performed by the electronic device in the above method embodiments.
[0440] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0441] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0442] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0443] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0444] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0445] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A scene recognition method, characterized in that, Applied to electronic devices, including: Display a first interface; determine a first region and a second region in the first interface, wherein the second region includes at least a portion of the region in the first interface other than the first region, and the area of the second region is greater than or equal to the area of the first region; Scene recognition is performed on the image corresponding to the second region to obtain the first scene type corresponding to the second region; When the first scene type is a critical scene, scene recognition is performed on the image corresponding to the first region to determine whether the first interface is in the target scene.
2. The scene recognition method according to claim 1, characterized in that, After performing scene recognition on the image corresponding to the second region to obtain the first scene type corresponding to the second region, the method further includes: When the first scene type is a non-critical scene, scene recognition is performed only on the image corresponding to the second region based on a preset first frequency in the first region and the second region, and scene recognition is not performed on the image corresponding to the first region.
3. The scene recognition method according to claim 2, characterized in that, When the first scene type is a critical scene, the step of performing scene recognition on the image corresponding to the first region includes: Scene recognition is performed on the image corresponding to the first region based on a preset second frequency, wherein the second frequency is greater than the first frequency.
4. The scene recognition method according to claim 2 or 3, characterized in that, The first interface for display includes: In response to a first operation on the game application, the game application is run and the first interface is displayed, which is the game running interface.
5. The scene recognition method according to claim 2 or 3, characterized in that, Determining the first region and the second region in the first interface includes: Based on the first image displayed in the first interface, the first image is divided into multiple first image blocks; Based on preset first boundary coordinates, the first region enclosed by the first boundary coordinates is determined, wherein the first region is composed of a portion of a plurality of first image blocks; Based on the second boundary coordinates corresponding to the first image and the first boundary coordinates, the difference region between the first image and the first region is obtained, and the difference region is determined as the second region, wherein the second region is composed of another part of a plurality of the first image blocks.
6. The scene recognition method according to claim 5, characterized in that, The step of performing scene recognition on the image corresponding to the second region to obtain the first scene type corresponding to the second region includes: The preset image is divided into multiple second image blocks, wherein each second image block corresponds to the position of a first image block; Obtain each of the first image blocks in the second region, and the first average distance between the second image blocks corresponding to each of the first image blocks; If the minimum value of the first average distance is less than the first average distance threshold, then the first scene type corresponding to the second region is determined to be a critical scene; If the minimum value of the first average distance is greater than or equal to the first average distance threshold, then the first scene type corresponding to the second region is determined to be a non-critical scene.
7. The scene recognition method according to claim 6, characterized in that, When the first scene type is a critical scene, the step of performing scene recognition on the image corresponding to the first region to determine whether the first interface is in the target scene includes: Obtain a first image feature of the first region image corresponding to the first region, wherein the first region image is a part of the first image, or is a part of a second image, and the second image is used to display after the first image is displayed on the first interface; In the preset image, a second image feature in the target region is obtained, wherein the target region corresponds to the position of the first region; Obtain the first Euclidean distance between the first image feature and the second image feature; If the first Euclidean distance is less than the preset first Euclidean distance threshold, then the first region scene recognition is determined to be successful, and the first interface is determined to be in the target scene.
8. The scene recognition method according to claim 7, characterized in that, The first image feature includes at least one of scale-invariant feature transform (SIFT) features, orientation fast and rotation-invariant ORB features; The second image feature includes at least one of scale-invariant feature transform (SIFT) features, orientation fast and rotation-invariant ORB features; The first image feature and the second image feature have the same feature type.
9. The scene recognition method according to claim 5, characterized in that, Before displaying the first interface, the method further includes: Set a preset image, and preset the first boundary coordinates based on the preset image.
10. The scene recognition method according to claim 9, characterized in that, The step of presetting the first boundary coordinates based on the preset image includes: At least one third image is compared with the preset image using cluster analysis to determine the coordinates of the first boundary, wherein the third image is used for display in the key scene.
11. The scene recognition method according to claim 10, characterized in that, The step of comparing at least one third image with the preset image based on cluster analysis to determine the coordinates of the first boundary includes: In each of the third images, the third image is divided into a plurality of third image blocks, and the preset image is divided into a plurality of second image blocks, wherein each of the third image blocks in each of the third images corresponds to the position of a second image block; In each of the third images, a second average distance is obtained between each third image block and each second image block to determine the target average distance corresponding to each third image, wherein the target average distance includes the second average distance with the largest value among the first N values, and N is a positive integer greater than 0; In each of the third images, a first number of third image blocks corresponding to the average distance to the target is obtained; The third image corresponding to the first quantity with the minimum value is determined as the target image; In the target image, the boundary coordinates of the region where the third image block corresponding to the average distance of the target is located are determined as the first boundary coordinates.
12. The scene recognition method according to claim 11, characterized in that, In each of the third images, obtaining the second average distance between each third image block and each of the second image blocks includes: In each of the third images, the plurality of third image blocks are divided into a first cluster of image blocks and a second cluster of image blocks, wherein the first cluster of image blocks includes a portion of the third image blocks, and the second cluster of image blocks includes another portion of the third image blocks; Based on each first pixel in each of the third image blocks in the first cluster of image blocks, and the second pixel in the second image block corresponding to the first pixel, a second Euclidean distance between each first pixel and each second pixel is determined; Based on each third pixel in each of the third image blocks in the second cluster of image blocks, and the fourth pixel in the second image block corresponding to the third pixel, a third Euclidean distance between each third pixel and each fourth pixel is determined; Based on all the second Euclidean distances corresponding to each of the third image blocks in the first cluster of image blocks, determine each of the third image blocks in the first cluster of image blocks, and the third average distance between each of the third image blocks and its corresponding second image blocks; Based on all the third Euclidean distances corresponding to each of the third image blocks in the second cluster of image blocks, a fourth average distance is determined between each of the third image blocks in the second cluster of image blocks and its corresponding second image blocks, wherein the second average distance includes the third average distance and the fourth average distance.
13. The scene recognition method according to claim 12, characterized in that, In each of the third images, dividing the plurality of third image blocks into a first cluster of image blocks and a second cluster of image blocks includes: Two cluster centers are arbitrarily determined among the plurality of the third image blocks; Determine the target Euclidean distance between each of the third image patches and each of the cluster centers; Based on the two target Euclidean distances corresponding to each third image block, each third image block is assigned to a cluster center with the smallest target Euclidean distance to obtain a first cluster target image block corresponding to one of the cluster centers, and a second cluster target image block corresponding to the other cluster center; Determine a first average error between the first cluster of target image patches and their corresponding cluster centers, and determine a second average error between the second cluster of target image patches and their corresponding cluster centers; If the first average error and the second average error are different, then the two cluster centers are updated, and the target Euclidean distance between each third image patch and each updated cluster center is determined again, thereby obtaining the updated first average error and the updated second average error. If the updated first average error and the updated second average error are the same, then the updated first cluster target image block is determined as the first cluster image block, and the updated second cluster target image block is determined as the second cluster image block.
14. The scene recognition method according to claim 12, characterized in that, The step of determining the second Euclidean distance between each first pixel and each second pixel based on each first pixel in each of the third image blocks in the first cluster of image blocks, and the second pixel in the second image block corresponding to the first pixel, includes: The first color value corresponding to each first pixel is set to (R1, G1, B1) based on the red, green and blue RGB coordinates, and the second color value corresponding to each second pixel is set to (R2, G2, B2) based on the RGB coordinates. Obtain the first difference between the first red component R1 in the first color value and the second red component R2 in the second color value, the second difference between the first green component G1 in the first color value and the second green component G2 in the second color value, and the third difference between the first blue component B1 in the first color value and the second blue component B2 in the second color value. The square root of the sum of the squares of the first difference, the second difference, and the third difference is determined as the second Euclidean distance.
15. The scene recognition method according to claim 14, characterized in that, The step of determining the third Euclidean distance between each third pixel and each fourth pixel based on each third pixel in each of the third image blocks in the second cluster of image blocks, and the fourth pixel in the second image block corresponding to the third pixel, includes: Based on the RGB coordinates, the third color value corresponding to each third pixel is set to (R3, G3, B3), and the fourth color value corresponding to each fourth pixel is set to (R4, G4, B4) based on the RGB coordinates. Obtain the fourth difference between the third red component R3 in the third color value and the fourth red component R4 in the fourth color value, the fifth difference between the third green component G3 in the third color value and the fourth green component G4 in the fourth color value, and the sixth difference between the third blue component B3 in the third color value and the fourth blue component B4 in the fourth color value. The square root of the sum of the squares of the fourth difference, the fifth difference, and the sixth difference is determined as the third Euclidean distance.
16. The scene recognition method according to claim 12, characterized in that, The step of determining each third image block in the first cluster of image blocks, and the third average distance between it and its corresponding second image block, based on all the second Euclidean distances corresponding to each third image block in the first cluster of image blocks, includes: In each of the third image blocks in the first cluster of image blocks, the sum of all the second Euclidean distances in each third image block is divided by the total number of all the first pixels in each first image block to obtain the third average distance between each third image block and its corresponding second image block.
17. The scene recognition method according to claim 12, characterized in that, The step of determining each third image block in the second cluster of image blocks, and the fourth average distance between each third image block and its corresponding second image block, based on all the third Euclidean distances corresponding to each third image block in the second cluster of image blocks, includes: In each of the third image blocks in the second cluster of image blocks, the sum of all the third Euclidean distances in each third image block is divided by the total number of all the third pixels in each third image block to obtain the fourth average distance between each third image block and its corresponding second image block.
18. The scene recognition method according to claim 11, characterized in that, Determining the average target distance corresponding to each of the third images includes: In each of the third images, each of the second average distances is compared with a preset second average distance threshold, so that N second average distances greater than the second average distance threshold are determined as the target average distance.
19. The scene recognition method according to claim 12, characterized in that, Determining the average target distance corresponding to each of the third images includes: The third and fourth average distances are sorted from largest to smallest, and the average distances of the top N largest values are determined as the target average distance.
20. An electronic device, characterized in that, include: The electronic device includes a processor and a memory, the memory storing program instructions that, when executed by the processor, cause the electronic device to perform the scene recognition method as described in any one of claims 1-19.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on an electronic device, cause the electronic device to perform the scene recognition method as described in any one of claims 1-19.