Scene recognition method and apparatus

By combining images acquired from multiple cameras at multiple azimuth angles with azimuth information, and using a two-layer neural network model for scene recognition, the problem of limited field of view and shooting angle of a single camera is solved, thereby improving the accuracy of recognition and reducing false judgments.

CN115049909BActive Publication Date: 2026-04-14HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-25
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing image-based scene recognition methods are not robust in complex scenes, have a small field of view and limited shooting angles, resulting in low recognition accuracy and easy misjudgment.

Method used

Multiple cameras are used to acquire images from different azimuth angles. The azimuth angle information of each image is combined to perform scene recognition through a two-layer neural network model, and local and global features are extracted using a competition mechanism.

Benefits of technology

It improves the accuracy of scene recognition, reduces false judgments, and achieves more comprehensive scene information acquisition and accurate recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115049909B_ABST
    Figure CN115049909B_ABST
Patent Text Reader

Abstract

The application relates to a scene recognition method and device. The method comprises the following steps: a terminal device collects images under a same scene from multiple azimuth angles through multiple cameras, wherein the azimuth angle of each camera when collecting the images is the azimuth angle corresponding to the images, and the azimuth angle is the included angle between the direction vector of each camera when collecting the images and the gravity unit vector; and the terminal device recognizes the same scene according to the images and the azimuth angles corresponding to the images, and obtains a scene recognition result. According to the embodiment of the application, the scene is recognized in combination with multiple images and the azimuth angle corresponding to each image, the accuracy of the scene recognition based on the images can be improved due to the obtained more comprehensive scene information, the problem that the scene recognition range and the shooting angle of a single camera are limited is solved, and more accurate recognition is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a scene recognition method and apparatus. Background Technology

[0002] Scene recognition technology based on mobile terminals such as smartphones is an important basic perception capability that can serve a variety of businesses such as smart mobility, direct service access, intent decision-making, smart noise reduction for headphones, search, and recommendation.

[0003] Existing scene recognition technologies mainly include: image-based scene recognition methods, sensor-based (e.g., WiFi, Bluetooth, location sensors, etc.) location-based scene recognition methods, or signal fingerprint-based scene recognition methods.

[0004] Image-based scene recognition methods are significantly impacted by factors such as camera field of view, shooting angle, and object occlusion, posing a major challenge to their robustness in complex scenes and making it difficult to achieve seamless real-time scene recognition for the user. Specifically, image-based scene recognition methods use front or rear cameras to capture images for identification. However, due to limited field of view, few effective features, arbitrary shooting angles, or the presence of numerous object features within the same image, noise can overwhelm key features, all of which can easily lead to misjudgments. For example, the presence of a ceiling in an image might be mistakenly identified as indoors, when in reality it could be on a subway or airplane. Therefore, the technical problem this application aims to solve is how to improve the accuracy of image-based scene recognition methods. Summary of the Invention

[0005] In view of this, a scene recognition method and device are proposed, which combines multiple images and the azimuth angle corresponding to each image to recognize the scene. Since more comprehensive scene information is obtained, the accuracy of image-based scene recognition can be improved, and the problem of limited field of view and shooting angle of single-camera scene recognition is solved, resulting in more accurate recognition.

[0006] In a first aspect, embodiments of this application provide a scene recognition method, the method comprising: a terminal device acquiring images of the same scene from multiple azimuth angles using multiple cameras, wherein the azimuth angle at which each camera acquires the image is the azimuth angle corresponding to the image, and the azimuth angle is the angle between the direction vector of each camera acquiring the image and the unit gravity vector; the terminal device recognizing the same scene based on the image and the azimuth angle corresponding to the image, thereby obtaining a scene recognition result.

[0007] The scene recognition method of this application uses multiple cameras to capture images of the same scene from multiple azimuth angles, and combines multiple images and the azimuth angle corresponding to each image to recognize the scene. Since more comprehensive scene information is obtained, the accuracy of image-based scene recognition can be improved, solving the problem of limited field of view and shooting angle of single-camera scene recognition, and the recognition is more accurate.

[0008] According to the first aspect, in a first possible implementation, the terminal device identifies the same scene based on the image and the azimuth angle corresponding to the image to obtain a scene recognition result, including: the terminal device extracts the azimuth angle feature corresponding to the image from the azimuth angle corresponding to the image; and uses a scene recognition model to identify the same scene based on the image and the azimuth angle feature corresponding to the image to obtain a scene recognition result, wherein the scene recognition model is a neural network model.

[0009] By assigning different weights (azimuth features) to the convolution kernels in the neural network model, we can extract the features most relevant to the scene under different azimuth angles and obtain more accurate prediction results.

[0010] According to the first possible implementation of the first aspect, in the second possible implementation, the scene recognition model includes multiple pairs of first feature extraction layers and first layer models. Each pair of first feature extraction layers and first layer models is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of the azimuth angle to obtain a first recognition result. The azimuth angle corresponding to the image of the azimuth angle is the azimuth angle corresponding to the first recognition result. The first feature extraction layer is used to extract features from the image of the azimuth angle to obtain a feature vector. The first layer model is used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle. The scene recognition model further includes a second layer model, which is used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

[0011] The scene recognition method in this application adopts a two-layer scene recognition model, collects images from multiple angles and combines the azimuth angles of the images, and uses a competition mechanism to perform scene recognition. It considers both local and global features, which can improve the accuracy of scene recognition in a user-unnoticed manner and reduce misjudgments.

[0012] According to the second possible implementation of the first aspect, in the third possible implementation, the first feature extraction layer is used to extract features of the image of the azimuth angle to obtain multiple feature vectors; the first layer model is used to calculate the first weight corresponding to each of the multiple feature vectors based on the azimuth angle corresponding to the image of the azimuth angle; the first layer model is used to obtain the first recognition result based on each feature vector and the first weight corresponding to each feature vector.

[0013] According to the second possible implementation of the first aspect, in the fourth possible implementation, the second layer model is used to calculate the second weight of the first identification result based on the azimuth angle corresponding to the first identification result; the second layer model is used to obtain the scene identification result based on the first identification result and the second weight of the first identification result.

[0014] According to the second possible implementation of the first aspect, in the fifth possible implementation, the second layer model presets a third weight corresponding to each of the first recognition results, and the second layer model is used to obtain the scene recognition result based on the first recognition result and the third weight corresponding to the first recognition result.

[0015] According to the second possible implementation of the first aspect, in the sixth possible implementation, the second layer model is used to determine the fourth weight corresponding to each of the first identification results based on the azimuth angle and the preset rule; wherein, the preset rule is a weight set based on the azimuth angle, different azimuth angles correspond to different weight sets, and each weight set includes the fourth weight corresponding to each of the first identification results; the second layer model is used to obtain the scene identification result based on the first identification result and the fourth weight corresponding to the first identification result.

[0016] For the second-layer model, which uses a time-limited method with preset weights or weight mapping functions, only the other parts of the neural network model need to be trained during training, and the second-layer model does not need to be trained, which can improve training efficiency.

[0017] According to the first aspect, in the seventh possible implementation, the method further includes: the terminal device acquiring the acceleration of the gravity sensor on the coordinate axes of the three-dimensional Cartesian coordinate system corresponding to each camera capturing the image, and obtaining the direction vector of each camera capturing the image; wherein the three-dimensional Cartesian coordinate system corresponding to each camera capturing the image has each camera as its origin, the z-direction is the direction along which the camera captures the image, and x and y are directions perpendicular to the z-direction, and the planes containing x and y are perpendicular to the z-direction; the azimuth angle is calculated based on the direction vector and the gravity unit vector.

[0018] According to the first possible implementation of the first aspect, in the eighth possible implementation, before using the scene recognition model to identify the same scene based on the image and the azimuth features corresponding to the image, the method further includes: the terminal device preprocessing the image; wherein, the preprocessing includes one or more combinations of the following processes: converting the image format, converting the image channels, unifying the image size, and normalizing the image, where converting the image format refers to converting a color image to a black and white image, converting the image channels refers to converting the image to the red, green, and blue RGB channels, unifying the image size refers to adjusting multiple images to have the same length and width, and image normalization refers to normalizing the pixel values ​​of the image.

[0019] Secondly, embodiments of this application provide a scene recognition device, the device comprising: an image acquisition module, configured to acquire images of the same scene from multiple azimuth angles using multiple cameras, wherein the azimuth angle at which each camera acquires the image is the azimuth angle corresponding to the image, and the azimuth angle is the angle between the direction vector of each camera acquiring the image and the unit gravity vector; and a recognition module, configured to recognize the same scene based on the image and the azimuth angle corresponding to the image, and obtain a scene recognition result.

[0020] The scene recognition device in this application uses multiple cameras to capture images of the same scene from multiple azimuth angles, and combines multiple images and the azimuth angle corresponding to each image to recognize the scene. Since more comprehensive scene information is obtained, the accuracy of image-based scene recognition can be improved, solving the problem of limited field of view and shooting angle for single-camera scene recognition, and making the recognition more accurate.

[0021] According to the second aspect, in a first possible implementation, the scene recognition module includes: an azimuth feature extraction module, used to extract the azimuth features corresponding to the image from the azimuth corresponding to the image; and a scene recognition model, used to recognize the same scene based on the image and the azimuth features corresponding to the image, and obtain a scene recognition result, wherein the scene recognition model is a neural network model.

[0022] According to the first possible implementation of the second aspect, in the second possible implementation, the scene recognition model includes multiple pairs of first feature extraction layers and first layer models. Each pair of first feature extraction layers and first layer models is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of the azimuth angle to obtain a first recognition result. The azimuth angle corresponding to the image of the azimuth angle is the azimuth angle corresponding to the first recognition result. The first feature extraction layer is used to extract features from the image of the azimuth angle to obtain a feature vector. The first layer model is used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle. The scene recognition model further includes a second layer model, which is used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

[0023] The scene recognition device in this application adopts a two-layer scene recognition model, collects images from multiple angles and combines the azimuth angles of the images, and uses a competition mechanism to perform scene recognition. It considers both local and global features, which can improve the accuracy of scene recognition in a user-unnoticed manner and reduce misjudgments.

[0024] According to the second possible implementation of the second aspect, in the third possible implementation, the first feature extraction layer is used to extract features of the image of the azimuth angle to obtain multiple feature vectors;

[0025] The first layer model is used to calculate the first weight of each feature vector among the multiple feature vectors based on the azimuth angle corresponding to the image of the azimuth angle;

[0026] The first layer model is used to obtain the first recognition result based on each feature vector and the first weight corresponding to each feature vector.

[0027] According to the second possible implementation of the second aspect, in the fourth possible implementation, the second layer model is used to calculate the second weight of the first identification result based on the azimuth angle corresponding to the first identification result; the second layer model is used to obtain the scene identification result based on the first identification result and the second weight of the first identification result.

[0028] According to the second possible implementation of the second aspect, in the fifth possible implementation, the second layer model presets a third weight corresponding to each of the first recognition results, and the second layer model is used to obtain the scene recognition result based on the first recognition result and the third weight corresponding to the first recognition result.

[0029] According to the second possible implementation of the second aspect, in the sixth possible implementation, the second layer model is used to determine the fourth weight corresponding to each of the first identification results based on the azimuth angle and the preset rule; wherein, the preset rule is a weight set based on the azimuth angle, different azimuth angles correspond to different weight sets, and each weight set includes the fourth weight corresponding to each of the first identification results; the second layer model is used to obtain the scene identification result based on the first identification result and the fourth weight corresponding to the first identification result.

[0030] According to the second aspect, in the seventh possible implementation, the device further includes: an azimuth angle acquisition module, used to acquire the acceleration of the gravity sensor on the coordinate axes of the three-dimensional Cartesian coordinate system corresponding to each camera acquiring the image, and to obtain the direction vector of each camera acquiring the image; wherein, the three-dimensional Cartesian coordinate system corresponding to each camera acquiring the image has each camera as its origin, the z-direction is the direction along which the camera captures the image, and x and y are directions perpendicular to the z-direction, and the planes containing x and y are perpendicular to the z-direction; the azimuth angle is calculated based on the direction vector and the gravity unit vector.

[0031] According to the first possible implementation of the second aspect, in the eighth possible implementation, the apparatus further includes: an image preprocessing module for preprocessing the image; wherein the preprocessing includes one or more combinations of the following processes: converting image format, converting image channels, unifying image size, and image normalization; converting image format refers to converting a color image to a black and white image; converting image channels refers to converting an image to the red, green, and blue RGB channels; unifying image size refers to adjusting multiple images to have the same length and width; and image normalization refers to normalizing the pixel values ​​of the image.

[0032] Thirdly, embodiments of this application provide a terminal device that can execute one or more of the scene recognition methods described in the first aspect or various possible implementations of the first aspect.

[0033] Fourthly, embodiments of this application provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code is run in an electronic device, the processor in the electronic device executes one or more of the scene recognition methods described in the first aspect or various possible implementations of the first aspect.

[0034] Fifthly, embodiments of this application provide a non-volatile computer-readable storage medium storing computer program instructions thereon, characterized in that, when the computer program instructions are executed by a processor, they implement one or more of the scene recognition methods described in the first aspect or various possible implementations of the first aspect.

[0035] These and other aspects of this application will become more apparent in the description of the following embodiments(s). Attached Figure Description

[0036] The accompanying drawings, which are included in and form part of this specification, illustrate exemplary embodiments, features, and aspects of this application together with the specification and serve to explain the principles of this application.

[0037] Figure 1a and Figure 1b Schematic diagrams of application scenarios according to an embodiment of this application are shown respectively.

[0038] Figure 2a This diagram illustrates scene recognition using a neural network model according to an embodiment of the present application.

[0039] Figure 2b A flowchart illustrating a scene recognition method according to an embodiment of this application is shown.

[0040] Figure 3 A schematic diagram illustrating an application scenario according to an embodiment of this application is shown.

[0041] Figure 4a and Figure 4b Schematic diagrams are shown for an azimuth angle determination method according to an embodiment of this application.

[0042] Figure 5 A block diagram showing the structure of a neural network model according to an embodiment of this application is provided.

[0043] Figure 6 A schematic diagram of the structure of a first-layer model according to an embodiment of this application is shown.

[0044] Figure 7 A flowchart illustrating a scene recognition method according to an embodiment of this application is shown.

[0045] Figure 8 A flowchart illustrating step S701 of a method according to an embodiment of this application is shown.

[0046] Figure 9 A block diagram of a scene recognition device according to an embodiment of this application is shown.

[0047] Figure 10 A schematic diagram of the structure of a terminal device according to an embodiment of this application is shown. Detailed Implementation

[0048] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.

[0049] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.

[0050] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.

[0051] Glossary

[0052] Azimuth: The angle between the camera's direction vector and the unit gravity vector.

[0053] Camera direction vector: A three-dimensional rectangular coordinate system is established with the camera as the origin. The z-direction is the direction along which the camera is shooting. The x and y directions are perpendicular to the z-direction, and the planes containing the x and y directions are perpendicular to the z-direction. The vector formed by the accelerations of the gravity sensor in the x, y, and z directions is the camera direction vector.

[0054] Gravity unit vector: (0,0,1).

[0055] Therefore, the technical problem this application aims to solve is how to improve the recognition accuracy of image-based scene recognition methods. Existing image-based scene recognition methods can be applied in the following scenarios: 1. Recognition is performed by acquiring images using only a single camera, resulting in a small field of view, few observed effective features, and low scene recognition recall; 2. Since scene recognition is imperceptible to the user, the camera's shooting angle is arbitrary, making it easy to misjudge similar objects or features. For example, if a ceiling feature appears in an image, it may be identified as indoors, when in reality it could be on a subway or airplane; 3. A large amount of object feature information exists in the same image, and noise information may overwhelm the main features, causing misjudgment of the main features.

[0056] To address the aforementioned technical issues, this application provides a scene recognition method that utilizes multiple cameras to capture images of the same scene from multiple azimuth angles, obtaining the azimuth angles at which the multiple cameras captured (acquired) the images. The azimuth angle at which each camera captured (acquired) the image is the azimuth angle corresponding to that image. By combining multiple images and the azimuth angle corresponding to each image, the scene is recognized. Since more comprehensive scene information is obtained, the accuracy of image-based scene recognition can be improved, solving the problem of limited field of view and shooting angle for single-camera scene recognition, resulting in more accurate recognition.

[0057] In one possible implementation, the multiple cameras in this application embodiment can be mounted on a terminal device. For example, the multiple cameras can be front-facing and rear-facing cameras on a mobile phone, multiple cameras mounted at different locations on a vehicle body, multiple cameras mounted at different directions on a drone, and so on. It should be noted that the above application scenarios are merely some examples of this application, and this application is not limited thereto.

[0058] Figure 1a and Figure 1b Schematic diagrams illustrating application scenarios according to an embodiment of this application are shown respectively. For example... Figure 1a As shown, a mobile phone can be equipped with a front-facing camera and a rear-facing camera. These two cameras can capture images from two different angles and also provide their azimuth angles. Combining these two images with their corresponding azimuth angles allows for more accurate scene recognition. For example... Figure 1b As shown, multiple cameras can be installed on the body of an autonomous vehicle, and these cameras can be placed in different locations, for example... Figure 1b As shown, cameras can be installed at the front, rear, sides, and roof of the vehicle. The direction of each camera can be adjusted individually. Autonomous vehicles can also be equipped with a controller that can connect to multiple cameras. It should be noted that autonomous vehicles can also be equipped with other sensors, such as GPS, radar, accelerometers, and gyroscopes. All sensors and cameras are connected to the controller. The controller can collect images from different angles through multiple cameras and obtain the azimuth angle of each camera. Based on multiple images and the corresponding azimuth angle of each image, the scene can be identified more accurately.

[0059] It should be noted that, Figure 1a and Figure 1b This is merely an example of an application scenario provided in this application, and the application is not limited thereto. For instance, this application can also be applied to scenarios where scene recognition is performed using images acquired by drones.

[0060] The scene recognition method provided in this application can be applied to terminal devices. For example, the terminal devices in this application can be smartphones, netbooks, tablets, laptops, wearable electronic devices (such as smart bracelets, smartwatches, etc.), TVs, virtual reality devices, speakers, e-ink devices, etc. Figure 10 This diagram illustrates the structure of a terminal device according to an embodiment of the present application. Taking a mobile phone as an example, the terminal device is shown below. Figure 10 A schematic diagram of the structure of mobile phone 200 is shown, and the details can be found in the following description.

[0061] In one possible implementation of this application, scene recognition is performed based on the front and rear cameras and accelerometer of a mobile phone. Specifically, the front and rear cameras simultaneously capture images from different azimuth angles, and the phone's accelerometer is used to extract the angle between the current orientation of the phone's camera and the direction of gravity; this angle can be used as the camera's azimuth angle.

[0062] In one possible implementation, the scene recognition method provided in this application embodiment can be implemented using a neural network model (scene recognition model). Figure 2a This diagram illustrates scene recognition using a neural network model according to an embodiment of this application. Figure 2a As shown, the neural network model used in this application embodiment may include: multiple pairs of first feature extraction layers and first layer models. Each pair of first feature extraction layers and first layer models is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of an azimuth angle to obtain a first recognition result. The azimuth angle corresponding to the image of an azimuth angle is the azimuth angle corresponding to the first recognition result. The first feature extraction layer is used to extract features from the image to obtain a feature vector (feature map), such as... Figure 2a The feature extraction layer 1 and feature extraction layer 2 shown are used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle.

[0063] like Figure 2a As shown, the neural network model may also include a second layer model, which is used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

[0064] Figure 2a The example shown includes two pairs of first feature extraction layers and a first-layer model. Figure 2a The application scenario shown can be applied to dual-camera (front and rear dual-camera) scenarios. Image 1 and Image 2 can be images captured by cameras at different angles. For example, Image 1 is captured by the front camera of the phone, and Image 2 is captured by the rear camera of the phone.

[0065] In one possible implementation, the number of logarithms of the first feature extraction layer and the first layer model included in the neural network model can be configured according to the number of camera angles set in the specific application scenario. For example, the number of logarithms of the first feature extraction layer and the first layer model can be greater than or equal to the number of angles.

[0066] In one possible implementation, the first feature extraction layer can be implemented using a Convolutional Neural Network (CNN). The CNN extracts features from the input image to obtain a feature map. For example, VGG (Visual Geometry Group Network), Inception, MobileNet, ResNet, DenseNet, Transformer, and other convolutional neural network models can be used as the first feature extraction layer to extract the feature map. Alternatively, a custom convolutional neural network structure can be used as the first feature extraction layer; this application does not limit this approach. Both the first and second layer models can be implemented based on attention mechanisms. The second layer model can also be implemented by pre-setting weights for each azimuth angle or by pre-setting a weight mapping function based on the azimuth angle; this application does not limit this approach either.

[0067] The scene recognition method provided in this application uses a two-layer scene recognition model based on a competition mechanism (attention mechanism): The first-layer model assigns different weights to the convolution result (the feature vector obtained by CNN convolution operation on the input image) according to the azimuth angle, activates neurons that identify local features of different scenes, extracts key features of the image at that azimuth angle, performs the first scene classification, and obtains the first recognition result; the second-layer model calculates the weights of the scene classification results at different azimuth angles under different scenes, and sums them to obtain the classification results of multiple images from different perspectives, thus obtaining the final scene recognition result. Using a two-layer scene recognition model based on a competition mechanism and combining azimuth angle information can identify key features and filter irrelevant information, effectively reducing the probability of misidentification. For example, from an upward perspective, it is impossible to distinguish between an airplane and a high-speed train (the ceiling features are similar and difficult to distinguish), while from a side perspective, they can be distinguished (the round window of an airplane and the square window of a high-speed train are easily distinguishable). The scene recognition method in this application helps to reduce misjudgment.

[0068] Therefore, by adopting a two-layer scene recognition model, collecting images from multiple angles and combining the azimuth angles of the images, and using a competition mechanism for scene recognition, the results of both local and global features are considered. This approach can improve the accuracy of scene recognition in a user-unnoticed manner and reduce misjudgments.

[0069] Figure 2aThe azimuth feature extraction shown (dashed box) can be implemented by a neural network. That is, the neural network model provided in this application embodiment can also include a second feature extraction layer, which can also be implemented by a convolutional neural network model. Figure 2a The azimuth feature extraction shown (dashed box) can also be obtained by calculating the azimuth using existing functions, and this application does not limit this.

[0070] Figure 2a The sensors shown can be accelerometers, gyroscopes, etc. The attitude of the terminal device can be obtained by collecting motion data from the sensors. The direction of the camera can be determined based on the attitude of the terminal device and the position of the camera. The azimuth angle of the camera can be determined based on the direction of the camera and the direction of gravity.

[0071] Figure 2b A flowchart of a scene recognition method according to an embodiment of this application is shown below. Figure 2a and Figure 2b The process of the image processing method of this application is described in detail.

[0072] 1. Image Acquisition

[0073] Multiple cameras positioned at different locations on the terminal device capture images of the same scene, allowing the terminal device to obtain images of the same scene from different perspectives (azimuth angles).

[0074] For example, simultaneously using the phone's front and rear cameras to capture images of the same scene from different perspectives can yield images of the same scene. The image captured by the front camera can be a single image or a composite image of images from multiple cameras; similarly, the image captured by the rear camera can also be a single image or a composite image of images from multiple cameras. In the scenario where images from multiple cameras are combined into a single image, the azimuth angle of this composite image is the same as that of a single camera.

[0075] In the embodiments of this application, the image captured by the camera can be a black and white image, an RGB (Red, Green, Blue) color image, an RGB-D (RGB-Depth) depth image (D refers to depth information), or an infrared image; this application does not limit the type of image captured.

[0076] Figure 3 A schematic diagram illustrating an application scenario according to an embodiment of this application is shown. For example... Figure 3As shown, taking the subway as an example, when a user looks at their phone on the subway, the phone is tilted at a certain angle to the subway floor. Therefore, the phone's front camera can capture the image of the subway ceiling, and the phone's rear camera can capture the image of the subway floor.

[0077] 2. Azimuth angle acquisition

[0078] The terminal device can acquire the azimuth angles of multiple camera angles simultaneously while capturing images. The azimuth angle of a camera can refer to the angle between the camera's direction vector and the unit vector of gravity.

[0079] In the embodiments of this application, the orientation vector of the camera can be obtained through sensors, such as gravity sensors, accelerometers, gyroscopes, etc., and this application does not limit this to any particular sensor. The unit vector of gravity is g. gravity = (0,0,1). Therefore, embodiments of this application acquire the direction vector of the camera, and based on the direction vector of the camera and g gravity The azimuth angle of the camera can be calculated by using (0,0,1).

[0080] Taking a mobile phone as an example, the orientation vector of the front-facing camera can be obtained through a gravity sensor. Figure 4a and Figure 4b Schematic diagrams are shown for each embodiment of an azimuth angle determination method according to this application. For example... Figure 4a As shown, a three-dimensional Cartesian coordinate system can be established with the front-facing camera as the origin. The z-direction is along the direction the front-facing camera is shooting and perpendicular to the phone's plane, while x and y are directions parallel to the phone's frame and perpendicular to the z-direction, respectively. The direction vector g of the front-facing camera can be obtained from the acceleration of the gravity sensor in the x, y, and z directions. camera = (Acc_x, Acc_y, Acc_z), assuming the azimuth angle of the front camera is θ, then the azimuth angle of the rear camera can be π-θ. Therefore, according to g camera and the unit vector of gravity g gravity = (0,0,1) can be used to calculate the azimuth angle θ of the front camera, which in turn gives the azimuth angles of the front and rear cameras of the phone.

[0081] Specifically, such as Figure 4b As shown, the direction vector g of the phone's front-facing camera camera With the unit vector of gravity g gravity The included angle θ satisfies formula (1):

[0082]

[0083] Therefore, the azimuth angle of the front-facing camera can be calculated using formula (2):

[0084]

[0085] Therefore, the azimuth angle of the phone's rear camera is

[0086] 3. Image preprocessing

[0087] The terminal device can preprocess the images captured by each camera separately. The preprocessing can include one or more of the following methods: image format conversion, image size unification, image normalization, and image channel conversion.

[0088] Image format conversion can refer to converting a color image to a black and white image.

[0089] Image channel conversion can refer to converting a color image to RGB three-channel, with the three channels being Red, Green, and Blue in sequence.

[0090] Image size standardization can refer to standardizing the length and width of images captured by each camera. For example, after standardization, the image length is 800 pixels and the width is 600 pixels.

[0091] The purpose of image normalization is to ensure that the mean of the extracted feature maps (feature vectors) is near 0. Therefore, for black and white images, image normalization can mean subtracting the mean of 127.5 from the pixel values ​​of the black and white image and then dividing by 128. That is, p_Normalized = (p - 127.5) / 128, where p can represent the pixel values ​​of the black and white image, and p_Normalized can represent the normalized pixel values. For color images, such as images represented in RGB format, the pixel values ​​can be subtracted from the mean [103.939, 116.779, 123.68], i.e.:

[0092] p_R_Normalized=p_R-103.939;

[0093] p_G_Normalized=p_G-116.779;

[0094] p_B_Normalized=p_B-123.68;

[0095] Where p_R, p_G, and p_B represent the pixel values ​​before normalization, and p_R_Normalized, p_G_Normalized, and p_B_Normalized represent the pixel values ​​after normalization.

[0096] The acquired images can be processed using one or more of the preprocessing methods described above. For example, in one possible implementation, the preprocessing procedure may include:

[0097] Step 1: Image format conversion. Convert the acquired color image to a black and white image. The conversion of a color image to a black and white image can be done by using the relevant grayscale formula method or the average value method to calculate the value of the pixel after conversion based on the value of the pixel before conversion.

[0098] Step 2: Standardize image size. For example, standardize the image size to 800 pixels long and 600 pixels wide.

[0099] Step 3: Image normalization. For a black and white image, p_Normalized = (p-127.5) / 128.

[0100] In another possible implementation, the preprocessing procedure may include:

[0101] Step 1: Image channel conversion. For color images, they can be uniformly converted to RGB three channels, namely Red, Green, and Blue in sequence. If the images before preprocessing are all in RGB format, the image channel conversion process can be omitted.

[0102] Step 2: Standardize image size. For example, standardize the image size to 800 pixels long and 600 pixels wide.

[0103] Step 3: Image normalization. For RGB images, the normalization process is as follows:

[0104] R channel: p_R_Normalized = (p_R - 103.939) / 1.0;

[0105] G channel: p_G_Normalized = (p_G - 116.779) / 1.0;

[0106] Channel B: p_B_Normalized = (p_B - 123.68) / 1.0.

[0107] In the embodiments of this application, the preprocessing methods used by the terminal device for images captured by the camera at each angle can be the same or different, and this application does not limit this. Preprocessing the captured images can unify the image format, which is beneficial for subsequent feature extraction and scene recognition processes.

[0108] 4. Azimuth feature extraction

[0109] The azimuth angles acquired in step 2 are used to extract azimuth features. Azimuth feature extraction can be achieved by combining one or more of the following processing methods: numerical normalization, discretization, one-bit effective encoding, trigonometric function transformation, etc. This application provides several different feature extraction methods; several examples are provided below.

[0110] Example 1: The terminal device can discretize the acquired azimuth angle and then perform one-hot encoding (one effective bit encoding). Discretization maps individual data points to a finite space without changing the relative size of the data. For example, for the azimuth angle in this embodiment, it can be discretized into four intervals: [0°, 45°), [45°, 90°), [90°, 135°), and [135°, 180°]. The interval to which the azimuth angle is mapped is encoded as 1, and the other intervals are encoded as 0. The feature vectors corresponding to the four intervals are [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 1, 0], and [0, 0, 0, 1]. For example, if the azimuth angle θ is 30° and it is mapped to the interval [0°, 45°], then the corresponding azimuth angle feature is [1, 0, 0, 0].

[0111] Example 2: The terminal device can directly perform trigonometric function transformations on the azimuth angle, and the resulting value is normalized to the [0,1] interval as the azimuth angle feature. The trigonometric function transformations can refer to sinθ, cosθ, tanθ, etc.

[0112] Example 3: The normalized trigonometric function values ​​in the interval [0,1] are discretized into four intervals: [0,0.25), [0.25,0.5), [0.5,0.75), and [0.75,1.0]. The terminal device can first perform trigonometric function transformation and normalization on the azimuth angle, and then determine the azimuth angle characteristics based on the interval to which the normalized trigonometric function values ​​are mapped. For example, if the azimuth angle θ is 30°, sinθ = 1 / 2, which maps to the interval [0.5,0.75), then the azimuth angle characteristics of θ are [0,0,1,0].

[0113] By assigning different weights (azimuth features) to the convolution kernels in the neural network model, we can extract the features most relevant to the scene under different azimuth angles and obtain more accurate prediction results.

[0114] 5. Scene Recognition

[0115] In the embodiments of this application, scene recognition is performed using a neural network model, combined with... Figure 2a The neural network model shown illustrates the process of scene recognition.

[0116] The terminal device uses the preprocessed image from step 3 as input data for multiple first feature extraction layers, and the azimuth features obtained from step 4 as input data for multiple first-layer models. Furthermore, the image and azimuth received by each pair of first feature extraction layers and first-layer models are associated. In other words, the image captured by the camera at the same azimuth angle and the azimuth feature corresponding to that camera are used as input data for a pair of first feature extraction layers and first-layer models, respectively.

[0117] For example, a mobile phone's front-facing camera captures an image, and after preprocessing the image, image 1 is obtained. The azimuth angle of the front-facing camera is θ1, and the azimuth feature of θ1 is C1. The neural network model includes a feature extraction layer 1 and a first-layer model 1. The output of feature extraction layer 1 is the input of first-layer model 1. The neural network model may also include a feature extraction layer 2 and a first-layer model 2. The output of feature extraction layer 2 is the input of first-layer model 2. The terminal device can use image 1 as the input of feature extraction layer 1 and C1 as the input of first-layer model 1, or the terminal device can use image 1 as the input of feature extraction layer 2 and C1 as the input of first-layer model 2. That is to say, the sequence numbers of image 1, image 2, feature extraction layer 1, feature extraction layer 2, first-layer model 1, and first-layer model 2 in this application are not intended to limit the order or correspondence, but are merely numbered to distinguish different modules and are not construed as limiting this application.

[0118] In this way, the first feature extraction layer extracts the features of the image to obtain a feature map (feature vector). The feature map (feature vector) is then input into the first layer model. The first layer model identifies (classifies) the scene of the image based on the feature map and the corresponding azimuth features of the image, and thus obtains the first recognition result.

[0119] The terminal device can also use azimuth features as input data for the second-layer model. Based on the first recognition result output by the first-layer model and the corresponding azimuth features, the second-layer model can further identify (classify) the scene and obtain the scene recognition result.

[0120] Figure 5 A block diagram showing the structure of a neural network model according to an embodiment of this application is provided. Figure 6 A schematic diagram of the structure of a first-layer model according to an embodiment of this application is shown.

[0121] Assuming that in the embodiments of this application, the number of feature vectors extracted by CNN is n, and the terminal device of this application performs image acquisition at J angles (azimuth angles), that is, it uses cameras at J angles to acquire images at J angles.

[0122] like Figure 5 As shown, assuming in Figure 5In the example, J is 2. Two cameras at different angles captured images 1 and 2 respectively. The azimuth angle corresponding to image 1 is θ1, and the azimuth angle corresponding to image 2 is θ2. The azimuth feature of azimuth angle θ1 is C1, and the azimuth feature of azimuth angle θ2 is C2. Image 1 serves as the input data for the upper CNN. The upper CNN extracts features from the input image 1 to obtain the feature vector y. i , where i is a positive integer from 1 to n. Image 2 serves as the input data for the lower CNN. The lower CNN extracts features from the input image 2 to obtain the feature vector x. i .

[0123] The terminal device will use the feature vector y i The azimuth feature C1 is used as the input data for the first layer model on the upper side. For example... Figure 6 As shown, the first layer model is based on the feature vector y i The first identification result Z can be calculated from the azimuth feature C1. j Where j is a positive integer from 1 to J, and J represents the number of angles (azimuth angles) at which the image is acquired. In this example, J equals 2. Figure 6 In the example, j equals 1. The first-layer model can be implemented based on an attention mechanism; specifically, it can include... Figure 6 The activation function (tanh), softmax function, and weighted averaging process are shown. Tanh is just one example of an activation function; this application is not limited to it and other activation functions can also be used, such as sigmoid, ReLU, LReLU, ELU, etc.

[0124] The specific calculation method is shown in the following formula (3):

[0125]

[0126] Among them, C j This represents the azimuth feature of the image corresponding to this angle, y i W is the feature vector output by the CNN. i and b i These are the weights and biases of the activation function, respectively. [C j y i The symbol represents the vector obtained by concatenating the feature vector and the azimuth feature. This represents the parameters of the softmax function, m, calculated using the tanh function. i It can determine whether the corresponding neuron is activated, extract the features corresponding to the tanh function as the basis for classification, and the softmax function is applied to m. i Normalization is performed to obtain the weight of each feature vector. Therefore, the azimuth feature can affect the calculated feature vector y. iweights s i If the calculated s i If the value of s is relatively large, then the feature vector will have a significant impact on the classification result. i If the feature vector is relatively small, then its impact on the classification result is relatively small. Therefore, the scene recognition model provided in this application can identify key features, filter irrelevant information, reduce noise, improve recognition accuracy, and achieve scene recognition that is imperceptible to the user based on the azimuth angle.

[0127] The weights s calculated according to formula (3) i The weight values ​​can be represented by different features extracted from the image. By summing the feature vectors and their corresponding weights, the first recognition result Z1 of the first layer model can be obtained.

[0128] The terminal device will use the feature vector x i (i is a positive integer from 1 to n', n' and n can be equal or not, and this application does not limit this) and azimuth feature C2 are used as input data for the first layer model on the lower side. The first recognition result Z2 can be calculated in the same way as the first layer model on the upper side.

[0129] The neural network model in this embodiment extracts feature maps using CNN, assigns different weights to different convolution results of the image based on the azimuth angle, and the first layer model identifies key features and filters irrelevant information through a competitive mechanism, which can effectively reduce the probability of misidentification.

[0130] In one possible implementation, the structure of the second-layer model is as follows: Figure 5 As shown, f1 and f2 represent the activation function tanh. The number of tanh functions in the second-layer model is related to the number of angles in the acquired images; the number of tanh functions can be equal to or greater than the number of angles. The input to the tanh function includes the output Z of the first-layer model. j and azimuth feature C j The second-layer model also includes a softmax function, which calculates the output Z of the first-layer model based on the result of the tanh function. j Weight S j Finally, the second-layer model uses the calculated weights S j The output Z of the first layer model j Calculate the final scene recognition result Z. The specific calculation method is shown in the following formula (4):

[0131]

[0132] Among them, [C j Z j] represents the first recognition result Z of image 1. j and the corresponding azimuth feature C j The concatenated vector, W j and b j These represent the weights and biases of the tanh function, respectively. This represents the parameters of the softmax function, M, calculated from the tanh function. j It can determine whether the corresponding neuron is activated, extract the features corresponding to the tanh function (the first recognition result) as the basis for classification, and use the softmax function on M. j Normalization is performed to obtain the weight of each first identification result. Therefore, azimuth features can affect the calculated first identification result Z. j Weight S j If the calculated S j If the value is relatively large, then the first recognition result has a significant impact on the classification result. If the calculated S... j If the angle is relatively small, then the initial recognition result has a smaller impact on the classification result. Therefore, the scene recognition model provided in this application can extract key angle features and filter irrelevant angles based on the azimuth angle, thereby improving recognition accuracy and achieving scene recognition that is imperceptible to the user.

[0133] The weight S is calculated according to formula (4). j This can represent the weight values ​​assigned to different first recognition results. By summing the first recognition results and their corresponding weights, the scene recognition result Z of the second-layer model can be obtained.

[0134] Among them, W i b i , W j b j and The uniformity is used as the model parameter, and sample data can be used to compare it. Figure 5 The parameter values ​​are obtained by training the neural network model shown. The neural network model of this application can be trained using the training methods in the relevant prior art. The training process will not be described in detail here.

[0135] In another possible implementation, the second-layer model can also be implemented by pre-weighting each azimuth angle, voting, or by pre-setting a weight mapping function based on the azimuth angle. That is, in another embodiment of this application, the neural network model includes multiple pairs of feature extraction layers and a first-layer model, and also includes a second-layer model, which is implemented by pre-weighting each azimuth angle or by pre-setting a weight mapping function based on the azimuth angle.

[0136] For example, the second-layer model can pre-weight each azimuth angle, and for each first recognition result Z output by the first-layer model... j The preset weight is S. j The second-layer model is based on each first recognition result Z. j and the corresponding weight S j The calculation can yield the scene recognition results.

[0137] The first-layer model can also preset a weight mapping function based on the azimuth angle. That is, different azimuth angles correspond to different preset weight reassemblies. The preset weight reassemblies can include each first recognition result Z. j The corresponding weights. For example, suppose a terminal device captures images from two angles using its front and rear cameras, and identifies these images to obtain two initial recognition results, Z1 and Z2, with corresponding weights S1 and S2 respectively. In one example, the weight mapping function preset based on the azimuth angle can be as follows:

[0138] (1) If the azimuth angle θ of the front camera belongs to the interval [0°, 45°), S1 = 1.0, S2 = 0.0;

[0139] (2) If the azimuth angle θ of the front camera belongs to the interval [45°, 90°), S1 = 0.7, S2 = 0.3;

[0140] (3) If the azimuth angle θ of the front camera belongs to the interval [90°, 135°), S1 = 0.3, S2 = 0.7;

[0141] (4) If the azimuth angle θ of the front camera belongs to the interval [135°, 180°], S1 = 0.0, S2 = 1.0.

[0142] The above are just some examples of how the second-layer model can be implemented, and this application is not limited to them.

[0143] For the second-layer model, which uses a time-limited method with preset weights or weight mapping functions, only the other parts of the neural network model need to be trained during training, and the second-layer model does not need to be trained, which can improve training efficiency.

[0144] 6. Output Results

[0145] The terminal device can output the final recognition result based on the scene recognition result and the preset strategy. The preset strategy may include filtering the scene recognition result according to the confidence threshold and then outputting the final recognition result, or merging multiple categories into a major category and outputting the final recognition result, etc.

[0146] Example 1: Assuming a confidence threshold of 0.8, when the confidence of the category corresponding to the scene recognition result is greater than or equal to the threshold, the category corresponding to the scene recognition result can be predicted as the final recognition result.

[0147] Example 2: Suppose that the scene recognition result includes 100 categories. The terminal device can merge some of these categories into a larger category for output. For example, cars and buses can be merged into the car category as the final recognition result output.

[0148] Based on the examples above, this application provides a scene recognition method. Figure 7 A flowchart illustrating a scene recognition method according to an embodiment of this application is shown. Figure 7 As shown, the scene recognition method may include the following steps:

[0149] Step S700: The terminal device acquires images of the same scene from multiple azimuth angles using multiple cameras. The azimuth angle at which each camera acquires the image is the azimuth angle corresponding to the image. The azimuth angle is the angle between the direction vector of each camera when acquiring the image and the unit vector of gravity.

[0150] Step S701: The terminal device identifies the same scene based on the image and the azimuth angle corresponding to the image, and obtains the scene recognition result.

[0151] The scene recognition method of this application uses multiple cameras to capture images of the same scene from multiple azimuth angles, and combines multiple images and the azimuth angle corresponding to each image to recognize the scene. Since more comprehensive scene information is obtained, the accuracy of image-based scene recognition can be improved, solving the problem of limited field of view and shooting angle of single-camera scene recognition, and the recognition is more accurate.

[0152] In the embodiments of this application, the image captured by the camera can be a black and white image, an RGB (Red, Green, Blue) color image, an RGB-D (RGB-Depth) depth image (D refers to depth information), or an infrared image; this application does not limit the type of image captured.

[0153] In this application embodiment, an image of an azimuth angle can be an image captured by a single camera, or it can be a composite image of images captured by multiple cameras. For example, a mobile phone may include multiple rear cameras, and the image captured by the rear cameras can be a composite image of images captured by multiple cameras.

[0154] In one possible implementation, before identifying the same scene based on the image and the azimuth angle corresponding to the image, the method may further include:

[0155] The terminal device preprocesses the image; wherein the preprocessing includes one or more of the following processes in combination: converting the image format, converting the image channels, unifying the image size, and normalizing the image. Converting the image format means converting a color image to a black and white image; converting the image channels means converting the image to the red, green, and blue RGB channels; unifying the image size means adjusting multiple images to have the same length and width; and normalizing the image means normalizing the pixel values ​​of the image.

[0156] For details, please refer to the process shown in Part 3 above. The same preprocessing method can be used for each azimuth angle, or different preprocessing methods can be used according to the images acquired at each azimuth angle. This application does not limit this.

[0157] In this embodiment, the terminal device obtains the acceleration of the gravity sensor on the coordinate axes of the three-dimensional Cartesian coordinate system corresponding to each camera capturing the image, and can obtain the direction vector of each camera capturing the image; wherein, the three-dimensional Cartesian coordinate system corresponding to each camera capturing the image has each camera as its origin, the z-direction is along the direction of the camera's shooting, and x and y are the directions perpendicular to the z-direction; the azimuth angle is calculated based on the direction vector and the gravity unit vector. For details, please refer to Part 2 above, which will not be repeated here.

[0158] For step S701, the terminal device can preset the weight corresponding to each azimuth angle, and weight the recognition results of the image at each azimuth angle according to the weight of each azimuth angle to obtain the final scene recognition result. Alternatively, the image and the corresponding azimuth angle can be input into a trained neural network model to recognize the same scene and obtain the scene recognition result. This application does not limit the specific method of scene recognition.

[0159] Figure 8 A flowchart illustrating step S701 of a method according to an embodiment of this application is shown. Figure 8 As shown, in one possible implementation, step S701, where the terminal device identifies the same scene based on the image and the azimuth angle corresponding to the image, and obtains a scene identification result, may include:

[0160] Step S7010: The terminal device extracts the azimuth feature corresponding to the image from the azimuth angle corresponding to the image;

[0161] Step S7011: Using a scene recognition model, the same scene is identified based on the image and the azimuth features corresponding to the image to obtain a scene recognition result, wherein the scene recognition model is a neural network model.

[0162] For step S7010, please refer to the description in Part 4 above; it will not be repeated here. Figure 2a As shown, after extracting the azimuth features, the image and the corresponding azimuth features can be input into the scene recognition model. The scene recognition model is then used to identify the same scene based on the image and the corresponding azimuth features to obtain the scene recognition result.

[0163] In the embodiments of this application, the scene recognition model can be implemented using various different neural network structures. One possible implementation is as follows: Figure 2a As shown, the scene recognition model includes multiple pairs of first feature extraction layers and first layer models. Each pair of first feature extraction layers and first layer models is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of the azimuth angle to obtain a first recognition result.

[0164] Examples of the first feature extraction layer can be as follows: Figure 2a The image feature extraction layer 1 and feature extraction layer 2 are shown. Figure 2a The example only shows two pairs of first feature extraction layers and first model layers. This application is not limited to this. The number of pairs of first feature extraction layers and first model layers included in the neural network model can be configured according to the number of camera angles set in the specific application scenario. For example, the number of pairs of first feature extraction layers and first model layers can be greater than or equal to the number of angles.

[0165] Wherein, the azimuth angle corresponding to the image of the azimuth angle is the azimuth angle corresponding to the first recognition result, the first feature extraction layer is used to extract the features of the image of the azimuth angle to obtain a feature vector, and the first layer model is used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle.

[0166] like Figure 2a As shown, the azimuth feature 1 extracted from the azimuth of image 1 can be used as the input of the first layer model 1, and the azimuth feature 2 extracted from the azimuth of image 2 can be used as the input of the first layer model 2. Feature extraction layer 1 can extract feature vector 1 from image 1 and output it to the first layer model 1, and feature extraction layer 2 can extract feature vector 2 from image 2 and output it to the first layer model 2. The first layer model 1 can combine feature vector 1 and azimuth feature 1 to perform scene recognition and obtain the first recognition result 1, and the first layer model 2 can combine feature vector 2 and azimuth feature 2 to perform scene recognition and obtain the first recognition result 2.

[0167] like Figure 2aAs shown, the scene recognition model further includes a second-layer model. First-layer model 1 and first-layer model 2 can output the first recognition result 1 and the first recognition result 2 to the second-layer model, respectively. The azimuth feature 1, extracted from the azimuth angle corresponding to image 1, can be used as input to the second-layer model. Azimuth feature 1 corresponds to the first recognition result 1. Similarly, the azimuth feature 2, extracted from the azimuth angle corresponding to image 2, can be used as input to the second-layer model. Azimuth feature 2 corresponds to the first recognition result 2. The second-layer model can be used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

[0168] The scene recognition method in this application adopts a two-layer scene recognition model, collects images from multiple angles and combines the azimuth angles of the images, and uses a competition mechanism to perform scene recognition. It considers both local and global features, which can improve the accuracy of scene recognition in a user-unnoticed manner and reduce misjudgments.

[0169] In one possible implementation, the first feature extraction layer is used to extract features from the image at the given azimuth angle, obtaining multiple feature vectors; the first layer model is used to calculate a first weight corresponding to each of the multiple feature vectors based on the azimuth angle corresponding to the image at the given azimuth angle; the first layer model is used to obtain the first recognition result based on each feature vector and its corresponding first weight. The second layer model is used to calculate a second weight of the first recognition result based on the azimuth angle corresponding to the first recognition result; the second layer model is used to obtain the scene recognition result based on the first recognition result and its second weight.

[0170] like Figure 5 and Figure 6 As shown, the first-layer model can include an activation function and a softmax function. The activation function can be the tanh function, but other types of activation functions can also be used, and are not limited to these. Figure 5 and Figure 6 The example shown can also use the Sigmoid activation function or the ReLU activation function. The number of activation functions in the first layer model can be set according to the number of feature vectors extracted by the feature extraction layer, and can be greater than or equal to the number of feature vectors extracted. The activation function is used to determine whether to activate the corresponding neuron based on the feature vector and azimuth feature. The feature corresponding to the activation function is extracted as the basis for classification. The activation function and the softmax function are used to calculate the first weight corresponding to each feature vector, such as the s calculated by formula (3) above. i The specific process can be found above and will not be repeated here.

[0171] Azimuth features can affect the calculated feature vector y. i weights s i If the calculated s i If the value of s is relatively large, then the feature vector will have a significant impact on the classification result. i If the feature vector is relatively small, then its impact on the classification result is relatively small. Therefore, the scene recognition model provided in this application can identify key features, filter irrelevant information, reduce noise, improve recognition accuracy, and achieve scene recognition that is imperceptible to the user based on the azimuth angle.

[0172] Similarly, the second-layer model can include activation functions and softmax functions. The activation function can be the tanh function, but other types of activation functions can also be used, and are not limited to these. Figure 5 and Figure 6 The example shown illustrates this. The number of activation functions in the second-layer model is related to the number of angles in the acquired images; the number of activation functions can be equal to or greater than the number of angles. The input to the activation functions includes the output Z of the first-layer model. j and azimuth feature C j The second-layer model also includes a softmax function, which calculates the output Z of the first-layer model based on the results of the activation function. j Weight S j The specific calculation process can be found in formula (4) and the description in Part 5 above, and will not be repeated here.

[0173] Azimuth features can affect the calculated first identification result Z. j Weight S j If the calculated S j If the value is relatively large, then the first recognition result has a significant impact on the classification result. If the calculated S... j If the angle is relatively small, then the initial recognition result has a smaller impact on the classification result. Therefore, the scene recognition model provided in this application can extract key angle features and filter irrelevant angles based on the azimuth angle, thereby improving recognition accuracy and achieving scene recognition that is imperceptible to the user.

[0174] In one possible implementation, the second-layer model pre-sets a third weight corresponding to each of the first recognition results. The second-layer model is used to obtain the scene recognition result based on the first recognition result and the third weight corresponding to the first recognition result. For example, the second-layer model can pre-set weights for each azimuth angle. For each first recognition result Z output by the first-layer model... j The preset weight is S. j The second-layer model is based on each first recognition result Z.j and the corresponding weight S j The calculation can yield the scene recognition results.

[0175] In one possible implementation, the second-layer model is used to determine the fourth weight corresponding to each of the first identification results based on the azimuth angle and a preset rule; wherein the preset rule is a weight set based on the azimuth angle, different azimuth angles correspond to different weight sets, and each weight set includes the fourth weight corresponding to each of the first identification results; the second-layer model is used to obtain the scene identification result based on the first identification result and the fourth weight corresponding to the first identification result.

[0176] For example, the first-layer model can also pre-define a weight mapping function based on the azimuth angle. That is, different azimuth angles correspond to different pre-define weight reassemblies, and the pre-define weight reassemblies can include each first recognition result Z. j The corresponding weights. For example, suppose a terminal device captures images from two angles using its front and rear cameras, and identifies these images to obtain two initial recognition results, Z1 and Z2, with corresponding weights S1 and S2 respectively. In one example, the weight mapping function preset based on the azimuth angle can be as follows:

[0177] (1) If the azimuth angle θ of the front camera belongs to the interval [0°, 45°), S1 = 1.0, S2 = 0.0;

[0178] (2) If the azimuth angle θ of the front camera belongs to the interval [45°, 90°), S1 = 0.7, S2 = 0.3;

[0179] (3) If the azimuth angle θ of the front camera belongs to the interval [90°, 135°), S1 = 0.3, S2 = 0.7;

[0180] (4) If the azimuth angle θ of the front camera belongs to the interval [135°, 180°], S1 = 0.0, S2 = 1.0.

[0181] The above are just some examples of how the second-layer model can be implemented, and this application is not limited to them.

[0182] For the second-layer model, which uses a time-limited method with preset weights or weight mapping functions, only the other parts of the neural network model need to be trained during training, and the second-layer model does not need to be trained, which can improve training efficiency.

[0183] This application also provides a scene recognition device. Figure 9 A block diagram of a scene recognition device according to an embodiment of this application is shown. Figure 9As shown, the device may include:

[0184] An image acquisition module is used to acquire images of the same scene from multiple azimuth angles using multiple cameras, wherein the azimuth angle at which each camera acquires the image is the azimuth angle corresponding to the image, and the azimuth angle is the angle between the direction vector of each camera when acquiring the image and the unit gravity vector;

[0185] The scene recognition module is used to identify the same scene based on the image and the azimuth angle corresponding to the image, and obtain the scene recognition result.

[0186] The scene recognition device in this application uses multiple cameras to capture images of the same scene from multiple azimuth angles, and combines multiple images and the azimuth angle corresponding to each image to recognize the scene. Since more comprehensive scene information is obtained, the accuracy of image-based scene recognition can be improved, solving the problem of limited field of view and shooting angle for single-camera scene recognition, and making the recognition more accurate.

[0187] In one possible implementation, the scene recognition module includes:

[0188] An azimuth feature extraction module is used to extract the azimuth features corresponding to the image from the azimuth angle corresponding to the image.

[0189] A scene recognition model is used to identify the same scene based on the image and the azimuth features corresponding to the image, and to obtain a scene recognition result, wherein the scene recognition model is a neural network model.

[0190] In one possible implementation, the scene recognition model includes multiple pairs of first feature extraction layers and first-layer models. Each pair of first feature extraction layers and first-layer models is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of the azimuth angle to obtain a first recognition result. The azimuth angle corresponding to the image of the azimuth angle is the azimuth angle corresponding to the first recognition result. The first feature extraction layer is used to extract features from the image of the azimuth angle to obtain a feature vector. The first-layer model is used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle. The scene recognition model further includes a second-layer model, which is used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

[0191] The scene recognition device in this application adopts a two-layer scene recognition model, collects images from multiple angles and combines the azimuth angles of the images, and uses a competition mechanism to perform scene recognition. It considers both local and global features, which can improve the accuracy of scene recognition in a user-unnoticed manner and reduce misjudgments.

[0192] In one possible implementation, the first feature extraction layer is used to extract features from the image of the azimuth angle to obtain multiple feature vectors; the first layer model is used to calculate a first weight corresponding to each of the multiple feature vectors based on the azimuth angle corresponding to the image of the azimuth angle; the first layer model is used to obtain the first recognition result based on each feature vector and the first weight corresponding to each feature vector.

[0193] In one possible implementation, the second-layer model is used to calculate a second weight of the first recognition result based on the azimuth angle corresponding to the first recognition result; the second-layer model is used to obtain the scene recognition result based on the first recognition result and the second weight of the first recognition result.

[0194] In one possible implementation, the second-layer model pre-defines a third weight corresponding to each of the first recognition results, and the second-layer model is used to obtain the scene recognition result based on the first recognition result and the third weight corresponding to the first recognition result.

[0195] In one possible implementation, the second-layer model is used to determine the fourth weight corresponding to each of the first identification results based on the azimuth angle and a preset rule; wherein the preset rule is a weight set based on the azimuth angle, different azimuth angles correspond to different weight sets, and each weight set includes the fourth weight corresponding to each of the first identification results; the second-layer model is used to obtain the scene identification result based on the first identification result and the fourth weight corresponding to the first identification result.

[0196] In one possible implementation, the device further includes: an azimuth angle acquisition module, used to acquire the acceleration of the gravity sensor on the coordinate axes of the three-dimensional Cartesian coordinate system corresponding to each camera acquiring the image, and to obtain the direction vector of each camera acquiring the image; wherein, the three-dimensional Cartesian coordinate system corresponding to each camera acquiring the image has each camera as its origin, the z-direction is the direction along which the camera captures the image, and x and y are directions perpendicular to the z-direction, and the planes containing x and y are perpendicular to the z-direction; the azimuth angle is calculated based on the direction vector and the gravity unit vector.

[0197] In one possible implementation, the apparatus further includes: an image preprocessing module for preprocessing the image; wherein the preprocessing includes one or more combinations of the following processes: converting image format, converting image channels, unifying image size, and image normalization; converting image format refers to converting a color image to a black and white image; converting image channels refers to converting an image to the red, green, and blue (RGB) channels; unifying image size refers to adjusting multiple images to have the same length and width; and image normalization refers to normalizing the pixel values ​​of the image.

[0198] In one possible implementation, the device further includes a result output module for outputting scene recognition results.

[0199] Figure 10 This diagram illustrates the structure of a terminal device according to an embodiment of the present application. Taking a mobile phone as an example, the terminal device is shown below. Figure 10 A structural schematic diagram of mobile phone 200 is shown.

[0200] Mobile phone 200 may include a processor 210, an external memory interface 220, an internal memory 221, a USB interface 230, a charging management module 240, a power management module 241, a battery 242, antenna 1, antenna 2, a mobile communication module 251, a wireless communication module 252, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone jack 270D, a sensor module 280, buttons 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a SIM card interface 295, etc. The sensor module 280 may include a gyroscope sensor 280A, an accelerometer sensor 280B, a proximity sensor 280G, a fingerprint sensor 280H, and a touch sensor 280K (of course, mobile phone 200 may also include other sensors, such as a temperature sensor, a pressure sensor, a proximity sensor, a magnetic sensor, an ambient light sensor, a barometric pressure sensor, a bone conduction sensor, etc., not shown in the figure).

[0201] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the mobile phone 200. In other embodiments of this application, the mobile phone 200 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0202] Processor 210 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller may serve as the central nervous system and command center of the mobile phone 200. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0203] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0204] The processor 210 can run the scene recognition method provided in the embodiments of this application, so as to combine multiple images and the azimuth angle corresponding to each image to identify the scene, obtain more comprehensive scene information, improve the accuracy of image-based scene recognition, and solve the problem of limited field of view and shooting angle of single-camera scene recognition, resulting in more accurate recognition. The processor 210 may include different devices. For example, when integrating a CPU and a GPU, the CPU and GPU can cooperate to execute the scene recognition method provided in the embodiments of this application. For example, some algorithms in the scene recognition method are executed by the CPU, and other algorithms are executed by the GPU to obtain faster processing efficiency.

[0205] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, mobile phone 200 may include one or N displays 294, where N is a positive integer greater than 1. Display screen 294 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces (GUIs). For example, display screen 294 can display photos, videos, web pages, or documents, etc. As another example, display screen 294 can display a graphical user interface. The graphical user interface (GUI) includes a status bar, a hideable navigation bar, a time and weather widget, and application icons, such as browser icons. The status bar includes the carrier name (e.g., China Mobile), mobile network (e.g., 4G), time, and remaining battery power. The navigation bar includes a back button icon, a home button icon, and a forward button icon. Furthermore, it is understood that in some embodiments, the status bar may also include Bluetooth icons, Wi-Fi icons, and external device icons. It is also understood that in other embodiments, the GUI may include a Dock bar, which may include frequently used application icons. When the processor 210 detects a touch event from a user's finger (or stylus, etc.) on an application icon, in response to the touch event, it opens the user interface of the application corresponding to that application icon and displays the application's user interface on the display 294.

[0206] In this embodiment of the application, the display screen 294 can be an integral flexible display screen, or it can be a splicing display screen composed of two rigid screens and a flexible screen located between the two rigid screens.

[0207] Camera 293 (front-facing camera and rear-facing camera, both of which may include one or more cameras) is used to capture still images or videos. Typically, camera 293 may include a photosensitive element such as a lens assembly and an image sensor. The lens assembly includes multiple lenses (convex or concave lenses) for collecting light signals reflected from the object to be photographed and transmitting the collected light signals to the image sensor. The image sensor generates an original image of the object to be photographed based on the light signals. In the embodiments of this application, images of the same scene are captured from multiple azimuth angles using multiple cameras. This allows for the combination of multiple images and the corresponding azimuth angle of each image to identify the scene, obtaining more comprehensive scene information, improving the accuracy of image-based scene recognition, and solving the problem of limited field of view and shooting angle for single-camera scene recognition, resulting in more accurate recognition.

[0208] Internal memory 221 can be used to store computer executable program code, which includes instructions. Processor 210 executes various functional applications and data processing of mobile phone 200 by running the instructions stored in internal memory 221. Internal memory 221 may include a program storage area and a data storage area. The program storage area can store the operating system, application code (such as camera applications, WeChat applications, etc.), etc. The data storage area can store data created during the use of mobile phone 200 (such as images and videos captured by the camera application, etc.).

[0209] The internal memory 221 may also store one or more computer programs 1310 corresponding to the scene recognition method provided in the embodiments of this application. The one or more computer programs 1304 are stored in the memory 221 and configured to be executed by the one or more processors 210. Each computer program 1310 includes instructions that can be used to execute various steps in the scene recognition method provided in the embodiments of this application. The computer program 1310 may include: an image acquisition module for acquiring images of the same scene from multiple azimuth angles using multiple cameras; a scene recognition module for recognizing the same scene based on the images and the corresponding azimuth angles, and obtaining a scene recognition result; an azimuth angle acquisition module for acquiring the acceleration of the gravity sensor on the coordinate axes of the three-dimensional rectangular coordinate system corresponding to each camera acquiring the image, obtaining the direction vector of each camera acquiring the image, and calculating the azimuth angle based on the direction vector and the gravity unit vector; and an image preprocessing module for preprocessing the images.

[0210] In addition, the internal memory 221 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0211] Of course, the code for the scene recognition method provided in this application embodiment can also be stored in external memory. In this case, the processor 210 can run the code for the scene recognition method stored in external memory through the external memory interface 220.

[0212] The functions of sensor module 280 are described below.

[0213] The gyroscope sensor 280A can be used to determine the motion attitude of the mobile phone 200. In some embodiments, the gyroscope sensor 280A can determine the angular velocity of the mobile phone 200 about three axes (i.e., the x, y, and z axes). That is, the gyroscope sensor 280A can be used to detect the current motion state of the mobile phone 200, such as whether it is shaking or stationary.

[0214] When the display screen in this embodiment is a foldable screen, the gyroscope sensor 280A can be used to detect folding or unfolding operations performed on the display screen 294. The gyroscope sensor 280A can report the detected folding or unfolding operations as events to the processor 210 to determine the folding or unfolding state of the display screen 294.

[0215] Accelerometer 280B can detect the magnitude of acceleration of mobile phone 200 in various directions (generally three axes). That is, gyroscope sensor 280A can be used to detect the current motion state of mobile phone 200, such as whether it is shaking or stationary. When the display screen in this embodiment is a foldable screen, accelerometer 280B can be used to detect folding or unfolding operations acting on display screen 294. Accelerometer 280B can report the detected folding or unfolding operations as events to processor 210 to determine the folding or unfolding state of display screen 294.

[0216] In the embodiments of this application, the terminal device can obtain the acceleration on the coordinate axis of the three-dimensional Cartesian coordinate system corresponding to each camera capturing the image through the accelerometer 280B, obtain the direction vector of each camera capturing the image, and calculate the azimuth angle based on the direction vector and the gravity unit vector.

[0217] The proximity sensor 280G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The mobile phone emits infrared light outward through the LED. The mobile phone uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the mobile phone. When insufficient reflected light is detected, the mobile phone can determine that there is no object near the mobile phone. When the display screen in this embodiment is a foldable screen, the proximity sensor 280G may be disposed on the first screen of the foldable display screen 294. The proximity sensor 280G may detect the size of the folding angle or unfolding angle between the first screen and the second screen based on the optical path difference of the infrared signal.

[0218] The gyroscope sensor 280A (or accelerometer 280B) can send the detected motion state information (such as angular velocity) to the processor 210. The processor 210 determines whether the current state is handheld or tripod based on the motion state information (for example, if the angular velocity is not 0, it means that the phone 200 is handheld).

[0219] The fingerprint sensor 280H is used to collect fingerprints. The mobile phone 200 can use the collected fingerprint characteristics to achieve fingerprint unlocking, app access lock, fingerprint photography, fingerprint answering of incoming calls, etc.

[0220] Touch sensor 280K, also known as a "touch panel," can be located on display screen 294. The touch sensor 280K and display screen 294 together form a touchscreen, also known as a "touchscreen." Touch sensor 280K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 294. In other embodiments, touch sensor 280K may also be located on the surface of mobile phone 200, in a different position than display screen 294.

[0221] For example, the display screen 294 of the mobile phone 200 displays the main interface, which includes icons for multiple applications (such as a camera application, a WeChat application, etc.). The user taps the camera application icon on the main interface using the touch sensor 280K, triggering the processor 210 to launch the camera application and open the camera 293. The display screen 294 displays the camera application's interface, such as the viewfinder. The display screen 294 can also be used to display scene recognition results.

[0222] The wireless communication function of mobile phone 200 can be implemented through antenna 1, antenna 2, mobile communication module 251, wireless communication module 252, modem processor and baseband processor.

[0223] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 200 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0224] The mobile communication module 251 can provide solutions for wireless communication applications including 2G / 3G / 4G / 5G on the mobile phone 200. The mobile communication module 251 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 251 can receive electromagnetic waves via the antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to the modem processor for demodulation. The mobile communication module 251 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna 1. In some embodiments, at least some functional modules of the mobile communication module 251 may be housed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 251 and at least some modules of the processor 210 may be housed in the same device.

[0225] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 270A, receiver 270B, etc.) or displays images or videos through the display screen 294. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 210 and may be housed in the same device as the mobile communication module 251 or other functional modules.

[0226] The wireless communication module 252 can provide solutions for wireless communication applications on the mobile phone 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 252 can be one or more devices integrating at least one communication processing module. The wireless communication module 252 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 210. The wireless communication module 252 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2. In this embodiment, the wireless communication module 252 is used to transmit data between other terminal devices under the control of processor 210.

[0227] In addition, mobile phone 200 can implement audio functions through audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and application processor, such as music playback and recording. Mobile phone 200 can receive key input 290, generating key signal inputs related to user settings and function control. Mobile phone 200 can use motor 291 to generate vibration alerts (such as vibration alerts for incoming calls). The indicator 292 in mobile phone 200 can be an indicator light, used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. The SIM card interface 295 in mobile phone 200 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to achieve contact and separation with mobile phone 200.

[0228] It should be understood that in practical applications, mobile phone 200 can include more than Figure 10 The number of components shown is not limited in this application embodiment. The illustrated mobile phone 200 is merely an example, and the mobile phone 200 may have more or fewer components than shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0229] The software system of a terminal device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application uses the layered architecture Android system as an example to illustrate the software structure of the terminal device.

[0230] An embodiment of this application provides a scene recognition device, including: a processor and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions.

[0231] Embodiments of this application provide a non-volatile computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the above-described method.

[0232] Embodiments of this application provide a computer program product including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.

[0233] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital video disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing.

[0234] The computer-readable program instructions or code described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0235] The computer program instructions used to perform the operations of this application may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information from computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of this application.

[0236] Various aspects of this application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0237] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0238] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0239] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.

[0240] It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented using hardware (such as circuits or ASICs (Application Specific Integrated Circuits)) that performs the corresponding function or action, or using a combination of hardware and software, such as firmware.

[0241] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, disclosure, and appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0242] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A scene recognition method characterized by, The method includes: The terminal device acquires images of the same scene from multiple azimuth angles using multiple cameras. The azimuth angle at which each camera acquires the image is the azimuth angle corresponding to the image. The azimuth angle is the angle between the direction vector of each camera when acquiring the image and the unit vector of gravity. The terminal device uses a scene recognition model to identify the same scene based on the image and the azimuth angle corresponding to the image, and obtains the scene recognition result; The scene recognition model includes multiple pairs of first feature extraction layers and first-layer models. Each pair of first feature extraction layers and first model layers is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of the azimuth angle to obtain a first recognition result; Wherein, the azimuth angle corresponding to the image of the azimuth angle is the azimuth angle corresponding to the first recognition result, the first feature extraction layer is used to extract the features of the image of the azimuth angle to obtain a feature vector, and the first layer model is used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle; The scene recognition model further includes a second-layer model, which is used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

2. The method of claim 1, wherein, The terminal device identifies the same scene based on the image and the corresponding azimuth angle, and obtains a scene recognition result, including: The terminal device extracts the azimuth feature corresponding to the image from the azimuth angle corresponding to the image; The same scene is identified using a scene recognition model based on the image and the corresponding azimuth features, to obtain a scene recognition result. The scene recognition model is a neural network model.

3. The method according to claim 1, characterized in that, The first feature extraction layer is used to extract features from the image of the azimuth angle to obtain multiple feature vectors; The first layer model is used to calculate the first weight of each feature vector among the multiple feature vectors based on the azimuth angle corresponding to the image of the azimuth angle; The first layer model is used to obtain the first recognition result based on each feature vector and the first weight corresponding to each feature vector.

4. The method according to claim 1, characterized in that, The second layer model is used to calculate the second weight of the first identification result based on the azimuth angle corresponding to the first identification result; The second-layer model is used to obtain the scene recognition result based on the first recognition result and the second weight of the first recognition result.

5. The method according to claim 1, characterized in that, The second-layer model pre-defines a third weight corresponding to each of the first recognition results. The second-layer model is used to obtain the scene recognition result based on the first recognition result and the third weight corresponding to the first recognition result.

6. The method according to claim 1, characterized in that, The second layer model is used to determine the fourth weight corresponding to each first identification result based on the azimuth angle and the preset rules; wherein, the preset rules are weight reassemblies set according to the azimuth angle, different azimuth angles correspond to different weight reassemblies, and each weight reassembly includes the fourth weight corresponding to each first identification result; The second-layer model is used to obtain the scene recognition result based on the first recognition result and the fourth weight corresponding to the first recognition result.

7. The method of claim 1, wherein, The method further includes: The terminal device obtains the acceleration of the gravity sensor on the coordinate axis of the three-dimensional Cartesian coordinate system when each camera captures the image, and obtains the direction vector of each camera when capturing the image; In this system, the three-dimensional Cartesian coordinate system corresponding to each camera when capturing the image has each camera as its origin, the z-direction is the direction along which the camera captures the image, and the x and y directions are perpendicular to the z-direction, and the planes containing the x and y directions are perpendicular to the z-direction. The azimuth angle is calculated based on the direction vector and the gravity unit vector.

8. The method of claim 2, wherein, Before using a scene recognition model to identify the same scene based on the image and the corresponding azimuth features, the method further includes: The terminal device preprocesses the image; The preprocessing includes one or more of the following processes in combination: image format conversion, image channel conversion, image size unification, and image normalization. Image format conversion refers to converting a color image to a black and white image. Image channel conversion refers to converting an image to the red, green, and blue (RGB) channels. Image size unification refers to adjusting the length and width of multiple images to be the same. Image normalization refers to normalizing the pixel values ​​of an image.

9. A scene recognition apparatus characterized by comprising: The device includes: An image acquisition module is used to acquire images of the same scene from multiple azimuth angles using multiple cameras, wherein the azimuth angle at which each camera acquires the image is the azimuth angle corresponding to the image, and the azimuth angle is the angle between the direction vector of each camera when acquiring the image and the unit gravity vector; The scene recognition module is used to identify the same scene based on the image and the azimuth angle corresponding to the image using a scene recognition model, and obtain the scene recognition result; The scene recognition model includes multiple pairs of first feature extraction layers and first-layer models. Each pair of first feature extraction layers and first model layers is used to process an image of an azimuth angle and the azimuth angle corresponding to the image of the azimuth angle to obtain a first recognition result; Wherein, the azimuth angle corresponding to the image of the azimuth angle is the azimuth angle corresponding to the first recognition result, the first feature extraction layer is used to extract the features of the image of the azimuth angle to obtain a feature vector, and the first layer model is used to obtain the first recognition result based on the feature vector and the azimuth angle corresponding to the image of the azimuth angle; The scene recognition model further includes a second-layer model, which is used to obtain the scene recognition result based on the first recognition result and the azimuth angle corresponding to the first recognition result.

10. The apparatus of claim 9, wherein, The scene recognition module includes: An azimuth feature extraction module is used to extract the azimuth features corresponding to the image from the azimuth angle corresponding to the image. A scene recognition model is used to identify the same scene based on the image and the azimuth features corresponding to the image, and to obtain a scene recognition result, wherein the scene recognition model is a neural network model.

11. The apparatus according to claim 9, characterized in that, The first feature extraction layer is used to extract features from the image of the azimuth angle to obtain multiple feature vectors; The first layer model is used to calculate the first weight of each feature vector among the multiple feature vectors based on the azimuth angle corresponding to the image of the azimuth angle; The first layer model is used to obtain the first recognition result based on each feature vector and the first weight corresponding to each feature vector.

12. The apparatus according to claim 9, characterized in that, The second layer model is used to calculate the second weight of the first identification result based on the azimuth angle corresponding to the first identification result; The second-layer model is used to obtain the scene recognition result based on the first recognition result and the second weight of the first recognition result.

13. The apparatus according to claim 9, characterized in that, The second-layer model pre-defines a third weight corresponding to each of the first recognition results. The second-layer model is used to obtain the scene recognition result based on the first recognition result and the third weight corresponding to the first recognition result.

14. The apparatus according to claim 9, characterized in that, The second-layer model is used to determine the fourth weight corresponding to each of the first identification results based on the azimuth angle and preset rules; The preset rule is a weighted reassembly set according to the azimuth angle. Different azimuth angles correspond to different weighted reassemblies. Each weighted reassembly includes a fourth weight corresponding to each first identification result. The second-layer model is used to obtain the scene recognition result based on the first recognition result and the fourth weight corresponding to the first recognition result.

15. The apparatus of claim 9, wherein, The device further includes: The azimuth angle acquisition module is used to obtain the acceleration of the gravity sensor on the coordinate axis of the three-dimensional rectangular coordinate system when each camera acquires the image, and to obtain the direction vector when each camera acquires the image; In this system, the three-dimensional Cartesian coordinate system corresponding to each camera when acquiring the image has each camera as its origin, the z-direction as the direction along which the camera captures the image, and the x and y directions as perpendicular to the z-direction, and the plane containing the x and y directions is perpendicular to the z-direction; the azimuth angle is calculated based on the direction vector and the gravity unit vector.

16. The apparatus of claim 10, wherein, The device further includes: An image preprocessing module is used to preprocess the image; The preprocessing includes one or more of the following processes in combination: image format conversion, image channel conversion, image size unification, and image normalization. Image format conversion refers to converting a color image to a black and white image. Image channel conversion refers to converting an image to the red, green, and blue (RGB) channels. Image size unification refers to adjusting the length and width of multiple images to be the same. Image normalization refers to normalizing the pixel values ​​of an image.

17. A computer program product comprising computer-readable code, wherein when the computer-readable code is run in an electronic device, a processor in the electronic device performs the method of any one of claims 1-8.

18. A non-transitory computer readable storage medium having stored thereon computer program instructions, wherein, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1-8.

19. A terminal device, comprising: include: processor; Memory used to store processor-executable instructions; The processor is configured to implement the method of any one of claims 1-8 when executing the instructions.

Citation Information

Patent Citations

  • Method and terminal for scanning identification based on wide angle view

    CN109117693A