Mobile device identification method, apparatus, device, storage medium and program product

By generating multi-view images and detecting global and local features, the problem of information loss in monocular monitoring systems is solved, improving the accuracy and adaptability of vehicle recognition.

CN122435563APending Publication Date: 2026-07-21ZHE JIANG SHEN XIANG ZHI NENG KE JI YOU XIAN GONG SI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHE JIANG SHEN XIANG ZHI NENG KE JI YOU XIAN GONG SI
Filing Date
2026-03-13
Publication Date
2026-07-21

Smart Images

  • Figure CN122435563A_ABST
    Figure CN122435563A_ABST
Patent Text Reader

Abstract

The application relates to a mobile device identification method, device, equipment, storage medium and program product, wherein the method comprises the following steps: acquiring a monocular photographic image of a mobile device; generating a multi-view image of the mobile device based on the monocular photographic image; detecting a global feature and a local feature of the mobile device in the multi-view image; and outputting an identification result of the mobile device based on the global feature and the local feature. The application enhances the identification accuracy of the mobile device in a monocular photographic scene, and improves the robustness and adaptability of the identification method in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a mobile device identification method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the deepening of smart city construction, monocular surveillance cameras are increasingly widely used in scenarios such as auto insurance claims, intelligent parking management, and traffic monitoring. Compared with multi-camera systems, monocular surveillance cameras have significant advantages such as lower cost, simpler deployment, and easier maintenance, making them more competitive in large-scale deployment scenarios. Taking a smart parking lot as an example, a medium-sized parking lot typically requires the deployment of dozens of monitoring points. If a multi-camera system is used, the equipment cost will increase by 3-5 times, and the installation complexity will increase significantly. In contrast, a monocular surveillance system can achieve basic vehicle model recognition functions at a lower cost.

[0003] However, monocular monitoring systems face core challenges in practical applications, such as the lack of multi-view information and fixed viewing angles. For example, in smart parking scenarios, because monitoring cameras are usually installed at a relatively high position, such as 5-8 meters, vehicles are far from the camera, resulting in vehicles occupying a very small proportion of the image, and there are also serious problems with occlusion and lighting variations. Monocular monitoring images often only provide two-dimensional images from a top-down perspective, and cannot directly obtain the true size and three-dimensional structural information of the vehicle, leading to very low vehicle recognition accuracy. Summary of the Invention

[0004] To overcome the problems existing in related technologies, this application provides a mobile device identification method, apparatus, device, storage medium and program product, which enhances the identification accuracy of mobile devices in monocular photography scenarios and improves the robustness and adaptability of the identification method in complex scenarios.

[0005] In a first aspect, embodiments of this application provide a mobile device identification method, comprising: acquiring a monocular photographic image of the mobile device; generating a multi-view image of the mobile device based on the monocular photographic image; detecting global features and local features of the mobile device in the multi-view image; and outputting an identification result of the mobile device based on the global features and the local features.

[0006] Secondly, embodiments of this application provide a mobile device identification device, comprising:

[0007] The acquisition module is used to acquire monocular photographic images from mobile devices;

[0008] A generation module is used to generate multi-view images of the mobile device based on the monocular photographic image;

[0009] The detection module is used to detect the global and local features of the mobile device in the multi-view images;

[0010] The output module is used to output the recognition result of the mobile device based on the global features and the local features.

[0011] Thirdly, embodiments of this application provide an electronic device, including:

[0012] At least one processor; and

[0013] A memory that is communicatively connected to the at least one processor;

[0014] The memory stores instructions executable by the at least one processor, which is configured to execute the instructions to implement the method described in any of the above aspects.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method described in any of the above aspects.

[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the above aspects.

[0017] Sixthly, embodiments of this application provide a cloud device, including:

[0018] At least one processor; and

[0019] A memory that is communicatively connected to the at least one processor;

[0020] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, cause the cloud device to perform the method described in any of the above aspects.

[0021] In a seventh aspect, embodiments of this application provide a mobile device identification system, including a server and a terminal, wherein the terminal and the server interact via data to implement the method described in any of the above aspects.

[0022] The mobile device identification method, apparatus, device, storage medium, and program product provided in this application generate multi-view images from monocular photographic images, thus overcoming the inherent defects of information loss under monocular perspective, simulating the visual effect of multi-angle observation, and providing more comprehensive spatial and structural information. On this basis, by detecting global and local features in multi-view images, global features can capture the overall shape and structural outline of the mobile device, while local features provide refined representation of specified areas. The collaborative extraction of global and local features makes the feature expression richer and more discriminative. The identification results are output based on global and local features, which enhances the identification accuracy of mobile devices in monocular photography scenarios and improves the robustness and adaptability of the identification method in complex scenarios. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are some embodiments of this application, and that those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0024] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0025] Figure 2 This is a schematic diagram illustrating an application scenario of a mobile device identification system provided in an embodiment of this application;

[0026] Figure 3 A flowchart illustrating a mobile device identification method provided in an embodiment of this application;

[0027] Figure 4 A schematic diagram illustrating the spatial relationship between a monocular camera and a vehicle, provided as an embodiment of this application;

[0028] Figure 5 A schematic diagram illustrating the positional relationship between a real lens and a virtual lens, provided in an embodiment of this application;

[0029] Figure 6 This is a schematic diagram illustrating the functional principle of a feature retrieval module provided in an embodiment of this application;

[0030] Figure 7 A schematic diagram illustrating a process for determining initial identification results based on a confidence selection mechanism, provided for an embodiment of this application;

[0031] Figure 8 A flowchart illustrating a mobile device identification method provided in an embodiment of this application;

[0032] Figure 9 This is a schematic diagram of the structure of a mobile device identification device provided in an embodiment of this application;

[0033] Figure 10 This is a schematic diagram of the structure of a cloud device provided in an embodiment of this application.

[0034] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0035] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0036] In this article, the term "and / or" is used to describe the relationship between related objects. Specifically, it means that there can be three kinds of relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, or B exists alone.

[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0038] To clearly describe the technical solutions of the embodiments of this application, the terms involved in this application are first defined as follows:

[0039] ID: Identifier, identification code, identifier.

[0040] YOLO: You Only Look Once is a single-stage object detection model architecture.

[0041] Faster R-CNN: Faster Region-based Convolutional Neural Networks.

[0042] U-Net: Convolutional Networks for Biomedical Image Segmentation.

[0043] CGAN: Conditional Generative Adversarial Network.

[0044] PatchGAN: Patch Generative Adversarial Network, a local image patch discriminator.

[0045] FC: Fully Connected Layer.

[0046] STN: Spatial Transformer Networks.

[0047] AIGC: Artificial Intelligence Generated Content.

[0048] MobileNet: Lightweight terminal convolutional network.

[0049] BCE: Binary Cross-Entropy Loss.

[0050] ResNet50: Residual Network with 50 layers.

[0051] SNR: Signal-to-Noise Ratio.

[0052] COCO: Common Objects in Context, a large-scale, multi-task image dataset for computer vision research.

[0053] Faster R-CNN: Faster Region-based Convolutional Neural Networks.

[0054] SSD: Single Shot MultiBox Detector.

[0055] Transformer: A machine learning model.

[0056] The mobile device identification method of this application embodiment can be applied to any field that requires automatic target identification.

[0057] With the deepening of smart city construction, monocular surveillance cameras are increasingly widely used in scenarios such as auto insurance claims, intelligent parking management, and traffic monitoring. Compared with multi-camera systems, monocular surveillance cameras have significant advantages such as lower cost, simpler deployment, and easier maintenance, making them more competitive in large-scale deployment scenarios. Taking a smart parking lot as an example, a medium-sized parking lot typically requires the deployment of dozens of monitoring points. If a multi-camera system is used, the equipment cost will increase by 3-5 times, and the installation complexity will increase significantly. In contrast, a monocular surveillance system can achieve basic vehicle model recognition functions at a lower cost.

[0058] However, monocular monitoring systems face core challenges in practical applications, such as the lack of multi-view information and fixed viewing angles. For example, in car insurance claims scenarios, it is necessary to accurately identify the model, year, and color of the accident vehicle. However, monocular monitoring images often only provide two-dimensional images from a top-down perspective, failing to directly obtain the vehicle's true size and three-dimensional structural information, resulting in a significant drop in vehicle model recognition accuracy. In smart parking scenarios, because monitoring cameras are typically installed at a height of 5-8 meters, vehicles are far from the camera, occupying only 10%-20% of the image. Furthermore, severe occlusion and lighting variations exist, causing the accuracy of general vehicle model recognition models to typically fall below 60%, failing to meet actual business needs.

[0059] To address at least one of the aforementioned problems, embodiments of this application provide a mobile device recognition scheme. By generating multi-view images from monocular photographic images, it overcomes the inherent information loss inherent in monocular perspectives, simulates the visual effects of multi-angle observation, and provides more comprehensive spatial and structural information. Furthermore, by detecting global and local features in the multi-view images, global features capture the overall shape and structural outline of the mobile device, while local features provide refined representations of specific regions. The collaborative extraction of global and local features results in richer and more discriminative feature representation. The recognition results are output based on both global and local features, enhancing the accuracy of mobile device recognition in monocular photography scenarios and improving the robustness and adaptability of the recognition method in complex scenes.

[0060] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.

[0061] like Figure 1 As shown, this embodiment provides an electronic device 1, including: at least one processor 11 and a memory 12. Figure 1 Taking a processor as an example, the processor 11 and the memory 12 are connected via a bus 10. The memory 12 stores instructions that can be executed by the processor 11. The instructions are executed by the processor 11 to enable the electronic device 1 to perform all or part of the process of the method in the following embodiments, so as to enhance the recognition accuracy of mobile devices in monocular photography scenarios and improve the robustness and adaptability of the recognition method in complex scenarios.

[0062] In one embodiment, the electronic device 1 may be an in-vehicle device, such as an in-vehicle controller, or a mobile phone, tablet computer, laptop computer, desktop computer, or a large computing system composed of multiple computers.

[0063] Figure 2 This is a schematic diagram illustrating an application scenario 200 of a mobile device identification system provided in an embodiment of this application. For example... Figure 2 As shown, the system includes: a server 210 and a terminal 220, wherein:

[0064] Server 210 can be a data center providing mobile device identification services, such as a vehicle monitoring service data center. In a real-world scenario, a vehicle monitoring service data center may have multiple servers 210. Figure 2 Taking a single server (210) as an example.

[0065] Terminal 220 can be an electronic device that interacts with the vehicle monitoring service data center, such as a computer, mobile phone, tablet, or other device used to access the vehicle monitoring service data center. There can also be multiple terminals 220. Figure 2 The following is an example using a single terminal 220.

[0066] The mobile device identification scheme of this application embodiment can be deployed on server 210, on terminal 220, or partially on server 210 and partially on terminal 220. The choice can be made based on actual needs in a real-world scenario, and this embodiment does not impose any limitations.

[0067] When the mobile device identification scheme is deployed entirely or partially on server 210, the call interface can be opened to terminal 220 to provide algorithm support to terminal 220.

[0068] The method provided in this application embodiment can be implemented by electronic device 1 executing corresponding software code, and is achieved through data interaction with a server. Electronic device 1 can be a local terminal device. When the method runs on a server, it can be implemented and executed based on a cloud interaction system, which includes a server and client devices.

[0069] In one possible implementation, the method provided in this application provides a graphical user interface through a terminal device, wherein the terminal device may be the aforementioned local terminal device or a client device in the aforementioned cloud interaction system.

[0070] Please refer to Figure 3 This is a mobile device identification method according to an embodiment of this application. The method can be derived from... Figure 1 The electronic device 1 shown is used to perform this action and can be applied to... Figure 2 In the mobile device recognition application scenario shown, the aim is to enhance the recognition accuracy of mobile devices in monocular photography scenarios and improve the robustness and adaptability of the recognition method in complex scenarios. This embodiment takes a terminal as the execution end as an example, and the method includes the following steps:

[0071] Step 301: Acquire a monocular photographic image from the mobile device.

[0072] In this step, the mobile device can be a vehicle or a movable mechanical device. The monocular photographic image can be an image captured by a monocular camera on the mobile device, such as an image of a vehicle captured by a monocular camera in a parking lot.

[0073] Step 302: Generate multi-view images of the mobile device based on the monocular photographic images.

[0074] In this step, by generating multi-view images from monocular photographic images, the inherent defects of missing information under monocular view can be made up for, simulating the visual effect of multi-angle observation, and thus providing mobile devices with more comprehensive spatial and structural information.

[0075] In one embodiment, step 302 may specifically include: determining the target camera used when capturing the monocular photographic image, and obtaining the pre-calibrated camera parameters of the target camera. Based on the camera parameters, determining the 3D information of the mobile device in the real scene and the shooting angle corresponding to the monocular photographic image. Offsetting the shooting angle according to a preset angle offset to obtain the target angle of at least one virtual lens. Based on the monocular photographic image, camera parameters, target angle, and 3D information, generating virtual perspective images of the mobile device under different target angles. The multi-view images include monocular photographic images and virtual perspective images.

[0076] In this embodiment, camera parameters include, but are not limited to, camera intrinsic parameters and camera extrinsic parameters, wherein the camera intrinsic parameters may include focal length. Principal point coordinates ( , ) and camera mounting tilt angle Camera extrinsic parameters can include camera mounting position and pose information, such as camera mounting height. To address the issue of missing appearance information from multiple perspectives in monocular cameras, a geometrically constrained multi-view generation network can be constructed based on camera intrinsic and extrinsic parameters to extend monocular photographic images into multi-view images.

[0077] The target camera refers to the camera used to capture monocular images. Utilizing the pre-calibrated camera parameters of the target camera provides a precise geometric constraint basis for the multi-view image generation process, ensuring the measurement accuracy and physical realism of subsequent steps. Based on the camera parameters of the target camera, the 3D information of the mobile device in the real-world scene and the original shooting perspective are determined, enabling the estimation and recovery of 3D spatial information from a monocular 2D image, providing spatial data support for perspective transformation. By offsetting the original shooting perspective according to a preset perspective offset, the target perspective of at least one virtual lens is obtained, and then virtual images originating from different viewing directions are synthesized, simulating the effect of multi-camera surround shooting. Then, based on the original monocular image, camera parameters, target perspective, and the 3D information of the mobile device, virtual perspective images are generated. Thus, without increasing any physical hardware costs or deployment complexity, a synthetic view covering multiple perspectives of the device can be derived from a single image. This provides a multi-angle image foundation for subsequent simultaneous detection of global and local features. It achieves the goal of compensating for hardware configuration limitations through software algorithms, improving the accuracy and reliability of mobile device identification, classification, or verification.

[0078] Optionally, the camera calibration process can occur during the target camera installation process. Using a checkerboard calibration board or a reference object of known size, multiple (e.g., at least 10) calibration images are taken from different angles. The Zhang calibration method is then used to analyze the projection positions of the checkerboard corner points in the images from different angles, and the target camera parameters are calculated. ,in, , Let f be the x-component of the camera focal length f in the camera coordinate system. Let f be the component of the camera focal length f along the y-axis in the camera coordinate system. , () are the coordinates of the main point. Set the camera's pitch angle, where h is the camera's installation height. Store the calibration results in a preset table as system configuration parameters. If there are multiple cameras, each camera can be numbered and stored in a preset table, which includes the mapping relationship between each camera number and its parameters. When inputting a monocular photographic image... At that time, the monocular photographic image was determined. The target camera number is determined, and then the camera parameters corresponding to the target camera number are read from the preset table.

[0079] In one embodiment, step 302, determining the 3D information of the mobile device in the real-world scene based on camera parameters in a monocular photographic image, includes: determining the top-down angle of the target camera relative to the mobile device when capturing the monocular photographic image; determining the lens depth of the target camera based on the top-down angle; and determining the position coordinates of the mobile device in the world coordinate system based on the camera parameters and the lens depth. The 3D information of the mobile device includes the lens depth and the position coordinates.

[0080] In this embodiment, the top-down view refers to the angle between the line connecting the mobile device to the target camera and the horizontal direction. Determining the top-down view of the target camera relative to the mobile device based on camera parameters provides a crucial geometric constraint for depth estimation. The lens depth is further determined based on the top-down view, achieving an accurate mapping from 2D image pixels to real-world distances. Combining the camera parameters and lens depth, the position coordinates of the mobile device in the world coordinate system are calculated, thus fully acquiring 3D information including depth and position. This achieves reliable and accurately measured 3D information recovery solely from monocular images, laying a solid geometric foundation for subsequent generation of virtual viewpoint images.

[0081] like Figure 4 The diagram shown illustrates the spatial relationship between a monocular camera and a vehicle according to an embodiment of this application. Taking the vehicle as the mobile device to be identified as an example, to ensure correct perspective, camera parameters are used. The target camera's downward angle β relative to the vehicle is derived using the following formula (1):

[0082] (1)

[0083] Furthermore, by combining β, the lens depth Z of the target camera can be derived, and the three-dimensional information of the vehicle can be obtained by calculating it using the following formula (2):

[0084]

[0085] in, These are the coordinates of the center of the vehicle detection bounding box in the monocular image, which can be obtained by performing vehicle detection on the monocular image. Let the principal point coordinates of the target camera be... Let (X, Y) be the lens depth of the target camera. Let (X, Y) be the vehicle's position coordinates in the world coordinate system. The vehicle's three-dimensional information can be represented as (X, Y, Z).

[0086] like Figure 5 The diagram shown illustrates the positional relationship between a real lens and a virtual lens according to an embodiment of this application. Taking a vehicle as the mobile device to be identified as an example, it illustrates the shooting angle corresponding to the monocular photographic image of the vehicle taken by the target camera. As a benchmark, assuming The target viewpoint of multiple virtual cameras is set by preset viewpoint offset. This can be achieved using the following formula (3):

[0087] (3)

[0088] in, For example, the preset view offset. More view offsets can be set according to needs. This avoids generating unpredictable appearances. Further adjustments can be made based on... Control the generation of vehicle exterior images from different shooting angles.

[0089] In one embodiment, step 302, which generates virtual perspective images of the mobile device under different target perspectives based on a monocular photographic image, camera parameters, target viewpoint, and 3D information, includes: inputting the monocular photographic image, camera parameters, target viewpoint, and 3D information into a preset generation model, so that the preset generation model outputs virtual perspective images of the mobile device under different target perspectives. The preset generation model is used to generate conditional vectors based on camera parameters and target viewpoint, and to generate geometric vectors based on 3D information. Based on the conditional vectors and geometric vectors, the monocular photographic image is reconstructed into a target feature map. Based on the 3D information, a spatial transformation is performed on the target feature map to generate virtual perspective images of the mobile device under different target perspectives.

[0090] In this embodiment, a structured pre-defined generation model is introduced to generate multi-view images from monocular images. The pre-defined generation model generates condition vectors based on camera parameters and the target viewpoint, and geometric vectors based on the 3D information of the mobile device. This achieves separate encoding and targeted representation of the two core types of information: "observation viewpoint conditions" and "object 3D structure," ensuring that subsequent processing can fully utilize prior knowledge of different properties. Then, under the joint constraints of the condition vectors and geometric vectors, the monocular photographic image is reconstructed into a target feature map. Within the feature space, based on the observation conditions of the target viewpoint and the object's true geometric structure, the original image features are guided and supplemented, preparing the feature level for viewpoint synthesis. Finally, based on the 3D information of the mobile device, a spatial transformation is performed on the target feature map, mapping it back to the image space. This ensures that each generated virtual viewpoint image strictly follows the true geometric relationships and perspective rules of the object in 3D space, thereby outputting a sequence of virtual viewpoint images with high geometric consistency and visual realism. This improves the visual quality and spatial accuracy of the generated virtual viewpoint images.

[0091] As mentioned above Figure 4 and Figure 5 Taking the vehicle monitoring scenario shown as an example, monocular camera images of the vehicle are used. Target camera parameters The target perspective of the virtual camera The system generates virtual viewpoint images using the vehicle's 3D information (X, Y, Z). Optionally, a multi-viewpoint generation network is built, such as one based on the CGAN backbone network. The overall network architecture can include a generator and a discriminator. The generator can be a U-Net-style encoder-decoder structure, and the discriminator can use a standard PatchGAN discriminator to determine whether an image is real or generated. The multi-viewpoint generation network is pre-trained to obtain a preset generation model. Alternatively, other generative model architectures such as StyleGAN (Style-based Generator Adversarial Network) or Diffusion Model can be used instead of CGAN, as long as they can achieve multi-viewpoint generation based on geometric constraints of monitoring parameters. The number of virtual viewpoints can be set to be greater.

[0092] Optionally, conditional information embedding can be added to the input layer of the multi-view generative network. Specifically, this can be done as follows:

[0093] First, set the camera parameters and the target perspective of virtual cameras The data is concatenated into a 6-dimensional vector. This 6-dimensional vector is then mapped to a 128-dimensional conditional vector Q through a two-layer fully connected network (e.g., a fully connected layer FC1, an activation function ReLU, and a second fully connected layer FC2). Similarly, the vehicle's 3D information (X, Y, Z) is mapped to a 128-dimensional geometric vector through another two-layer fully connected network. Then the condition vector Q and the geometric vector Concatenate them into a 256-dimensional vector. ,Will The image is mapped back to 128 dimensions through a fully connected layer, and then processed under the constraint of vector F. Remodel into one The target feature map, where H is the monocular image. High, W represents a monocular photographic image. The target feature map is injected into the encoder part of the generator via skip connections. Simultaneously, an STN (Spatial Transformation Network) module is introduced into the encoder part of the generator. This module receives the feature map of the current level and uses the affine transformation matrix calculated using the vehicle's 3D information (X, Y, Z) as a control signal to perform a spatial transformation on the feature map, adapting it to the target's viewpoint.

[0094] In one embodiment, before inputting the monocular photographic image, camera parameters, target viewpoint, and 3D information into the preset generation model, the method further includes: acquiring a sample image set containing sample mobile devices, wherein the sample images are labeled with corresponding sample camera parameters, pixel coordinates of the sample mobile devices in the sample images, and 3D coordinates of the sample mobile devices in the world coordinate system. The sample images are configured with labeled images of the sample mobile devices from different viewpoints. A preset conditional generative adversarial network is trained based on the sample image set to obtain the preset generation model. The conditional generative adversarial network includes a generator and a discriminator, and the generator's loss function includes pixel loss and size loss between the generated image and the real image.

[0095] In this embodiment, a supervised model training framework based on a conditional generative adversarial network (GAN) is pre-constructed to train a preset generative model. First, a sample image set is constructed, with each sample image labeled with its precisely corresponding camera parameters, pixel coordinates, and 3D world coordinates. Each sample is also equipped with a real-world labeled image from a different perspective, forming a comprehensive supervised learning dataset. This dataset provides sufficient and high-quality supervision signals for the model to learn the complex mapping relationship from monocular images to multiple virtual perspective images, ensuring the accuracy and reliability of the training process. Furthermore, by employing a conditional generative adversarial network as the training architecture, leveraging the adversarial game characteristic of its generator and discriminator, the trained preset generative model can output visually highly realistic virtual perspective images, improving the visual quality and realism of the generated results.

[0096] Optionally, the generator's loss function includes both pixel loss and size loss. Pixel loss directly constrains the generated image to align with the real label image at the pixel level, ensuring the accuracy of texture, color, and other details. Size loss specifically constrains the size ratio of the mobile device in the generated image to be consistent with the theoretical size calculated based on 3D information, thus ensuring the geometric and scale accuracy of the generated image. By designing a composite loss function, the model is forced to strictly adhere to the geometric constraints of 3D space while learning visual realism, enabling the final trained preset generation model to possess both excellent visual realism and geometric consistency.

[0097] As mentioned above Figure 4 and Figure 5 Taking the vehicle monitoring scenario shown as an example, the training sample image set can include real vehicle images retrieved from business operations and image data synthesized by AIGC. For each sample image, the corresponding sample camera parameters are provided. Each vehicle in the sample image needs to have its detection bounding box center coordinates labeled. And calculate the three-dimensional coordinates of the sample vehicle. In addition, for each sample image, multiple label images with different virtual perspectives can be generated. These label images can be rendered using 3D modeling software or obtained by matching existing multi-view image sets.

[0098] During training, the loss function of this conditional generative adversarial network can include the following three parts: adversarial loss, pixel loss, and size loss. The adversarial loss is divided into generator loss and discriminator loss, and the loss function of the adversarial loss can be expressed as formula (4):

[0099] (4)

[0100] in, To generate an image. The image is a real image, which in this embodiment can be a monocular photographic image of the vehicle. . For generator loss, F is the condition vector Q and the geometric vector. The concatenated vector This is due to discriminator loss. This represents the output of the generator. This indicates the discriminator's judgment result on the generator's output image. This indicates that the discriminator recognizes the real image. The determination result, E(·), represents the mathematical expectation, which is the average of the random variables within the parentheses over their distribution.

[0101] Optionally, pixel-level reconstruction loss (i.e., pixel loss) Using L1 loss (a metric for measuring the difference between model predictions and actual values) to measure the pixel-level difference between the generated image and the real image can be expressed as the following formula (5):

[0102] (5)

[0103] Optionally, a size normalization loss (i.e., size loss) can be introduced during training. Calculate the width of the vehicle detection box in the generated image. Compared with the original image (i.e., monocular photographic image) The width of the vehicle detection frame in ) For comparison, the size normalization loss is defined by formula (6):

[0104] (6)

[0105] As part of the loss function, the difference between the width of the vehicle detection box in the generated image and the width of the detection box in the original image is calculated to ensure that the vehicle size in the generated image conforms to the real-world proportions and to prevent the generation of vehicle appearances that do not conform to physical laws.

[0106] Alternatively, the total loss function of the generator can be expressed as the following formula (7):

[0107] (7)

[0108] in, and These are hyperparameters that can be set based on actual needs to balance the weights of various losses.

[0109] As mentioned above Figure 4 and Figure 5 Taking the vehicle monitoring scenario shown as an example, assuming a preset viewing angle offset... The monocular photographic image of the vehicle Target camera parameters The target perspective of the virtual camera The virtual viewpoint image generated by inputting the vehicle's 3D information (X, Y, Z) into the trained generative model can be represented as follows: .

[0110] In one embodiment, prior to step 303, the method may further include: determining motion information of the mobile device based on the monocular photographic image, and adding motion blur information to the virtual viewpoint image based on the motion information. And / or, determining illumination information corresponding to the monocular photographic image, and adding illumination noise information to the virtual viewpoint image based on the illumination information.

[0111] In this embodiment, by actively and specifically augmenting the generated virtual perspective images before feature extraction, the synthetic data is transformed from an "ideal laboratory state" to a "complex real-world state," thereby significantly improving the robustness and generalization ability of the subsequent feature extraction and recognition models. Specifically, motion information of the mobile device can be determined based on monocular photographic images, and motion blur can be added accordingly to simulate the dynamic blurring effect caused by the relative motion of the mobile device in real shooting scenes. This results in the generated virtual perspective image sequence not only containing static spatial multi-view information but also embedding temporal continuity representations that conform to physical laws, enhancing its adaptability to moving scenes.

[0112] Optionally, the actual lighting information corresponding to the monocular photographic image can be determined, and lighting noise can be added accordingly to reproduce the complex and varied lighting conditions in the real world (such as shadows, reflections, uneven brightness, etc.), so that the lighting effect of the generated image is consistent with the on-site lighting environment of the original monocular photographic image. At the same time, reasonable perturbations are introduced, thereby improving the immunity of the feature detection module to lighting interference.

[0113] Alternatively, motion blur and lighting noise addition can be implemented individually or in combination. This flexibility allows the generated virtual viewpoint images to cover a wide range of scenes, from static uniform lighting to dynamic complex lighting.

[0114] As mentioned above Figure 4 and Figure 5 Taking the vehicle monitoring scenario shown as an example, the generation result of the generative model is analyzed. Add motion blur and low-light noise to simulate realistic surveillance scenarios. For example, motion blur can be simulated using a linear motion blur model. First, based on vehicle speed... (km / h) and direction of motion (which can be obtained by tracking consecutive frames of the vehicle detection box) are used to calculate the vehicle's speed during exposure time. (For example, world coordinate displacement within a fixed 20ms) Then the displacement Projecting back onto the image plane yields the vehicle's pixel-level motion vectors. It can be expressed using formula (8):

[0115] (8)

[0116] in, Let f be the x-component of the camera focal length f in the camera coordinate system. Let f be the component of the camera focal length f along the y-axis in the camera coordinate system. , ( ) are the coordinates of the principal point. Based on the calculated motion vector... Generate a linear motion fuzzy kernel The core The length is The direction angle is Finally, the virtual perspective image of the virtual camera generated by the aforementioned preset generation model. Convolution is performed to obtain an image with motion blur. The formula is shown in (9):

[0117] (9)

[0118] Alternatively, the low-light noise addition process can be achieved using the following formula (10):

[0119] (10)

[0120] in, This represents the virtual viewpoint image output by the preset generation model. Image with added lighting noise Input monocular photographic image Corresponding ambient light intensity noise variance , .

[0121] Alternatively, it can also be applied to images with motion blur. Add illumination noise information to the existing data.

[0122] In one embodiment, before step 303, the method further includes: performing a quality assessment on the virtual viewpoint image to determine the quality confidence level of each image in the virtual viewpoint image. Images with a quality confidence level less than a preset threshold are removed from the virtual viewpoint image to obtain the final virtual viewpoint image.

[0123] In this embodiment, the virtual viewpoint images are quality-assessed, and the quality confidence level of each image is determined. This allows the system to quantitatively judge the credibility and clarity of the generated images, identifying blurry, distorted, or information-missing images that may be caused by limitations of the generation model, poor original input quality, or extreme viewpoints. Furthermore, images with a quality confidence level below a preset threshold are directly deleted, ensuring that only images meeting the quality standards can enter the subsequent computationally intensive feature detection process. This avoids the noise and misleading information carried by low-quality images contaminating the global and local feature extraction process, thereby improving the efficiency and accuracy of feature extraction.

[0124] Optionally, MobileNetV2 can be used as the backbone of the quality assessment network to perform quality assessment on the aforementioned generated virtual viewpoint images, outputting the quality confidence score of each virtual viewpoint image (the confidence score can range from 0 to 1), and setting a preset threshold for the quality confidence score. (For example, experience points 0.6), optionally, The system outputs virtual view images with a quality confidence score greater than a preset threshold, while discarding those with a quality confidence score less than the preset threshold. This ensures that the generated results conform to the characteristics of the monitoring scene and filters out low-quality generated images.

[0125] Optionally, during the training of the quality assessment network, positive samples (label=1) are real multi-view vehicle image datasets, and negative samples can be obviously distorted images generated by the multi-view generation network in the early stage of training or when the hyperparameters are poor, as well as manually labeled low-quality sample images. The loss function can be BCE loss (binary cross-entropy loss), which can be calculated using the following formula (11):

[0126] (11)

[0127] in, For sample labels, The quality confidence level of the network output is used to assess the quality.

[0128] Step 303: Detect global and local features of the mobile device in multi-view images. Local features are used to characterize the features of a specified region on the mobile device.

[0129] In this step, global and local features are detected in multi-view images. Global features can capture the overall shape and structural outline of the mobile device, while local features provide a refined representation of a specified area. Taking a vehicle as an example, the specified area includes, but is not limited to, details such as brand logos and wheel hubs. The collaborative extraction of global and local features makes the feature expression richer and more discriminative.

[0130] Optionally, taking a vehicle monitoring scenario as an example, the vehicle detection module can use YOLOv8 as the backbone network and adopt a dual-branch detection structure. The main branch is used to detect the overall vehicle target bounding box (i.e., global features) and outputs the detection bounding box and vehicle confidence score. The local region branch is used to detect local regions such as vehicle logos and wheel hubs (i.e., local features) and outputs the detection bounding boxes and confidence scores of these local regions. The local regions are used for subsequent fine-grained feature extraction to improve the accuracy of vehicle model recognition in extreme cases such as incomplete vehicle bodies or poor image quality.

[0131] In one embodiment, step 303 may specifically include: determining the camera pitch angle corresponding to the multi-view image; generating a non-uniform spatial attention mask corresponding to the multi-view image based on the camera pitch angle; adjusting the perspective relationship of the multi-view image based on the non-uniform spatial attention mask; and outputting the global features of the mobile device in the multi-view image based on the adjusted feature map.

[0132] In this embodiment, an adaptive spatial attention mechanism based on camera pitch angle is introduced to optimize the global feature extraction process, improving the robustness and representativeness of features to changes in viewing angle. Specifically, the camera pitch angle corresponding to the multi-view images is determined to provide specific viewpoint geometry information during image acquisition, enabling the model to perceive perspective distortion differences caused by different shooting angles. A non-uniform spatial attention mask is generated based on this pitch angle, allowing the model to dynamically adjust the attention weights for different image regions. For example, at high pitch angles, attention is strengthened to the top features of the mobile device, while at low pitch angles, attention is focused on side features, thus adaptively focusing on the most discriminative region. Furthermore, this mask is used to adjust the perspective relationship of the multi-view images, effectively correcting the geometric distortion and spatial compression caused by viewpoint tilt in the original image, making the adjusted feature map more consistent with the standard geometric representation of the object, thereby reducing missed detections.

[0133] Optionally, taking the aforementioned vehicle monitoring scenario as an example, global feature detection of the vehicle is implemented based on a YOLOv8 network, with the input being vehicle images (including monocular photographic images). (and the virtual view image generated by the model) and the pitch angle of the target camera A geometry-aware module is embedded after the YOLOv8 backbone network, and a geometry scaling factor is applied. Defined as formula (12):

[0134] (12)

[0135] This geometric scaling factor This reflects the vertical compression effect of the image caused by the downward viewing angle: the larger the pitch angle (i.e., the lower the lens is), the more severely the vehicle is "flattened" in the image, and the lower the effective vertical resolution. The feature map of the input image in the last stage of the network... A non-uniform spatial attention mask is applied to the top edge of the image height dimension (H). The mask M is... The value of the control is larger at the bottom (near area) of the image and smaller at the top (far area), simulating the perspective law of near objects appearing larger and far objects appearing smaller in the real world. The final output feature map can be expressed as formula (13):

[0136] (13)

[0137] Where ⊙ represents element-wise multiplication, repeat() represents expanding the one-dimensional mask to the full channel and width dimensions, W represents the width of the input image, and C represents the number of color channels of the image.

[0138] The geometric perception module solves the problem of missing detection of small target vehicles at a distance, outputs vehicle detection boxes and vehicle confidence scores, and initially classifies the vehicles in the picture into categories such as motorcycles, cars, SUVs, buses, trucks and others.

[0139] Alternatively, the global feature detection network can be replaced by other object detection networks such as Faster R-CNN or SSD instead of YOLOv8. The non-uniform spatial attention mask M in the geometry perception module can also be implemented using an exponential decay function or other monotonically decreasing functions.

[0140] In one embodiment, step 303 may further include: determining the signal-to-noise ratio (SNR) of the monocular image; determining a first confidence threshold for monocular image adaptation based on the SNR, wherein the first confidence threshold is positively correlated with the SNR; detecting initial global features of the mobile device in multi-view images respectively, and selecting initial global features with a confidence level greater than the first confidence threshold as the final global features of the mobile device in the corresponding view images.

[0141] In this embodiment, an adaptive feature confidence screening mechanism based on source image quality is established, upgrading global feature detection from a static, uniform threshold judgment to a dynamic, precise quality control process, thereby improving the reliability of the extracted global features and the overall robustness of the system. Specifically, the signal-to-noise ratio (SNR) of the monocular image is first determined, and an appropriate first confidence threshold is determined based on this SNR, ensuring that this threshold is positively correlated with the SNR. This allows the strictness of feature screening to adapt to the quality of the input monocular image. That is, when the monocular image is clear (high SNR), a higher confidence threshold is used for strict screening to ensure that the retained features have extremely high determinism and discriminative power. When the monocular image quality is poor (low SNR), the threshold requirement is correspondingly reduced, thereby avoiding the loss of too much feature information while tolerating a certain degree of uncertainty, achieving a trade-off under quality fluctuations. Then, the initial global features of each multi-view image are detected separately, and uniform screening is performed according to the above-mentioned adaptive first confidence threshold to ensure that the final global feature set used to identify the mobile device has undergone standardized quality control that matches the quality of its source image, and those candidate features that are unreliable due to low image quality are eliminated.

[0142] Optionally, a dynamic confidence calibration mechanism can adjust the confidence threshold for target detection in real time based on the image signal-to-noise ratio (SNR) to avoid missed vehicle detection under low-resolution and low-light conditions. The image signal-to-noise ratio (SNR) is calculated using the following formula (14):

[0143] (14)

[0144] in, Monocular photographic images of vehicles grayscale image, mean ( ) indicates that the input image is calculated. The average of the squared pixel values, var(I) represents the calculation of the input image. The variance of pixel values. Different target detection confidence thresholds can be set based on the image signal-to-noise ratio.

[0145] For example, different first confidence thresholds can be set based on multi-level signal-to-noise ratios (SNR). Assuming there are two SNR levels of 30dB and 10dB, the system checks if the image SNR is greater than 30dB. If the image SNR is greater than 30dB, the first confidence threshold can be set to 0.5. If the image SNR is less than or equal to 30dB but greater than 10dB, the first confidence threshold is set to 0.3. If the image SNR is less than or equal to 10dB, a multi-frame temporal fusion mechanism is triggered. For example, it calls up several consecutive frames (e.g., 3 frames) of monocular images of the vehicle taken by the target camera. The weighted average of the SNR of these consecutive monocular images is used as the final input image SNR, and the first confidence threshold is set based on this final input image SNR.

[0146] Step 304: Output the recognition results of the mobile device based on global and local features.

[0147] In this step, the recognition results are output based on global and local features, which enhances the recognition accuracy of mobile devices in monocular photography scenarios and improves the robustness and adaptability of the recognition method in complex scenarios.

[0148] In one embodiment, step 304 may specifically include: for a single-view image in a multi-view image, identifying the first attribute information of the mobile device under a preset attribute category based on global features; and for the single-view image in the multi-view image, calculating the similarity between local features and reference features of each reference device in a preset feature library. The preset feature library includes multiple reference devices and reference features and second attribute information of each reference device under a preset attribute category. The reference device with the highest similarity is determined as a candidate device matching the mobile device, and the preset second attribute information of the candidate device is obtained. Based on the first and second attribute information corresponding to the multi-view image, the identification result of the mobile device is output.

[0149] In this embodiment, the preset attribute categories include, but are not limited to, one or more of the mobile device's brand, model, style, and color. By effectively combining the advantages of two technical approaches—"attribute recognition based on global features" and "instance retrieval based on local features"—high-precision and robust mobile device recognition is achieved. First, for single-view images, the first attribute information of the preset attribute category is identified based on global features. The strong representational ability of global features of the overall shape and macroscopic attributes of the mobile device enables rapid and direct attribute information discrimination. Simultaneously, for the same single-view image, the similarity between its local features and reference features in the preset feature library is calculated, performing refined instance retrieval and comparison, fully leveraging the high discriminative power of local features for detailed features. The reference device with the highest similarity is identified as a candidate device, and its preset second attribute information is obtained. Then, a fusion decision is made based on the first and second attribute information corresponding to multi-view images, providing the possibility of internal cross-validation for the recognition results (including global classification results and instance retrieval results) of each view image. For example, when the "first attribute information" determined by global features is unreliable due to occlusion at a certain viewpoint, it can be supplemented or corrected by the "second attribute information" retrieved from instances under that viewpoint. This effectively mitigates recognition errors or randomness under a single viewpoint and significantly improves the confidence and accuracy of the final recognition results by utilizing the complementarity between viewpoints.

[0150] Optionally, taking a vehicle monitoring scenario as an example, a vehicle attribute recognition model can be pre-trained to identify the first attribute information of a mobile device under a preset attribute category based on global features. Specifically, a sufficiently large dataset can be constructed for training the vehicle attribute recognition model. For example, this dataset can include 500,000 samples and more than 3,000 attribute category samples, where the number of samples for each attribute category needs to be expanded to at least a certain number (e.g., more than 500 images), and the ratio of the training set to the test set can be 9:1.

[0151] Optionally, the training dataset for the vehicle attribute recognition model can be obtained by acquiring vehicle model data in the following two ways:

[0152] a. Web crawler

[0153] Content was scraped from mainstream automotive websites such as Autohome, Dongchedi, and Yiche using web scraping technology. The scraped content included vehicle images (multi-angle, multi-scene) and vehicle parameters (brand, model, year, color). A lightweight detector (such as YOLOv5s) was used to filter out non-vehicle images from the scraped images, retaining only those with a confidence score > 0.9.

[0154] b. Multi-view AIGC data augmentation

[0155] For rare car models with insufficient online images (e.g., models with fewer than 500 samples), InstructPix2Pix (an AI technology for image editing based on natural language instructions) and / or Qwen-Image-Edit (a large language model supporting image editing) can be used for targeted data augmentation. By configuring prompts for different shooting angles, and focusing on configuring the generated different perspectives (such as front, rear, side, and top-down views), multi-view images of the car model can be generated. Finally, a quality evaluation network is used for evaluation, and the generated results with a confidence score > 0.75 are added to the dataset to alleviate the long-tail distribution problem of vehicle data.

[0156] The entire dataset construction process can be divided into three stages: Initially, approximately 100,000 basic images were crawled, covering common car models. In the middle stage, AIGC was used to generate 500,000 augmented images, focusing on supplementing rare car models. In the later stage, continuous optimization was carried out using data from real-world application scenarios.

[0157] Optionally, the vehicle attribute recognition model can adopt a multi-task learning architecture, such as using ResNet50 as the shared backbone network of the vehicle attribute recognition model (applicable to other models), outputting a global feature vector. Separate classification heads are configured for brand, model, year, and color tasks, outputting the corresponding class probabilities. Different weights are configured for the loss function, resulting in a total loss function for the vehicle attribute recognition model. This can be expressed as formula (15):

[0158] (15)

[0159] in, This indicates a loss of brand identity for the vehicle. This indicates the weight of the vehicle's brand attributes. This indicates the loss of vehicle model attributes. This indicates the weight of the vehicle's model attribute. This indicates the loss of the vehicle's model year attribute. This indicates the weight of the vehicle's model year attribute. This indicates the loss of the vehicle's color attributes. This represents the weight of the vehicle's color attribute. The weights of the loss values ​​for each attribute category can be set as follows: , , , This means prioritizing the accuracy of the brand and model number, and different loss weights can be set according to different scenarios.

[0160] The vehicle attribute recognition model can employ a two-stage training strategy: the first stage freezes the backbone network, training only the classifier heads to achieve rapid convergence. The second stage fine-tunes the entire network, while introducing label smoothing and Mixup (linear interpolation data augmentation) data augmentation to improve the model's generalization ability.

[0161] Optionally, the backbone network of the vehicle attribute recognition model can be replaced with other network architectures such as EfficientNet (an efficient convolutional neural network model) or ViT (Vision Transformer). The loss weights for multi-task learning can be dynamically adjusted, for example, adaptively adjusting the weight ratio of brand model and color tasks based on image quality.

[0162] Optionally, the virtual viewpoint image generated in step 302 is normalized and the original monocular photographic image of the vehicle is normalized. Then, the image is cropped according to the vehicle target bounding box. The cropped image is input into a trained vehicle attribute recognition model. The vehicle attribute recognition model outputs the probability of the vehicle's brand, model, year, color, and other preset attribute categories in each input image. The output of the vehicle attribute recognition model can be in the form of index numbers, with each index number representing an attribute category. A preset mapping table can be pre-configured, which includes attribute categories corresponding to different index numbers. According to the preset mapping table, the output of the vehicle attribute recognition model is mapped to the final first attribute information. ,as follows:

[0163]

[0164] in, The brand name corresponds to a first sub-confidence level of 1. . Indicates the model number, and the corresponding first sub-confidence level is... . Indicates the year, with the corresponding first sub-confidence level being... . Representing color, the corresponding first sub-confidence is . This indicates the first confidence level of the first attribute information, and .

[0165] In real-world scenarios, the overall vehicle features in monocular photographic images are easily affected by factors such as lighting, occlusion, and shooting angle. Vehicle attribute recognition models have limitations in recognizing the global features of such vehicles. This application proposes a feature retrieval module that uses local feature retrieval to find vehicle models and combines this with vehicle model attribute recognition to jointly determine the output of vehicle attribute information, thereby improving vehicle model recognition performance in extreme cases.

[0166] Optionally, such as Figure 6 As shown, the feature retrieval module can be divided into two parts: local feature library construction and feature retrieval.

[0167] 1. Construction of local feature library

[0168] First, acquire a dataset of vehicle images. This dataset can be the same one used in the training process of the vehicle attribute recognition model. Then, employ a local region detector from the vehicle attribute recognition model to detect specified local regions such as car logos and wheel rims in the vehicle dataset images (e.g., screenshots of the car logo and wheel rims). Input the detection results of these local regions into a pre-trained local feature extractor for feature extraction, obtaining the vehicle's local features. This local feature extractor can be a ResNet18 (18-layer residual network) feature extractor pre-trained on the COCO dataset. Finally, the extracted local features are inserted into a vehicle local feature library for subsequent feature retrieval. Optionally, each entry in the local feature library includes: vehicle model ID, car logo feature vector, wheel rim feature vector, image path, and generation timestamp. The vehicle model ID corresponds one-to-one with the vehicle model IDs in the overall dataset, ensuring consistent naming and synchronization with the overall dataset.

[0169] 2. Feature retrieval

[0170] In the retrieval phase, for the input vehicle image (including monocular photographic images and the virtual viewpoint image generated in step 302), the same local feature extractor as in step 1 is used to extract features from the cropped vehicle logo and wheel hub images, obtaining local feature vectors for the vehicle. The extracted local feature vectors are then compared with preset local feature vectors in the local feature library for similarity calculation. For example, cosine similarity can be used. The attribute information of the candidate vehicle (i.e., Top-1 model) corresponding to the preset local feature vector with the highest similarity is taken as the output result of the feature retrieval. This feature retrieval result (i.e., the second attribute information) can be represented as follows:

[0171]

[0172] in, The brand name is represented by the second sub-confidence level. . Indicates the model number, and the corresponding second sub-confidence level is . Indicates the year, with the corresponding second sub-confidence level being... . Representing color, the corresponding second sub-confidence is . The second confidence level represents the second attribute information. The value of the second confidence level is equal to the similarity between the extracted local feature vector and the preset local feature vector of the candidate vehicle. .

[0173] Alternatively, other feature extractors such as MobileNet or Swin Transformer (ShiftedWindow Transformer) can be used instead of ResNet18 for local feature extraction. Similarity can be calculated using other metrics such as Euclidean distance or Mahalanobis distance.

[0174] In one embodiment, step 304, based on the first and second attribute information corresponding to the multi-view images, outputs the recognition result of the mobile device, including: for a single-view image in the multi-view images, determining a first confidence level of the first attribute information and a second confidence level of the second attribute information in the single-view image; determining an initial recognition result of the mobile device in the single-view image based on the first and second confidence levels; for a single attribute category in the preset attribute categories, calculating the weighted frequency of each initial recognition result corresponding to the single attribute category in the multi-view images; determining the initial recognition result with the highest weighted frequency as the final recognition result of the mobile device under the single attribute category; and outputting the final recognition results of the mobile device under each attribute category in the preset attribute categories.

[0175] In this embodiment, a refined decision-making mechanism based on confidence weighting and cross-view multi-attribute statistical fusion is introduced to systematically integrate recognition information from multiple perspectives and sources, thereby achieving a final recognition output with high robustness and high accuracy. First, for each single-view image, the first confidence level of its first attribute information and the second confidence level of its second attribute information are determined, so that both the results from global feature classification and local feature retrieval are accompanied by a quantitative measure of reliability. Then, based on the two confidence levels, the initial recognition result for the single-view image is determined, and the dual-path recognition information is initially optimized or weighted and fused within that single-view image. For each attribute category, the weighted frequency of each initial recognition result corresponding to that attribute across all multi-viewpoints is calculated. The "weighting" is related to the confidence level used to generate the initial recognition result; for example, initial results with higher confidence contribute more weight. This ensures that the fusion process considers not only the "quantity" (frequency) of results from different views but also the "quality" (confidence) of each result, effectively amplifying the decision-making influence of high-quality, high-determinism perspectives while suppressing low-confidence noise results caused by poor image quality or severe occlusion in individual perspectives. The result with the highest weighted frequency is determined as the final recognition result for that attribute category, and the results for each attribute category are output separately. This achieves adaptive, fault-tolerant multi-view decision fusion, improving the overall fault tolerance, stability, and attribute-level recognition accuracy of mobile device recognition systems in complex real-world scenarios.

[0176] In one embodiment, determining the initial recognition result of the mobile device in a single-view image based on a first confidence level and a second confidence level includes: if the first confidence level is greater than a second confidence level threshold, determining the first attribute information as the initial recognition result of the mobile device in the single-view image; if the first confidence level is less than or equal to the second confidence level threshold, and the second confidence level is greater than the second confidence level threshold, determining the second attribute information as the initial recognition result of the mobile device in the single-view image; if both the first confidence level and the second confidence level threshold are less than the second confidence level threshold, determining the initial recognition result of the mobile device in the single-view image based on the first sub-confidence of different attribute categories in the first attribute information and the second sub-confidence of the corresponding attribute categories in the second attribute information.

[0177] In this embodiment, a clear three-level judgment logic is established. When the first confidence level (the attribute classification result derived from global features) is higher than the second confidence level threshold, the first attribute information derived from global features is preferentially determined as the initial recognition result. This fully utilizes the advantages of global features in direct and efficient discrimination when the information is complete, achieving fast and reliable decision-making. When the first confidence level is insufficient and the second confidence level (the instance retrieval matching result derived from local features) is higher than the second confidence level threshold, the second attribute information derived from local features is determined as the initial recognition result. Thus, when global features are unreliable due to image quality, occlusion, or other reasons, the system can automatically switch to a backup path that relies on local details for instance matching, ensuring the continuity of decision-making. When the confidence levels of both main information sources are insufficient, the system further delves into the attribute category sub-level. Based on the first sub-confidence of each attribute in the first attribute information and the second sub-confidence of the corresponding attribute in the second attribute information, a refined determination is made to independently evaluate and select different attribute categories such as "brand" and "model" at a finer granular level, thereby achieving a relatively reliable initial recognition result for each attribute component. It enhances the adaptability and fault tolerance of the initial recognition results generated from a single viewpoint, providing higher quality and more reliable intermediate results for subsequent cross-viewpoint fusion.

[0178] like Figure 7 The diagram shown is a flowchart illustrating a process for determining the initial recognition result based on a confidence selection mechanism, as provided in an embodiment of this application. Taking the aforementioned vehicle monitoring scenario as an example, to ensure the reliability of the final recognition result, for a single-view image, the output of the aforementioned vehicle attribute recognition model is used... and the output of the feature retrieval process and their corresponding confidence levels Based on the combined confidence levels of both methods, the initial vehicle identification result is selected and output. The output initial identification result... Prioritize the output of the vehicle attribute recognition model ,when First confidence level When the value is below the set second confidence threshold, determine Second confidence level :like Second confidence level If it exceeds the second confidence threshold, then take... This is the final identification result. If... If none of the confidence scores reach the second confidence threshold, the initial recognition result of the vehicle in the single-view image is determined by comparing the corresponding sub-confidence scores. Assume the second confidence threshold is 0.8. The specific process may include the following steps:

[0179] Step 701: Determine Is it greater than or equal to 0.8? If yes, proceed to step 702; otherwise, proceed to step 703.

[0180] Step 702: Determine Is it greater than or equal to 0.8? If yes, proceed to step 706; otherwise, proceed to step 703.

[0181] Step 703: Determine Is it greater than or equal to 0.8? If yes, proceed to step 705; otherwise, proceed to step 704.

[0182] Step 704: Determine Is it greater than or equal to? If yes, proceed to step 706; otherwise, proceed to step 707.

[0183] Step 705: Determine Is it greater than or equal to 0.8? If yes, proceed to step 707; otherwise, proceed to step 704.

[0184] Step 706: Determine the initial vehicle identification result = .

[0185] Step 707: Determine the initial vehicle identification result = .

[0186] Optionally, the above Figure 7 The corresponding confidence level selection mechanism can set multiple threshold ranges, corresponding to different decision-making logics.

[0187] Based on this, for each virtual viewpoint image generated in step 302 above, the vehicle attribute recognition model is run independently to obtain the initial recognition results of the vehicle in each single-viewpoint image. And its confidence level. Then, through a multi-view voting mechanism, for each attribute category (such as brand), the initial recognition results under all view images are statistically analyzed. The frequency of occurrence of specific attribute information is used to determine the final identification result of the vehicle under that attribute category, and the attribute information with the highest frequency is selected. Taking the aforementioned vehicle monitoring scenario as an example, the specific implementation process can be as follows:

[0188] 1) Input Preparation

[0189] Receive multi-view images of the vehicle. Assume the total input is images from N perspectives, for example, N=3, corresponding to -45°, 0°, and +45° perspectives.

[0190] For each viewpoint image, the vehicle attribute recognition model is invoked, and the corresponding initial recognition result and confidence score are output. The initial recognition results and confidence scores of virtual viewpoint images under different attribute categories can be represented as follows:

[0191] Initial brand identification results set: ,in Let be the initial brand recognition result in the i-th viewpoint image. The value of i ranges from [1, N].

[0192] Initial model identification result set: ,in The initial model identification result is given in the image from the i-th viewpoint.

[0193] Initial set of year identification results: ,in This represents the initial year identification result in the i-th viewpoint image.

[0194] Initial color recognition result set: ,in This represents the initial color recognition result in the i-th viewpoint image.

[0195] Corresponding confidence set: ,in denoted as the overall confidence level of the i-th perspective.

[0196] 2) Attribute category voting

[0197] Brand Voting: Statistical Collection The frequency of each brand appearing in the text is calculated based on the corresponding confidence level, and the calculation process can be achieved using formula (16):

[0198] (16)

[0199] in, For brand information The weighted frequency, that is, the initial identification result of all brands in set B is b. Sum the corresponding confidence levels. Let b be the confidence level of the initial brand recognition result in the i-th viewpoint image.

[0200] Model Voting: Similarly, for set S, calculate the weighted frequency of each model based on the corresponding confidence level. .

[0201] Year-based voting: Similarly, for set K, calculate the weighted frequency of each year based on the corresponding confidence level. .

[0202] Color voting: Similarly, for sets The weighted frequency of each color is calculated based on the corresponding confidence level. .

[0203] 3) The final identification result is determined and includes the following:

[0204] The final brand identity result: That is, the brand information with the highest weighted frequency is selected as the final identification result of the vehicle under the brand category.

[0205] Model Result: That is, the model information with the highest weighted frequency is selected as the final identification result of the vehicle under the model category.

[0206] Year / model year results: That is, the model year information with the highest weighted frequency is selected as the final identification result of the vehicle under the model year category.

[0207] Color result: That is, the color information with the highest weighted frequency is selected as the final recognition result of the vehicle under the color category.

[0208] Overall confidence level: , where N is the total number of viewpoints in the input image.

[0209] Alternatively, the multi-perspective voting mechanism can also adopt an index-weighted method based on confidence level, or introduce time series information for time-series fusion.

[0210] 4) Exception handling

[0211] Optionally, when the maximum weighted frequency is lower than a threshold (e.g., set to 0.5), a secondary verification mechanism can be triggered: The virtual perspective image generated in step 302 is invoked, and its features are compared with those of the vehicle's monocular photographic image. For example, the aforementioned local feature extractor can be used to extract features from both the virtual perspective image and the monocular photographic image. Based on the extracted features, the similarity between the two is calculated. The target virtual perspective image with the highest similarity to the monocular photographic image is selected from multiple virtual perspective images. If the similarity between the target virtual perspective image and the monocular photographic image is greater than a certain threshold (e.g., 0.6), then the attribute information of the vehicle in each attribute category in the target virtual perspective image is used as the final recognition result of the vehicle. Otherwise, "Uncertain" is returned with the reason noted.

[0212] like Figure 8 The diagram shown is a flowchart of a mobile device identification method provided in an embodiment of this application. Taking a vehicle monitoring scenario as an example, it may include the following steps:

[0213] Step 801: Acquire a monocular photographic image of the vehicle.

[0214] Step 802: Determine the target camera to be used when capturing monocular images and obtain the pre-calibrated camera parameters of the target camera.

[0215] Step 803: Determine the 3D information of the vehicle in the real scene and the shooting angle corresponding to the monocular photographic image based on the camera parameters.

[0216] Step 804: Shift the shooting angle according to the preset angle offset to obtain the target angle of at least one virtual lens.

[0217] Step 805: Based on the monocular photographic image, camera parameters, target viewpoint, and 3D information, generate virtual viewpoint images of the vehicle from different target viewpoints. The multi-view images include both monocular photographic images and virtual viewpoint images.

[0218] Step 806: Detect global and local features of vehicles in multi-view images.

[0219] Step 807: For a single-view image in a multi-view image, based on global features, identify the first attribute information of the vehicle under a preset attribute category. The preset attribute category includes one or more of the vehicle's brand, model, style, and color.

[0220] Step 808: For a single-view image within a multi-view image, calculate the similarity between local features and reference features of each reference device in a preset feature library. The preset feature library includes multiple reference devices, their reference features, and second attribute information under preset attribute categories.

[0221] Step 809: Identify the reference device with the highest similarity as the candidate device that matches the vehicle, and obtain the preset second attribute information of the candidate device.

[0222] Step 810: Based on the first attribute information and the second attribute information corresponding to the multi-view images, output the vehicle recognition result.

[0223] For details of each step of the above method, please refer to the relevant descriptions in the above embodiments, which will not be repeated here.

[0224] The aforementioned mobile device recognition method, based on artificial intelligence algorithms and surveillance cameras, enables vehicle model recognition under a monocular surveillance lens. In practical applications, it can generate virtual multi-view images of vehicles using a multi-view generation network based on the image from the monocular surveillance lens and the lens parameters, effectively solving the problem of missing depth in monocular cameras and enabling the monocular surveillance lens to generate multi-view images of vehicle appearance. Detection boxes for local areas such as vehicles, logos, and wheel rims are extracted from the multi-view images, cropped, and used for vehicle attribute classification and feature retrieval, respectively. The resulting vehicle classification and feature retrieval results are analyzed and integrated to output the final vehicle attribute table information, including but not limited to vehicle type, brand, model, year, and color. It can be applied to different scenarios, such as parking lot management, insurance claims, and vehicle after-sales analysis, requiring only adjustments to the multi-view generation parameters and the training data of the attribute recognition model for specific scenarios.

[0225] Please refer to Figure 9 This is a mobile device identification device 900 according to an embodiment of this application. This device can be applied to electronic device 1 and can also be applied to... Figure 2 The mobile device recognition application scenario shown aims to enhance the recognition accuracy of mobile devices in monocular photography scenarios and improve the robustness and adaptability of the recognition method in complex scenes. The device includes: an acquisition module 901, a generation module 902, a detection module 903, and an output module 904. The functional principles of each module are as follows:

[0226] The acquisition module 901 is used to acquire monocular photographic images from the mobile device.

[0227] The generation module 902 is used to generate multi-view images for mobile devices based on monocular photographic images.

[0228] The detection module 903 is used to detect global and local features of mobile devices in multi-view images.

[0229] Output module 904 is used to output the recognition results of mobile devices based on global and local features.

[0230] In one embodiment, the generation module 902 is used to determine the target camera used when capturing a monocular photographic image and to obtain the pre-calibrated camera parameters of the target camera. Based on the camera parameters, the 3D information of the mobile device in the real scene and the shooting angle corresponding to the monocular photographic image are determined. The shooting angle is offset according to a preset angle offset to obtain the target angle of at least one virtual lens. Based on the monocular photographic image, camera parameters, target angle, and 3D information, virtual perspective images of the mobile device under different target angles are generated. The multi-view images include the monocular photographic image and the virtual perspective images.

[0231] In one embodiment, the generation module 902 is used to determine the top-down angle of the target camera relative to the mobile device when capturing a monocular photographic image, based on camera parameters. The lens depth of the target camera is determined based on the top-down angle. The position coordinates of the mobile device in the world coordinate system are determined according to the camera parameters and the lens depth. The three-dimensional information of the mobile device includes the lens depth and position coordinates.

[0232] In one embodiment, the generation module 902 is used to input a monocular photographic image, camera parameters, target viewpoint, and 3D information into a preset generation model, so that the preset generation model outputs virtual viewpoint images of the mobile device under different target viewpoints. The preset generation model generates conditional vectors based on camera parameters and target viewpoints, and generates geometric vectors based on 3D information. Based on the conditional vectors and geometric vectors, the monocular photographic image is reconstructed into a target feature map. Based on the 3D information, a spatial transformation is performed on the target feature map to generate virtual viewpoint images of the mobile device under different target viewpoints.

[0233] In one embodiment, the apparatus further includes: a training module, configured to acquire a set of sample images containing sample mobile devices before inputting monocular photographic images, camera parameters, target viewpoints, and 3D information into a preset generative model. The sample images are labeled with corresponding sample camera parameters, pixel coordinates of the sample mobile device within the sample image, and 3D coordinates of the sample mobile device in the world coordinate system. The sample images are configured with labeled images of the sample mobile device from different viewpoints. A preset conditional generative adversarial network is trained based on the sample image set to obtain a preset generative model. The conditional generative adversarial network includes a generator and a discriminator, and the generator's loss function includes pixel loss and size loss between the generated image and the real image.

[0234] In one embodiment, the apparatus further includes: a post-processing module, configured to determine motion information of the mobile device based on a monocular photographic image before detecting global and local features of the mobile device in the multi-view image, and to add motion blur information to the virtual view image based on the motion information; and / or, to determine illumination information corresponding to the monocular photographic image, and to add illumination noise information to the virtual view image based on the illumination information.

[0235] In one embodiment, the detection module 903 is used to determine the camera pitch angle corresponding to the multi-view image. A non-uniform spatial attention mask is generated based on the camera pitch angle. The perspective relationship of the multi-view image is adjusted based on the non-uniform spatial attention mask, and the global features of the mobile device in the multi-view image are output based on the adjusted feature map.

[0236] In one embodiment, the detection module 903 is used to determine the signal-to-noise ratio (SNR) of a monocular photographic image. Based on the SNR, a first confidence threshold for monocular image adaptation is determined, and the first confidence threshold is positively correlated with the SNR. Initial global features of the mobile device in multi-view images are detected respectively, and initial global features with a confidence level greater than the first confidence threshold are selected as the final global features of the mobile device in the corresponding view images.

[0237] In one embodiment, the output module 904 is used to identify, based on global features, the first attribute information of the mobile device under a preset attribute category for a single-view image in a multi-view image. The preset attribute category includes one or more of the mobile device's brand, model, style, and color. For the single-view image in the multi-view image, the similarity between local features and reference features of each reference device in a preset feature library is calculated. The preset feature library includes multiple reference devices, reference features of each reference device, and second attribute information under the preset attribute category. The reference device with the highest similarity is determined as a candidate device matching the mobile device, and the preset second attribute information of the candidate device is obtained. Based on the first and second attribute information corresponding to the multi-view image, the identification result of the mobile device is output.

[0238] In one embodiment, the output module 904 is configured to, for a single-view image within a multi-view image, determine a first confidence level of a first attribute information and a second confidence level of a second attribute information in the single-view image. Based on the first and second confidence levels, an initial recognition result of the mobile device in the single-view image is determined. For a single attribute category within a preset attribute category, the weighted frequency of each initial recognition result corresponding to that single attribute category in the multi-view image is calculated. The initial recognition result with the highest weighted frequency is determined as the final recognition result of the mobile device under that single attribute category. The final recognition results of the mobile device under each attribute category within the preset attribute categories are output.

[0239] In one embodiment, the output module 904 is configured to determine the initial recognition result of the mobile device in a single-view image based on a first confidence level and a second confidence level, including: if the first confidence level is greater than a second confidence level threshold, determining the first attribute information as the initial recognition result of the mobile device in the single-view image; if the first confidence level is less than or equal to the second confidence level threshold, and the second confidence level is greater than the second confidence level threshold, determining the second attribute information as the initial recognition result of the mobile device in the single-view image; if both the first confidence level and the second confidence level threshold are less than the second confidence level threshold, determining the initial recognition result of the mobile device in the single-view image based on the first sub-confidence of different attribute categories in the first attribute information and the second sub-confidence of the corresponding attribute categories in the second attribute information.

[0240] For a detailed description of the mobile device identification device 900 described above, please refer to the description of the relevant method steps in the above embodiments. The implementation principle and technical effect are similar, and will not be repeated here.

[0241] Figure 10 This is a schematic diagram of the structure of a cloud device 100 provided as an exemplary embodiment of this application. The cloud device 100 can be used to run the methods provided in any of the above embodiments. Figure 10 As shown, the cloud device 100 may include: a memory 1004 and at least one processor 1005. Figure 10 Let's take a processor as an example.

[0242] The storage device 1004 is used to store computer programs and can be configured to store various other data to support operations on the cloud device 100. The storage device 1004 may be object storage (OSS).

[0243] The memory 1004 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0244] The processor 1005, coupled to the memory 1004, is used to execute the computer program in the memory 1004 to implement the solution provided in any of the above method embodiments. The specific functions and technical effects that can be achieved will not be described here.

[0245] Furthermore, such as Figure 10 The cloud device also includes other components such as firewall 1001, load balancer 1002, communication component 1006, and power supply component 1003. Figure 10 The diagram only shows some components and does not mean that cloud devices only include... Figure 10 The components shown.

[0246] In one embodiment, the above Figure 10The communication component 1006 is configured to facilitate wired or wireless communication between the device containing the communication component 1006 and other devices. The device containing the communication component 1006 can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, LTE (Long Term Evolution), 5G, or combinations thereof. In one exemplary embodiment, the communication component 1006 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component 1006 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra Wide Band (UWB), Bluetooth, and other technologies.

[0247] In one embodiment, the above Figure 10 The power supply component 1003 provides power to various components of the device in which it resides. The power supply component 1003 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.

[0248] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.

[0249] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.

[0250] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0251] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.

[0252] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The memory may include high-speed RAM (Random Access Memory), and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk, or optical disc, etc.

[0253] The aforementioned storage media can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage media can be any available medium accessible to general-purpose or special-purpose computers.

[0254] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.

[0255] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes that element.

[0256] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0257] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0258] The collection, storage, use, processing, transmission, provision, and disclosure of user data and other information involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0259] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for identifying mobile devices, characterized in that, include: Acquire monocular images from mobile devices; Based on the monocular photographic image, generate a multi-view image of the mobile device; Detect global and local features of the mobile device in the multi-view images; The identification result of the mobile device is output based on the global features and the local features.

2. The method according to claim 1, characterized in that, The step of generating multi-view images of the mobile device based on the monocular photographic image includes: Determine the target camera used when capturing the monocular photographic image, and obtain the pre-calibrated camera parameters of the target camera; Based on the camera parameters, determine the 3D information of the mobile device in the real scene in the monocular photographic image and the shooting angle corresponding to the monocular photographic image; The shooting angle is offset according to a preset angle offset to obtain the target angle of at least one virtual lens; Based on the monocular photographic image, the camera parameters, the target viewpoint, and the 3D information, virtual viewpoint images of the mobile device under different target viewpoints are generated, and the multi-view images include the monocular photographic image and the virtual viewpoint image.

3. The method according to claim 2, characterized in that, Determining the 3D information of the mobile device in the real-world scene from the monocular photographic image based on the camera parameters includes: Based on the camera parameters, determine the top-down angle of the target camera relative to the mobile device when capturing the monocular photographic image; The lens depth of the target camera is determined based on the overhead view. The position coordinates of the mobile device in the world coordinate system are determined based on the camera parameters and the lens depth. The three-dimensional information of the mobile device includes the lens depth and the position coordinates.

4. The method according to claim 2, characterized in that, The step of generating virtual view images of the mobile device under different target viewpoints based on monocular photographic images, camera parameters, target viewpoints, and 3D information includes: The monocular photographic image, the camera parameters, the target viewpoint, and the 3D information are input into a preset generation model so that the preset generation model outputs virtual viewpoint images of the mobile device under different target viewpoints; The preset generation model is used to generate a conditional vector based on the camera parameters and the target viewpoint, and to generate a geometric vector based on the three-dimensional information; based on the conditional vector and the geometric vector, the monocular photographic image is reshaped into a target feature map; and the target feature map is spatially transformed based on the three-dimensional information to generate virtual viewpoint images of the mobile device under different target viewpoints.

5. The method according to claim 4, characterized in that, Before inputting the monocular photographic image, the camera parameters, the target viewpoint, and the 3D information into the preset generation model, the method further includes: Obtain a set of sample images containing sample mobile devices. The sample images are labeled with corresponding sample camera parameters, the pixel coordinates of the sample mobile device in the sample image, and the three-dimensional coordinates of the sample mobile device in the world coordinate system. The sample images are configured with label images of the sample mobile device from different viewpoints. A preset conditional generative adversarial network is trained based on the sample image set to obtain the preset generative model; The conditional generative adversarial network includes a generator and a discriminator, and the loss function of the generator includes pixel loss and size loss between the generated image and the real image.

6. The method according to claim 2, characterized in that, Before detecting the global and local features of the mobile device in the multi-view image, the method further includes: Based on the monocular photographic image, determine the motion information of the mobile device, and add motion blur information to the virtual viewpoint image based on the motion information; And / or, determine the illumination information corresponding to the monocular photographic image, and add illumination noise information to the virtual viewpoint image based on the illumination information.

7. The method according to claim 1, characterized in that, The detection of global features of the mobile device in the multi-view image includes: Determine the camera pitch angle corresponding to the multi-view images; Generate a non-uniform spatial attention mask corresponding to the multi-view image based on the camera pitch angle; The perspective relationship of the multi-view image is adjusted according to the non-uniform spatial attention mask, and the global features of the mobile device in the multi-view image are output based on the adjusted feature map.

8. The method according to claim 1 or 7, characterized in that, The detection of global features of the mobile device in the multi-view image includes: Determine the signal-to-noise ratio of the monocular photographed image; Based on the signal-to-noise ratio, a first confidence threshold for the monocular image adaptation is determined, wherein the first confidence threshold is positively correlated with the signal-to-noise ratio; Initial global features of the mobile device in the multi-view images are detected respectively, and initial global features with confidence scores greater than the first confidence score threshold are selected as the final global features of the mobile device in the corresponding view images.

9. The method according to claim 1, characterized in that, The step of outputting the identification result of the mobile device based on the global features and the local features includes: For a single-view image in the multi-view image, based on the global features, the first attribute information of the mobile device under a preset attribute category is identified. The preset attribute category includes one or more of the mobile device's brand, model, style, and color. For a single-view image in the multi-view image, the similarity between the local features and the reference features of each reference device in the preset feature library is calculated; wherein, the preset feature library includes multiple reference devices and the reference features of each reference device and the second attribute information under the preset attribute category; The reference device with the highest similarity is determined as a candidate device that matches the mobile device, and the second attribute information preset by the candidate device is obtained; Based on the first attribute information and the second attribute information corresponding to the multi-view images, the recognition result of the mobile device is output.

10. The method according to claim 9, characterized in that, The step of outputting the recognition result of the mobile device based on the first attribute information and the second attribute information corresponding to the multi-view images includes: For a single-view image in the multi-view image, determine the first confidence level of the first attribute information and the second confidence level of the second attribute information in the single-view image; The initial recognition result of the mobile device in the single-view image is determined based on the first confidence level and the second confidence level; For a single attribute category in the preset attribute categories, the weighted frequency of each initial recognition result corresponding to the single attribute category in the multi-view image is calculated; The initial identification result with the highest weighted frequency is determined as the final identification result of the mobile device under the single attribute category; Output the final identification result of the mobile device under each attribute category in the preset attribute category.

11. The method according to claim 10, characterized in that, The step of determining the initial recognition result of the mobile device in the single-view image based on the first confidence level and the second confidence level includes: If the first confidence level is greater than the second confidence threshold, the first attribute information is determined as the initial recognition result of the mobile device in the single-view image; If the first confidence level is less than or equal to the second confidence level threshold, and the second confidence level is greater than the second confidence level threshold, the second attribute information is determined as the initial recognition result of the mobile device in the single-view image; If both the first confidence level and the second confidence level are less than the second confidence threshold, the initial recognition result of the mobile device in the single-view image is determined based on the first sub-confidence level of different attribute categories in the first attribute information and the second sub-confidence level of the corresponding attribute category in the second attribute information.

12. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which is configured to execute the instructions to implement the method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-10.

14. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-10.