Method, device, equipment, medium and program product for processing service request
By generating depth maps and calculating binocular parallax using a deep learning model, stereoscopic video is output, solving the problems of insufficient accuracy and immersion in identity verification in remote video banking, and improving security and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INDUSTRIAL AND COMMERCIAL BANK OF CHINA
- Filing Date
- 2026-03-16
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies struggle to accurately identify users in remote video banking, and are prone to misidentification due to image flattening operations, reducing the accuracy and reliability of identity verification. Furthermore, traditional methods require additional equipment or result in insufficient immersion.
By acquiring monocular video frame data, generating depth maps using deep learning models, calculating binocular parallax, and combining display device parameters to output stereoscopic video, a naked-eye 3D effect is achieved, avoiding dependence on dedicated 3D hardware.
It enhances the security and reliability of remote services, improves the user experience, and provides a more immersive and realistic remote interaction experience.
Smart Images

Figure CN122492099A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial technology or other related fields, and in particular to a method, apparatus, device, medium and program product for processing business requests. Background Technology
[0002] With the development of fintech, banking services are trending towards online, intelligent, and scenario-based operations, leading to a growing demand from customers for remote video banking services. For example, in remote account opening, loan application interviews, or identity verification, customers and account managers need to communicate in real time via video.
[0003] Currently, banks' remote video services are mainly based on traditional 2D planar video call technology, which requires customers to connect via video through mobile banking or web browsers to communicate face-to-face with account managers to complete identity verification.
[0004] However, existing technologies struggle to capture the depth information of images, making it difficult for account managers to accurately determine if the person is the actual user. This can easily lead to misidentification due to image flattening operations (such as replacing real users with photos or videos), reducing the accuracy and reliability of identity verification and potentially impacting the security of business transactions. Summary of the Invention
[0005] This application provides a method, apparatus, device, medium, and program product for processing business requests, which aims to improve the security and reliability of handling remote business.
[0006] Firstly, this application provides a method for processing business requests, including:
[0007] Obtain remote service requests, and receive monocular video frame data sent by the client based on the remote service requests. The monocular video frame data includes user identity information.
[0008] Input monocular video frame data into a deep learning model to generate a depth map corresponding to the monocular video frame data;
[0009] The binocular disparity is determined based on the depth map to generate left and right eye images corresponding to the depth map;
[0010] Based on the left and right eye images and the configuration parameters of the display device, the adapted stereoscopic video is output to the display device through full-screen rendering technology.
[0011] Secondly, this application provides a service request processing apparatus, comprising:
[0012] The sending module is used to acquire remote service requests and receive monocular video frame data sent by the client based on the remote service requests. The monocular video frame data includes user identity information.
[0013] The processing module is used to input monocular video frame data into a deep learning model to generate a depth map corresponding to the monocular video frame data.
[0014] The processing module is also used to determine the binocular parallax based on the depth map in order to generate left and right eye images corresponding to the depth map;
[0015] The processing module is also used to output the adapted stereoscopic video to the display device based on the left and right eye images and the configuration parameters of the display device, using full-screen rendering technology.
[0016] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0017] The memory stores instructions that the computer executes;
[0018] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0020] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0021] The business request processing method, apparatus, device, medium, and program product provided in this application acquires remote business requests by receiving monocular video frame data containing user identity information sent by the client; inputs the monocular video frame data into a deep learning model to generate a corresponding depth map; determines the binocular parallax based on the depth map to generate corresponding left and right eye images; and outputs the adapted stereoscopic video to the display device based on the left and right eye images and the configuration parameters of the display device through full-screen rendering technology. In this process, the deep learning model calculates the scene's depth information from the monocular video and uses this to simulate human stereoscopic vision to generate left and right eye images with accurate parallax. Then, the device-adaptive full-screen rendering technology presents a naked-eye 3D effect. This effectively solves the problems of low accuracy in identity verification, weak liveness detection and anti-counterfeiting capabilities, and insufficient immersion in communication caused by the lack of depth information in traditional 2D videos during remote business processing, thereby improving the security, reliability, and user experience of remote financial services. Attached Figure Description
[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0023] Figure 1 A schematic diagram illustrating a scenario for the business request processing method provided in this application;
[0024] Figure 2 A flowchart illustrating the processing method for the business requests provided in this application;
[0025] Figure 3 A schematic diagram of the structure of the device for processing business requests provided in this application;
[0026] Figure 4 A schematic diagram of the structure of the electronic device provided in this application.
[0027] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0029] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0030] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0031] It should be noted that the business request processing methods, apparatus, equipment, media and program products provided in this application can be used in the field of financial technology or other related fields, or in any field other than the field of financial technology or other related fields. The application fields of the business request processing methods, apparatus, equipment, media and program products in this application are not limited.
[0032] When customers conduct business remotely with staff via mobile banking or online banking (such as transferring funds, resetting passwords, or handling complex transactions), identity verification steps are involved to ensure the security of the transaction. During identity verification, key characteristics of the customer are checked via video to ensure that subsequent transactions can proceed only if the identity information is verified correctly.
[0033] With the rapid development of fintech, banking services are accelerating their transformation towards online, intelligent, and scenario-based approaches. Customer demand for remote video banking services is growing, especially in scenarios such as identity verification, business consultation, and complex financial operations. Existing technologies, based on 2D video calling, capture key user features through a camera and transmit them to the receiving screen. However, images captured in traditional 2D video calls lack depth information, making it difficult to determine the user's identity through stereoscopic vision. This makes identity verification susceptible to being deceived by photos or videos, failing to meet users' requirements for a realistic interactive experience and security. For example, in remote account opening, loan interviews, or identity verification, customers and account managers need to communicate in real-time via video. However, because 2D video cannot present depth information (e.g., three-dimensional information), account managers struggle to accurately determine the other party's identity. Misidentification can easily occur due to flattened image manipulation (e.g., replacing real users with photos or videos), reducing the accuracy and reliability of identity verification and potentially affecting the secure processing of transactions. Therefore, existing technologies are prone to misidentification, easily reducing the accuracy and reliability of identity verification. Furthermore, as users' demand for immersive interactive experiences is gradually increasing, the traditional 2D video methods used in existing technologies are prone to lacking a sense of depth, making it difficult for both parties in the call to perceive each other's spatial location information and facial micro-expressions, which can easily affect the realism of the communication and reduce the naturalness and trust in the communication.
[0034] Furthermore, existing video call technologies require users to wear 3D glasses (e.g., polarized glasses, active shutter glasses) or use VR devices (e.g., head-mounted displays) to generate left and right eye images through split-screen technology or parallax control on the device to achieve a stereoscopic effect. However, existing technologies require users to wear specialized equipment, increasing usage costs and operational complexity. Since 3D glasses can easily cause dizziness, and VR devices require a fixed wearing position, their use is limited by the scenarios in which they are used, and they also place high demands on the terminal's display capabilities (e.g., split-screen accuracy, parallax control). Existing methods are difficult to adapt to ordinary displays or mobile devices. Additionally, existing technologies typically rely on simple geometric modeling or template matching to convert 2D images into 3D images. This method struggles to dynamically calculate depth information in complex scenes, easily leading to distortion and latency issues in the resulting 3D images, reducing the real-time performance and accuracy of identity verification during business transactions.
[0035] The service request processing method of this application involves acquiring a remote service request, receiving monocular video frame data sent by the client based on the remote service request (the monocular video frame data includes user identity information), inputting the monocular video frame data into a deep learning model to generate a depth map corresponding to the monocular video frame data, determining the binocular parallax based on the depth map to generate left and right eye images corresponding to the depth map, and outputting the adapted stereoscopic video to the display device based on the left and right eye images and the configuration parameters of the display device using full-screen rendering technology. This solution employs a deep convolutional neural network to extract spatial features of video frames through multi-layer convolution and dynamically generate high-precision depth maps. To achieve end-to-end real-time processing, a real-time communication protocol is used for real-time transmission and processing of video frames. Simultaneously, full-screen rendering technology directly outputs the binocular images to a regular display, avoiding reliance on dedicated 3D hardware. Furthermore, by optimizing the binocular parallax calculation logic through algorithms, it is ensured that the generated left and right eye images conform to the stereoscopic perception law of the human eye. This enables video call technology that can generate stereoscopic visual effects without the need for additional equipment, providing users with a more immersive and realistic remote interaction experience, improving the security and credibility of handling remote business, and enhancing the user experience.
[0036] The method for processing business requests provided in this application is intended to solve the above-mentioned technical problems in the prior art.
[0037] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0038] Figure 1A schematic diagram illustrating the scenario of the business request processing method provided in this application, such as... Figure 1 As shown, the user initiates a remote service request through a user terminal (i.e., the client) and calls the camera to collect monocular video frame data. The user terminal establishes a bidirectional data connection with the server through the network to send monocular video frame data to the server. After receiving the video stream, the server schedules its internal data processing flow. After processing, the final stereoscopic video signal is output to the display device at the review end to display the naked-eye 3D stereoscopic video image to the bank's review personnel.
[0039] Figure 2 A flowchart illustrating the processing method for the business requests provided in this application, as shown below. Figure 2 As shown, the method includes:
[0040] S201. Obtain remote service requests and receive monocular video frame data sent by the client based on the remote service requests.
[0041] More specifically, it acquires remote service requests and receives monocular video frame data sent by the client based on the remote service requests. The monocular video frame data includes user identity information.
[0042] Optionally, monocular video frame data is two-dimensional image data acquired by a single camera. Monocular video frame data includes continuous image frames in a time series, such as an RGB format video stream acquired by the camera.
[0043] In one possible embodiment, a user triggers a remote transaction request when opening a remote account using their client (e.g., a smartphone with a mobile banking app installed). This request is sent to the bank's backend server. Upon receiving the request, the server sends an instruction to the client to initiate a video verification process. In response, the client uses its front-facing camera to capture a real-time RGB video stream containing the user's face and identification document at 30 frames per second. The client then encodes this video stream into continuous monocular video frames and sends them over the network to the server for further processing.
[0044] Optionally, before receiving monocular video frame data sent by the client, a bidirectional connection between the client and the server is established through a communication protocol; wherein, the communication protocol includes a protocol for realizing bidirectional real-time communication, or a protocol for segmented transmission of video frame data.
[0045] Alternatively, a communication protocol refers to standardized rules for data transmission between a client and a server. For example, the WebSocket protocol (i.e., a protocol for enabling bidirectional real-time communication) uses long-lived connections for bidirectional real-time communication. The TCP / IP protocol (i.e., a protocol for segmenting video frame data) transmits video frames by splitting them into multiple data packets.
[0046] Optionally, the client and server establish a bidirectional long connection through a communication protocol (e.g., WebSocket protocol). After the client captures a monocular video frame, it encapsulates the video frame data (e.g., RGB format) into a binary message and sends it to the server in real time.
[0047] In one possible implementation, after the user triggers a remote service request but before the video stream data transmission begins, the client first initiates a connection request to the server (e.g., a financial server). The server dynamically selects a communication protocol based on the current network load policy or the network latency information reported by the client. For example, if the current environment is determined to be a high-quality local area network, the WebSocket protocol is preferentially used to establish a bidirectional long connection to achieve low-latency real-time transmission of video frames; if the client is detected to be in an unstable mobile network environment, the TCP / IP protocol is used to divide each frame of video data into multiple data packets, number them, and then transmit them to ensure data integrity and order. The data packet reassembly step is performed only after all data packets have successfully arrived at the server.
[0048] For example, in high-latency networks, the TCP / IP protocol is selected for segmented transmission to ensure the integrity of transmitted data; in low-latency scenarios, the WebSocket protocol is selected for real-time transmission to improve the real-time performance of data transmission, thereby expanding the application scenarios of the technical solution.
[0049] This embodiment provides flexibility at the communication layer by establishing a bidirectional connection between the client and the financial server before video data transmission. Appropriate communication protocols can be selected based on different network environments and business needs, thereby improving the real-time performance and reliability of the video verification process under network fluctuations and ensuring the continuity and stability of financial services.
[0050] S202. Input the monocular video frame data into the deep learning model to generate the depth map corresponding to the monocular video frame data.
[0051] More specifically, deep learning models refer to algorithmic models based on multi-layer neural networks used to extract spatial features from monocular video frame data and generate depth maps. For example, deep convolutional neural networks. Depth maps represent the distance information of each pixel in an image from the observer, quantizing depth values using grayscale values or floating-point numbers; for example, using a range of 0-1 to represent near and far relationships.
[0052] Optionally, the monocular video frame data is input into a deep learning model to generate a depth map corresponding to the monocular video frame data. Specifically, this includes: inputting the monocular video frame data into a deep learning model, performing multi-layer convolution operations on the monocular video frame through the deep learning model to extract spatial features of the user's face and document area; generating a depth value for each pixel based on the extracted spatial features, and generating a depth map based on each depth value. The depth map is used to distinguish between real users and attack vectors.
[0053] Optionally, multi-layer convolution operation refers to extracting spatial features of an image step by step through multiple convolutional layers. For example, using convolutional kernels of different sizes such as 3×3 and 5×5 to extract local details and global structure.
[0054] Optionally, spatial features refer to geometric or textural information contained in an image, such as facial contours, background details, etc.
[0055] Optionally, the depth value is used to represent the distance information of the pixel from the observer, for example, using a 0-1 range to quantize near-far relationships.
[0056] In one possible implementation, a deep learning model is used to extract spatial features from monocular video frame data layer by layer through multi-layer convolutional operations. For example, edge information is extracted through the first convolutional layer, texture information through the second convolutional layer, and object contours through the third convolutional layer. Then, based on the extracted spatial features, depth values for each pixel are generated through fully connected layers or upsampling layers, and finally, these are integrated to generate a depth map. This process can dynamically adapt to depth changes in complex scenes, thereby improving the accuracy of the depth map.
[0057] In one possible embodiment, after the server receives monocular video frame data from the client terminal, it inputs the monocular video frame data into a pre-trained deep convolutional neural network model. This model first extracts primary edge information (e.g., the boundary between a face and the background) from the video frame through a first convolutional layer (e.g., using a 3x3 convolutional kernel). Then, it further extracts intermediate texture features (e.g., skin texture, printed patterns on an ID card) based on the edge information through a second convolutional layer (e.g., using a 5x5 convolutional kernel). Further, it integrates the features extracted from the first two layers through a third convolutional layer (e.g., using a 7x7 convolutional kernel) to extract high-level object contour features (e.g., complete facial contours, the overall three-dimensional shape of the ID card). Based on the spatial features extracted layer by layer, an upsampling layer restores the image to its original resolution, and a fully connected layer maps each pixel to generate a depth value between 0 and 1 (where 0 represents nearest and 1 represents farthest). Finally, the depth values of all pixels are integrated to form a high-precision depth map. This process enables the model to dynamically adapt to complex business scenarios with different users, lighting conditions, and backgrounds, ensuring the accuracy of depth estimation.
[0058] For example, in the scenario of remote bank account opening, a deep learning model can be used to accurately distinguish between a user's real face and a photo attack, thereby reducing the probability of misidentification. At the same time, detailed information is extracted through multi-layer convolution, which enhances the model's ability to perceive complex scenes, making the generated binocular images more consistent with the stereoscopic perception of the human eye, and improving the clarity and stability of the image.
[0059] This embodiment extracts spatial features of key regions such as the user's face and identification documents through multi-layer convolutional operations of a deep learning model, and generates a depth map based on these features to distinguish real users from attack vectors (such as photos or screenshots). By selectively extracting features from key regions, this solution enables the generated depth map to more accurately reflect the essential differences in three-dimensional structure between real faces and forged vectors, helping to improve the accuracy of real user detection and anti-fraud capabilities, and providing technical protection for the security of financial services.
[0060] Optionally, multi-layer convolution operations are performed on monocular video frames using a deep learning model. Specifically, this includes: extracting image features of different scales in parallel from the monocular video frames using multi-scale convolution kernels of the deep learning model; and integrating image features of different scales using a feature fusion layer in the deep learning model to obtain multi-scale information for depth computation by integrating local details and global structure.
[0061] Optionally, multi-scale convolution kernels: convolution kernels of different sizes (e.g., 3×3, 5×5, 7×7) are used to extract image features in parallel.
[0062] Optionally, the feature fusion layer refers to a network layer used to integrate features at different scales, such as skip connections or weighted summation.
[0063] In one possible embodiment, the step of processing the same input image frame in parallel using multi-scale convolutional kernels through a deep learning model includes: using a 3×3 convolutional kernel to focus on extracting local detail features (e.g., subtle textures of the user's eyes, microtexture on an ID card); using a 5×5 convolutional kernel to extract mid-scale features (e.g., the outline of the bridge of the nose, the corner structure of the ID card); and using a 7×7 convolutional kernel to capture global structural features (e.g., the pose of the entire head, the relative position of the ID card in the image). Subsequently, a feature fusion layer (e.g., using skip connection technology) concatenates and weights the feature maps of the above three different scales in the channel dimension, thereby integrating a comprehensive spatial feature containing multi-scale information from fine-grained details to macroscopic structure.
[0064] For example, in a bank identity verification scenario, a deep learning model accurately distinguishes between a user's real face and a photo attack, while retaining more detailed information, making the image more consistent with the human eye's stereoscopic perception. This embodiment enhances the model's ability to perceive complex scenes through multi-scale feature fusion.
[0065] This embodiment employs multi-scale convolutional kernels to extract features from monocular video frame data in parallel, and integrates multi-scale features through a feature fusion layer to simultaneously capture multi-level information in the image, from local details (e.g., skin texture, microtext on documents) to global structures (e.g., facial contours, overall shape of documents). This improves the comprehensiveness and robustness of the final generated feature representation, and helps to enhance the depth perception accuracy of deep learning models for complex scenes and targets of different scales.
[0066] Optionally, based on the extracted spatial features, a depth value for each pixel is generated, specifically including: inputting the spatial features into a fully connected layer of a deep learning model; processing the spatial features through a non-linear activation function in the fully connected layer, and outputting a depth value for each pixel.
[0067] Optionally, a fully connected layer refers to a layer in a neural network that connects all features of the previous layer with all neurons of the current layer. Fully connected layers are used to integrate spatial features to generate depth values.
[0068] Optionally, a nonlinear activation function refers to a mathematical function that introduces a nonlinear relationship, such as ReLU or Sigmoid. Nonlinear activation functions are used to enhance the expressive power of deep learning models.
[0069] In one possible embodiment, spatial features are integrated through a fully connected layer, where each neuron is connected to all features in the previous layer. After processing by a non-linear activation function, the depth value of each pixel is output. For example, when the input features are 128-dimensional vectors, the fully connected layer outputs 256-dimensional depth values, and finally, a depth map with the same size as the input image is generated through upsampling. This embodiment improves the accuracy of depth map generation by combining fully connected layers and non-linear activation functions. For example, by using a non-linear activation function, distortion caused by linear feature superposition is avoided, improving the fit between the depth values and the actual scene requirements, thereby optimizing the disparity calculation effect of binocular images. The spatial features (e.g., 128-dimensional feature vectors) extracted and flattened by the convolutional layer are input to the fully connected layer of the model. Each neuron in this fully connected layer is connected to each dimension of the input vector, and all spatial features are globally integrated through matrix operations. The integrated features are then processed by the non-linear activation function ReLU to introduce a non-linear transformation, avoiding simple linear superposition of features, thereby enhancing the model's ability to fit complex depth relationships. The processed output is a 256-dimensional vector, where each value represents the depth information of the corresponding region after image downsampling. Finally, this 256-dimensional depth vector is restored to the same size as the original input image through an upsampling layer (e.g., deconvolution operation) to generate the final depth map. This process ensures the accuracy of depth value calculation, making the generated depth map more closely resemble the real 3D scene.
[0070] This embodiment achieves accurate mapping from spatial features to depth values through a network structure using fully connected layers and nonlinear activation functions. By leveraging the feature integration capabilities of fully connected layers and the fitting capabilities of nonlinear activation functions to the complex relationships, the accuracy and robustness of depth value calculation are improved, thereby further enhancing the detail richness of the final generated depth map and improving the naturalness of depth map transitions.
[0071] Optionally, during the generation of the depth value of each pixel, the depth calculation weights of different regions are dynamically adjusted through an attention mechanism to obtain weight-optimized spatial features; the depth value of each pixel is generated based on the weight-optimized spatial features.
[0072] Optionally, the attention mechanism is used to highlight features of key regions through dynamic weight allocation.
[0073] In one possible implementation, the depth calculation process is dynamically adjusted by calculating the weight of each region based on an attention mechanism. For example, if the user's facial region has a higher weight and the background region has a lower weight, the deep learning model prioritizes enhancing the depth calculation accuracy of the face to suppress background interference.
[0074] In one possible implementation, an attention module is integrated into the decoder portion of the model. This attention module obtains global information for each feature channel by performing global average pooling on the spatial features, and then learns the importance weights for each channel through a fully connected network. Finally, these weights are multiplied back onto the original feature map to achieve dynamic weight allocation. In this process, feature channels related to key business areas such as the user's face and held identification are assigned higher weights, while feature channels related to cluttered backgrounds are suppressed. The deep learning model calculates depth values based on the spatial features with the optimized weights, thereby improving the accuracy of depth estimation for key business areas.
[0075] For example, in identity verification scenarios, by improving the depth calculation accuracy of key areas and enhancing the three-dimensionality of the user's face, the accuracy of customer managers' review of customers can be further improved, while reducing the interference of complex backgrounds on binocular images.
[0076] In this embodiment, an attention mechanism is introduced during the depth value generation process to dynamically adjust the depth calculation weights for different regions. This allows the model to automatically focus on key areas in the image (such as the user's face or a handheld document) while suppressing interference from irrelevant backgrounds. Thus, with limited computing resources, the accuracy and efficiency of depth estimation for key business target areas are improved, enhancing the three-dimensionality and realism of these areas.
[0077] Optionally, before generating the depth map corresponding to the monocular video frame data, temporal information processing is performed on the continuous video frames to correct local errors in the depth map caused by minor user movements, resulting in a temporally optimized video frame sequence; the temporally optimized video frame sequence is then input into a deep learning model to generate the corresponding depth map.
[0078] Optionally, temporal information processing is used to analyze the temporal consistency of consecutive video frames to optimize the processing results of the current frame.
[0079] In one possible implementation, before generating the depth map of the current frame, a deep learning model analyzes the depth information of the previous few frames, and a temporal information processing module (e.g., an LSTM network) predicts the depth change trend of the current frame data, correcting local depth errors caused by rapid movement or occlusion. For example, when the user moves quickly, the deep learning model predicts a reasonable depth value for the current frame based on the depth information of the previous frame to avoid jump distortion.
[0080] In one possible embodiment, when processing a real-time video stream, the server not only inputs the current frame into the model but also caches the depth maps generated from the previous N frames (e.g., the previous 4 frames) and their corresponding RGB frames, and inputs them together into the temporal information processing module. This LSTM network learns the motion patterns and depth variation patterns of the video over time and, based on the sequence information of previous frames, predicts and optimizes the depth value of the current frame. For example, when a user suddenly turns their head quickly, causing a temporary occlusion or motion blur in part of their face, the depth map calculated solely from the current frame's information is prone to errors or omissions. In this case, the temporal information processing module predicts and fills in the abnormal areas of the current frame based on the clear and stable facial depth information from the previous 4 frames, thereby effectively correcting local depth errors caused by rapid movement or occlusion. Here, N is a positive integer greater than 1.
[0081] For example, when the user moves quickly or the background changes suddenly, the model adaptively adjusts the depth values to avoid abrupt distortion in the generated binocular images and improve the smoothness of the image. Therefore, processing temporal information can improve the stability of depth maps in dynamic scenes.
[0082] Before generating the depth map, this embodiment incorporates temporal information processing of consecutive video frames to correct local errors in the depth map caused by minor user movements. This approach leverages the temporal continuity of video to effectively smooth out jitter or noise that may occur in single-frame depth estimation, improving the stability and coherence of the depth map in dynamic scenes. This results in smoother generated stereoscopic video footage, reducing the possibility of visual discomfort or misjudgment caused by jumps in depth information.
[0083] S203. Determine the binocular disparity based on the depth map to generate left and right eye images corresponding to the depth map.
[0084] More specifically, binocular parallax refers to the horizontal displacement between the left and right eye images. This application simulates human stereoscopic vision through binocular parallax, for example, by adjusting the offset of the left and right eye images through depth values.
[0085] Optionally, after receiving monocular video frame data on the server side, a pre-trained deep learning model is invoked to process the video frames frame by frame to extract spatial features and generate depth maps. Subsequently, the disparity between the left and right eye images is calculated based on the depth maps to generate left and right eye images.
[0086] In one possible embodiment, the depth map generated in step S202 is obtained through a disparity calculation module within the server. Based on preset virtual binocular camera parameters (e.g., interpupillary distance), a horizontal disparity offset is calculated for each pixel in the depth map. The calculation principle is that the smaller the depth value of a pixel (the closer the object), the larger the absolute value of the disparity offset between the left and right eye images, simulating the human eye's visual principle of "larger disparity for near objects." Then, using the original 2D video frame as a reference, the server offsets pixels to the left and right respectively according to the calculated disparity map, thereby generating two images representing the left and right eye views, i.e., left and right eye images. The generated left and right eye images are used to synthesize the final naked-eye 3D video stream and output it to the reviewer's display device, enabling the reviewer to observe a user image with a prominent stereoscopic effect, thus allowing for more accurate identity verification.
[0087] S204. Based on the left and right eye images and the configuration parameters of the display device, the adapted stereoscopic video is output to the display device through full-screen rendering technology.
[0088] Optionally, full-screen rendering technology refers to the technology of directly rendering image data to the full-screen area of the display device, for example, using the OpenCV library to implement horizontal offset rendering of binocular images.
[0089] Optionally, the server uses full-screen rendering technology (such as the OpenCV library) to render the left and right eye images to the display device in a horizontal offset manner. This simulates human stereoscopic vision by leveraging the parallax between the left and right eye images, achieving a naked-eye 3D effect. The entire process ensures real-time transmission and processing of video frames through communication protocols, dynamically adapts to depth changes in complex scenes using deep learning models, and adapts to ordinary displays through full-screen rendering technology, eliminating the need for dedicated 3D equipment and reducing costs.
[0090] More specifically, based on the left and right eye images and the configuration parameters of the display device, the adapted stereoscopic video is output to the display device through full-screen rendering technology. Specifically, this includes: dynamically adjusting the parallax offset of the left and right eye images based on the resolution and physical size of the display device to obtain adapted left and right eye images; and generating a binocular composite image based on the adapted left and right eye images through full-screen rendering technology, so as to output the binocular composite image to the display device.
[0091] Optionally, the parallax offset refers to the horizontal displacement between the left and right eye images, used to simulate human stereoscopic vision.
[0092] Optionally, a reasonable parallax offset can be calculated based on the display device's resolution (e.g., 1920×1080) and physical size (e.g., 24 inches). For example, for high-resolution screens, the parallax offset is automatically reduced to avoid excessive stereoscopic effect that could cause user discomfort; for low-resolution screens, the offset is appropriately increased to enhance immersion.
[0093] In one possible embodiment, the server's image rendering module acquires the left and right eye images generated in step S203, while simultaneously reading the configuration parameters of the target display device from the system configuration library (e.g., the reviewer is using a 27-inch monitor with a resolution of 2560x1440). Then, the rendering module dynamically calculates an optimal parallax offset suitable for the target display device based on its physical size and pixel density. More specifically, to avoid visual fatigue caused by excessive parallax on high PPI (pixels per inch) screens, the image rendering module reduces the final offset compared to a standard value and performs pixel-level horizontal displacement processing on the original left and right eye images based on the adjusted offset, thereby obtaining a new, adapted pair of left and right eye images. Finally, the server calls the OpenCV library's full-screen rendering interface to combine these adapted images into a single binocular composite image, which is then output to the reviewer's display device for full-screen display. When reviewers see this image, their eyes will perceive a strong sense of depth due to parallax, allowing them to more clearly judge the three-dimensional contours of the user's face and the authenticity of the document, thus improving the accuracy and immersive experience of remote verification.
[0094] This embodiment generates a binocular composite image by dynamically adjusting the parallax offset based on the physical parameters of the display device. It achieves adaptability to different terminal display devices (such as high-resolution teller screens, ordinary PC monitors, and mobile devices), generating comfortable, dizzying stereoscopic visual effects on various devices, thus expanding the applicability of this technology in various banking business channels and enhancing the user experience.
[0095] The business request processing method provided in this application involves acquiring monocular video containing identity information based on a remote business request, then generating a depth map using a deep learning model, and generating left and right eye images based on the depth map. Finally, it combines display device parameters to output an adapted stereoscopic video. This upgrades traditional two-dimensional planar video verification to stereoscopic video verification with depth information in remote financial transactions, providing customer service or auditing personnel with a more realistic and immersive visual experience. This enhances the ability to perceive user identity features (e.g., micro-expressions) and the authenticity of documents in a stereoscopic manner, thereby strengthening the security and credibility of remote transactions.
[0096] Figure 3 A schematic diagram of the structure of the service request processing device provided in this application is shown below. Figure 3As shown, the service request processing device 30 provided in this embodiment includes:
[0097] The sending module 301 is used to acquire remote service requests and receive monocular video frame data sent by the client based on the remote service requests. The monocular video frame data includes user identity information.
[0098] The processing module 302 is used to input monocular video frame data into a deep learning model to generate a depth map corresponding to the monocular video frame data.
[0099] The processing module 302 is also used to determine the binocular parallax based on the depth map in order to generate left and right eye images corresponding to the depth map;
[0100] The processing module 302 is also used to output the adapted stereoscopic video to the display device based on the left and right eye images and the configuration parameters of the display device through full-screen rendering technology.
[0101] Optionally, the processing module 302 is also used to input monocular video frame data into a deep learning model, and perform multi-layer convolution operations on the monocular video frame through the deep learning model to extract the spatial features of the user's face and document area.
[0102] Based on the extracted spatial features, a depth value is generated for each pixel, and a depth map is generated based on each depth value. The depth map is used to distinguish between real users and attack vectors.
[0103] Optionally, the processing module 302 is also used to input spatial features into the fully connected layer of the deep learning model;
[0104] Spatial features are processed by a non-linear activation function in a fully connected layer, and the depth value of each pixel is output.
[0105] Optionally, the processing module 302 is further configured to perform temporal information processing on the continuous video frames before generating the depth map corresponding to the monocular video frame data, so as to correct the local error of the depth map caused by the user's small actions and obtain a temporally optimized video frame sequence.
[0106] The time-optimized video frame sequence is input into the deep learning model to generate the corresponding depth map.
[0107] Optionally, the processing module 302 is also used to extract image features of different scales in a monocular video frame in parallel using the multi-scale convolutional kernels of a deep learning model;
[0108] By integrating image features at different scales through feature fusion layers in deep learning models, multi-scale information for deep computation can be obtained by integrating local details and global structure.
[0109] Optionally, the processing module 302 is also used to dynamically adjust the depth calculation weights of different regions through an attention mechanism during the process of generating the depth value of each pixel, so as to obtain the weight-optimized spatial features.
[0110] The depth value of each pixel is generated based on the weighted spatial features.
[0111] Optionally, the processing module 302 is also used to dynamically adjust the parallax offset of the left and right eye images based on the resolution and physical size of the display device to obtain adapted left and right eye images;
[0112] Based on the adapted left and right eye images, a binocular composite image is generated using full-screen rendering technology, and then output to the display device.
[0113] Optionally, the processing module 302 is also used to establish a bidirectional connection between the client and the financial service provider via a communication protocol before receiving monocular video frame data sent by the client;
[0114] The communication protocols include protocols for enabling bidirectional real-time communication or protocols for transmitting video frame data in segments.
[0115] The service request processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0116] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. Figure 4 As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.
[0117] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.
[0118] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0119] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0120] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0121] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0122] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0123] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0124] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0125] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0126] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0127] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0128] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0129] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.
[0130] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0131] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0132] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0133] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0134] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method of processing a service request, characterized by, include: Obtain a remote service request, and based on the remote service request, receive monocular video frame data sent by the client, wherein the monocular video frame data includes user identity information; The monocular video frame data is input into a deep learning model to generate a depth map corresponding to the monocular video frame data. The binocular disparity is determined based on the depth map to generate left and right eye images corresponding to the depth map; Based on the left and right eye images and the configuration parameters of the display device, the adapted stereoscopic video is output to the display device using full-screen rendering technology.
2. The method of claim 1, wherein, The monocular video frame data is input into a deep learning model to generate a depth map corresponding to the monocular video frame data, specifically including: The monocular video frame data is input into a deep learning model, and the deep learning model performs multi-layer convolution operations on the monocular video frame to extract the spatial features of the user's face and document area. Based on the extracted spatial features, a depth value for each pixel is generated, and a depth map is generated based on each depth value. The depth map is used to distinguish between real users and attack vectors.
3. The method according to claim 2, characterized in that, Based on the extracted spatial features, a depth value for each pixel is generated, specifically including: The spatial features are input into the fully connected layer of the deep learning model; The spatial features are processed by a non-linear activation function in the fully connected layer to output the depth value of each pixel.
4. The method according to claim 1, characterized in that, Also includes: Before generating the depth map corresponding to the monocular video frame data, the continuous video frames are processed for timing information to correct the local error in the depth map caused by the user's small actions, so as to obtain a timing-optimized video frame sequence. The time-optimized video frame sequence is input into the deep learning model to generate the corresponding depth map.
5. The method according to claim 2, characterized in that, The deep learning model is used to perform multi-layer convolution operations on the monocular video frame, specifically including: The deep learning model uses multi-scale convolutional kernels to extract image features of different scales in parallel from the monocular video frame; The feature fusion layer in the deep learning model integrates image features at different scales to obtain multi-scale information for deep computation by integrating local details and global structure.
6. The method according to claim 3, characterized in that, Also includes: During the generation of the depth value of each pixel, the depth calculation weights of different regions are dynamically adjusted through an attention mechanism to obtain the spatial features after weight optimization. The depth value of each pixel is generated based on the spatial features optimized by the weights.
7. The method according to claim 1, characterized in that, Based on the left and right eye images and the configuration parameters of the display device, the adapted stereoscopic video is output to the display device using full-screen rendering technology, specifically including: Based on the resolution and physical size of the display device, the parallax offset of the left and right eye images is dynamically adjusted to obtain adapted left and right eye images; Based on the adapted left and right eye images, a binocular composite image is generated using full-screen rendering technology, and then the binocular composite image is output to a display device.
8. The method according to claim 1, characterized in that, Also includes: Before receiving monocular video frame data sent by the client, a two-way connection between the client and the financial service provider is established through a communication protocol. The communication protocol includes a protocol for enabling bidirectional real-time communication or a protocol for segmented transmission of video frame data.
9. A service request processing apparatus, comprising: The sending module is used to acquire remote service requests and receive monocular video frame data sent by the client based on the remote service requests. The monocular video frame data includes user identity information. The processing module is used to input the monocular video frame data into a deep learning model to generate a depth map corresponding to the monocular video frame data. The processing module is further configured to determine binocular parallax based on the depth map in order to generate left and right eye images corresponding to the depth map; The processing module is also used to output the adapted stereoscopic video to the display device based on the left and right eye images and the configuration parameters of the display device using full-screen rendering technology.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 8.