Video call apparatus, method, and vehicle for a vehicle

By installing an image capture, processing, and display system in the vehicle, the facial images of infants and drivers are centered, solving the problem of infants' crying distracting the driver and improving driving safety and interactive experience.

CN122340233APending Publication Date: 2026-07-03MOBILITY ASIA SMART TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MOBILITY ASIA SMART TECH CO LTD
Filing Date
2025-01-03
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

In a vehicle, the crying of an infant can distract the driver and affect driving safety. Furthermore, the driver's attempts to soothe the infant may lead to poor driving behavior, increasing the risk of traffic accidents.

Method used

A video call device is provided, which captures facial images of occupants in a vehicle cabin through an image capture device, performs image correction by a processor and places the facial images in the center area, and displays the corrected image sequence on a display in the vehicle cabin, ensuring that the driver in the front seat and the infant in the back seat can clearly see each other's facial images.

Benefits of technology

It effectively reduces anxiety and distraction during driving, improves driving safety, enhances the interaction between the driver and infants, and ensures that the driver can monitor the status of infants in the back seat in real time, while also providing comfort to the infants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122340233A_ABST
    Figure CN122340233A_ABST
Patent Text Reader

Abstract

A video conferencing device, method, and vehicle for use in a vehicle are provided. The video conferencing device includes: an image capture device configured to capture an initial image sequence including facial images of occupants within the vehicle cabin; a processor communicatively connected to the image capture device and configured to generate a corrected image sequence based on the initial image sequence, wherein facial images of the occupants are located in the central region of each image in the corrected image sequence; and a display communicatively connected to the processor and configured to display the corrected image sequence to other occupants within the vehicle cabin.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and in particular to a video calling device, method, vehicle, computer equipment, non-transitory computer-readable storage medium, and computer program product for vehicles. Background Technology

[0002] Vehicle safety is a primary goal in the automotive industry. Factors affecting vehicle safety include not only those inherent to the vehicle itself but also those of the driver. For example, in situations where parents are driving and an infant is in the back seat or a rear-seat car seat, the infant's crying can distract the driver. The driver may turn their head or frequently check the rearview mirror to observe the infant's condition, thus failing to focus on the road. Furthermore, the infant's crying can trigger anxiety, tension, or frustration in the driver, potentially interfering with driving decisions and prolonging reaction time, making it difficult for the driver to remain calm and make rational judgments. Additionally, the driver's actions of reaching out to the back seat to soothe the infant can not only be distracting but also negatively impact steering wheel control. Moreover, in cases of severe crying, the driver may need to make an emergency stop to attend to the infant, which, in fast-moving traffic or where suitable parking is lacking, could increase the risk of a traffic accident.

[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0004] In view of this, the present disclosure provides a video calling device, method, vehicle, computer equipment, non-transitory computer-readable storage medium, and computer program product for vehicles to improve the problems in the related art.

[0005] According to a first aspect of this disclosure, a video call device for a vehicle is provided, comprising: an image capture device configured to capture an initial image sequence including facial images of occupants within the vehicle cabin; a processor communicatively connected to the image capture device and configured to generate a corrected image sequence based on the initial image sequence, wherein facial images of the occupants are located in the central region of each image in the corrected image sequence; and a display communicatively connected to the processor and configured to display the corrected image sequence to other occupants within the vehicle cabin.

[0006] According to a second aspect of this disclosure, a vehicle is provided, including the aforementioned video call device for a vehicle.

[0007] According to a third aspect of this disclosure, a video call method for a vehicle is provided, comprising: capturing an initial image sequence including facial images of occupants in the vehicle cabin; generating a corrected image sequence based on the initial image sequence, wherein facial images of the occupants are located in the central region of each image in the corrected image sequence; and displaying the corrected image sequence to other occupants in the vehicle cabin.

[0008] According to a fourth aspect of this disclosure, a computer device is provided, comprising: a memory, a processor, and a computer program stored on the memory, wherein the processor is configured to execute the computer program to implement the above-described video call method for a vehicle.

[0009] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the above-described video call method for a vehicle.

[0010] According to a sixth aspect of this disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the above-described video call method for a vehicle.

[0011] According to embodiments of this disclosure, it can be ensured that the driver in the front seat and the infant in the back seat can clearly see each other on the display. This technology allows the driver, even while sitting in the front of the vehicle, to monitor the infant's condition in the back seat in real time, effectively reducing anxiety and distraction during driving and improving driving safety. The infant in the back seat can also see the driver's face in real time on the display, thus receiving effective reassurance. Furthermore, centering the image enhances realism, especially for the infant's visual experience, and strengthens the sense of interaction between the two. Attached Figure Description

[0012] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0013] Figure 1 This is a block diagram illustrating a video calling device for a vehicle according to an exemplary embodiment of the present disclosure; Figure 2 This is a block diagram illustrating a video calling device for a vehicle according to another exemplary embodiment of the present disclosure; Figure 3 This is a flowchart illustrating a video call method for a vehicle according to an exemplary embodiment of the present disclosure; Figure 4 This is a structural block diagram of an exemplary computing device that can be applied to exemplary embodiments of the present disclosure. Detailed Implementation

[0014] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0015] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0016] As mentioned above, when parents are driving a vehicle and infants are in the back seat or in a rear child safety seat, the crying of infants in the back seat may distract the driver or interfere with driving decisions and prolong the driver's reaction time, which may increase the risk of traffic accidents and endanger vehicle driving safety.

[0017] In view of this, the present disclosure provides a video calling device, method, vehicle, computer equipment, non-transitory computer-readable storage medium, and computer program product for vehicles to improve the problems in the related art.

[0018] The following is a reference to the appendix. Figure 1 and Figure 2 A video call device for a vehicle according to an embodiment of the present disclosure will be described.

[0019] Figure 1 This is a block diagram illustrating a video calling device for a vehicle according to an exemplary embodiment of the present disclosure.

[0020] like Figure 1 As shown, the video calling device 100 for a vehicle includes: Image capture device 101 is configured to capture an initial image sequence, which includes facial images of occupants inside the vehicle cabin; Processor 102, communicatively connected to image capture device 101, is configured to generate corrected image sequences based on an initial image sequence, wherein facial images of the occupants are located in the central region of each image in the corrected image sequence; and Display 103 is communicatively connected to processor 102 and is configured to display a calibrated sequence of images to other occupants within the cabin.

[0021] The image capture device 101 may be, for example, a high-resolution camera module. The image capture device 101 can be positioned in the center of the vehicle cabin to capture images of the front and / or rear of the cabin. For example, the image capture device 101 may capture images of only the front of the cabin, only the rear of the cabin, or both. The camera in the camera module may, for example, have autofocus and image stabilization functions to cope with vibrations during vehicle movement and the need for clear imaging at different distances. For example, using a lens module based on optical image stabilization technology, combined with an autofocus algorithm, can quickly adjust the focus when the vehicle accelerates, decelerates, or experiences bumps, ensuring the clarity of each image in the initial image sequence. The camera may be equipped with an infrared fill light to clearly acquire image sequences of occupants even in low-light conditions (such as night driving), and the brightness of the infrared fill light should be automatically adjustable to avoid visual interference to the occupants.

[0022] In the example, processor 102 can first extract feature points from the facial image (e.g., key facial features such as eyes, nose, and mouth). The positions and descriptors of these feature points (used to describe pixel information around the feature points) can serve as the basis for subsequent matching and correction. For each image, processor 102 can match the extracted facial feature points with a preset standard facial template (containing the standard positions and distribution of facial feature points). A distance-based matching method, such as Euclidean distance, can be used. The distance between each extracted feature point and the template feature points is calculated, and the matching pair with the smallest distance is selected. Based on the matched feature point pairs, the transformation required to correct the facial image to the central region is determined by calculating an affine transformation matrix or homography matrix, thereby transforming the facial image from its current position to the corresponding position in the image's central region.

[0023] In the example, processor 102 can employ a deep learning-based image correction algorithm. First, face detection is performed on the initial input image sequence, using a pre-trained face detection model (such as a convolutional neural network-based face detection algorithm) to locate the position and bounding box of the face in the image. Then, based on the face's position information, the translation and rotation parameters required to move the face image to the center region of the image are calculated. For example, if the detected occupant's face image is located in the upper left corner of the image, the number of pixels needed for right and down translation, as well as the possible rotation angle (if the face is tilted), are determined by calculating the coordinate difference between the face image and the image center.

[0024] In addition, image interpolation algorithms can be used to fill and optimize pixels in translated and rotated images to ensure that image quality is not affected. For example, bilinear interpolation and bicubic interpolation can be used. Bilinear interpolation calculates the value of a new pixel by linearly weighting it to neighboring pixels, which has a relatively small computational load and is suitable for real-time processing scenarios. Bicubic interpolation, for example, can use information from the surrounding 16 pixels to perform more complex weighted calculations, resulting in higher quality image interpolation results. The appropriate interpolation algorithm can be selected based on the system's real-time requirements and hardware performance, or a combination of both can be used to maximize image quality while ensuring real-time performance.

[0025] In the example, processor 102 can also perform image enhancement processing on the corrected image sequence. This includes adjusting the brightness, contrast, and color saturation of the image to improve its visual appeal. For example, a histogram equalization algorithm is used to enhance the image's contrast, making facial features more clearly discernible. Regarding color saturation, it can be appropriately adjusted based on the overall tone of the image and the in-vehicle environment to avoid colors being too vibrant or too dull, ensuring that the image presents a natural and comfortable visual effect on the display.

[0026] In addition, to reduce latency in image transmission and display, the processor 102 can compress the image sequence to reduce data transmission bandwidth requirements and storage needs, thereby improving the overall response speed and real-time performance of the system in vehicle driving scenarios.

[0027] The display 103 can be connected to the processor 102 via wired or wireless means, and can be configured to display a corrected sequence of images, including facial images of rear passengers, to occupants in the front seat area of ​​the vehicle cabin, or a corrected sequence of front images, including facial images of front passengers, to occupants in the rear seat area. This ensures that the driver in the front seat and the infant in the rear seat can clearly see each other on the display. This technology allows the driver to monitor the infant's condition in the rear seat in real time, even while sitting in the front, effectively reducing anxiety and distraction during driving and improving driving safety. The infant in the rear seat can also see the driver's face in real time on the display, thus receiving effective reassurance. Furthermore, centering the images enhances realism, especially for the infant's visual experience, and strengthens the interaction between the two parties.

[0028] According to some embodiments, the initial image sequence includes a first image of a first occupant in the front row of the vehicle cabin and a second image of a second occupant in the rear row of the vehicle cabin. The processor 102 may be further configured to generate a corrected first image sequence and a corrected second image sequence based on the first and second images in the initial image sequence, respectively, wherein the facial image of the first occupant is located in the central region of each image in the corrected first image sequence, and the facial image of the second occupant is located in the central region of each image in the corrected second image sequence. The display 103 may include a front-row display, which is communicatively connected to the processor 102 and configured to display the corrected second image sequence to the first occupant in the front row of the vehicle cabin. The display 103 may also include a rear-row display, which is communicatively connected to the processor 102 and configured to display the corrected first image sequence to the second occupant in the rear row of the vehicle cabin.

[0029] In the example, the image capture device 101 can be positioned in the middle of the vehicle cabin and capture images of the front and rear of the vehicle cabin.

[0030] In the example, the front display can be mounted on the center console close to the driver's line of sight. The tilt angle of the front display is adjustable to allow the driver to view images of rear passengers without obstructing their view. The front display may be treated with anti-glare technology, using an anti-glare coating on the screen surface or a special optical design to reduce the impact of external light (such as direct sunlight or interior light reflections) on the display, ensuring clear display of the calibrated second image sequence under various lighting conditions.

[0031] In the example, the rear-seat display can be installed on the back of the front seat headrest (i.e., the rear side), featuring a rotatable and angle-adjustable design to allow passengers in different rear positions (such as the left, right, or middle seats) to easily view the calibrated first image sequence. The rear-seat display may also have anti-glare capabilities. Furthermore, to protect the eyesight of rear passengers, especially infants and young children, the rear-seat display can be equipped with a blue light filter function. This reduces the proportion of blue light emitted by the screen, minimizing eye strain from prolonged viewing. For example, hardware-level blue light filtering chips or software algorithms can be used to reduce blue light intensity to a safe range without significantly affecting image color and clarity.

[0032] According to some embodiments, the processor 102 may be further configured to: perform facial recognition on the initial image sequence; and in response to recognizing that the facial image of the second occupant in any frame of the initial image sequence has infant characteristics, control the rear display to start displaying the corrected first image sequence, and control the front display to start displaying the corrected second image sequence.

[0033] Therefore, when an infant or toddler is detected in the back seat, images of the front and rear seats can be automatically transmitted, improving both ease of use and driving safety.

[0034] In the example, processor 102 may include a face recognition module, which may employ a convolutional neural network model. This model can be pre-trained on a dataset containing a large number of facial images of different ages, genders, and ethnicities to learn a general facial feature representation. Furthermore, during training, image data labeled with face location, identity information, and age ranges (including categories such as infants and adults) can be used. The model's weight parameters are continuously optimized through a backpropagation algorithm, enabling the model to accurately identify faces in images and determine their age category.

[0035] When performing facial recognition on an initial image sequence, the model, in addition to locating the face, can also focus on facial information related to infant characteristics. For example, infants typically have relatively large head proportions, rounder facial contours, smaller eye spacing, and fuller cheeks. The model determines whether these features are characteristic of infants through quantitative analysis of these features. Methods based on feature point distance ratios can be used, such as calculating the ratio of eye spacing to face width, or the ratio of the area of ​​the mouth and nose region to the face area, and comparing these ratios with pre-set infant characteristic thresholds. If these ratios are within the infant characteristic threshold range, the facial image is determined to have infant characteristics. Simultaneously, skin texture analysis can be incorporated. Infants' skin is relatively smooth with less and shallower texture. By extracting texture features from the facial skin region of the image (e.g., using texture analysis methods such as gray-level co-occurrence matrix) and comparing them with adult skin texture models, the model can assist in determining whether the image belongs to an infant.

[0036] Once the processor 102 recognizes that the facial image of the second occupant in any frame of the initial image sequence has infant characteristics, it can immediately send a display start signal to the rear and front displays, while simultaneously transmitting relevant data information of the corrected first and second image sequences.

[0037] In some scenarios, infants and young children may undergo significant changes in position or posture when distressed. According to some embodiments, the processor 102 may be further configured to: perform image recognition on an initial image sequence; and, in response to recognizing that the magnitude of the position change or posture change of a second occupant in any two consecutive frames of the initial image sequence is greater than a magnitude threshold, control the rear display to begin displaying a corrected first image sequence, and control the front display to begin displaying a corrected second image sequence.

[0038] Therefore, when distressed behavior of infants and toddlers in the back seat is detected, images of both the front and back seats can be automatically transmitted, improving both ease of use and driving safety.

[0039] In the example, when performing image recognition on the initial image sequence, the processor 102 can first employ a feature extraction algorithm to obtain key information from the images. Human keypoint detection methods can be used to detect the position coordinates of major joints of the human body, such as the head, shoulders, elbows, wrists, hips, knees, and ankles. For each frame of the second image, information from these keypoints is extracted to construct a feature vector of the human posture. For two consecutive frames of the second image, the processor 102 can compare the coordinate changes of the infant's keypoints between the two frames. The displacement of each keypoint in the horizontal and vertical directions is calculated, and then the displacements of all keypoints are comprehensively analyzed to calculate the overall positional change of the infant.

[0040] When the position change or posture change calculated by the processor 102 exceeds the amplitude threshold, a display activation signal can be immediately sent to the rear and front displays, while simultaneously transmitting relevant data information of the corrected first and second image sequences. In this example, the amplitude threshold can be determined through extensive experimental data. For instance, by simulating various normal and abnormal (crying) movements of infants and toddlers in different scenarios, statistically analyzing their position and posture change ranges, and selecting a suitable upper limit as the threshold, the display function can be triggered promptly when the second passenger undergoes a significant change in movement, such as when a child in the rear seat suddenly changes from a quiet sitting posture to standing or twists their body dramatically.

[0041] According to some embodiments, the processor 102 may be further configured to: in response to receiving a signal indicating that the vehicle speed is greater than a speed threshold or a signal indicating that the vehicle is in complex road conditions, control the front display to pause the display of the corrected second image sequence.

[0042] In this example, processor 102 can connect to the vehicle's speed sensors and road condition monitoring system via an internal vehicle communication bus (such as a CAN bus). The speed sensors can continuously send real-time vehicle speed data to processor 102. For example, the speed sensors can be devices that calculate vehicle speed based on wheel rotation speed or GPS positioning information, and they can send data packets containing the current vehicle speed value to processor 102 at a certain frequency.

[0043] In this example, a traffic monitoring system can be used to collect traffic information about the area surrounding the vehicle. For instance, sensors such as cameras and radar can be used to detect road conditions in front of, behind, and to the sides of the vehicle, including road type (e.g., highway, city road, rural road), traffic flow, the presence of construction zones, and curve curvature. When the traffic monitoring system detects complex traffic conditions, such as traffic congestion ahead, entering a curve, or a construction zone, it can send corresponding indication signals to the processor 102.

[0044] When the vehicle speed data received by the processor 102 exceeds a speed threshold, it can control the front display to pause the display of the corrected second image sequence. This is because at higher speeds, the driver needs to focus more on road conditions, reducing unnecessary visual distractions. This further enhances driving safety.

[0045] Furthermore, in the example, to avoid frequent display switching caused by frequent speed fluctuations around a speed threshold, a buffer interval can be set. For example, the display should only be paused when the vehicle speed exceeds 60 km / h for more than 5 seconds.

[0046] The processor 102 can analyze signals from the traffic monitoring system. If the signal indicates that the vehicle is in complex road conditions, such as traffic congestion in urban areas with frequent stops and starts; or entering a curve requiring the driver to concentrate on steering to maintain vehicle stability; or there is a construction area ahead, where the road narrows and poses potential dangers, the processor 102 can control the front display to pause the display of the corrected second image sequence. Different priorities can be set for different types of complex road conditions. For example, entering a curve or construction area has a higher priority, and the display is immediately paused upon receiving such a signal; while traffic congestion can be assessed based on the degree of congestion (such as congestion length, average vehicle speed, etc.) to determine whether to pause the display.

[0047] According to some embodiments, the image capture device (e.g., a camera) can be disposed on the inner top of the vehicle cabin, and the processor 102 can be further configured to: crop out a first image and a second image from an initial image sequence; and input the cropped first and second images, as well as the initial image sequence, into a trained model to obtain a corrected first image sequence and a corrected second image sequence output by the model. The frontal facial image of the first occupant is located in the central region of each image in the corrected first image sequence, and the frontal facial image of the second occupant is located in the central region of each image in the corrected second image sequence.

[0048] Placing the image capture device on the inner top of the vehicle cabin provides a wider field of view. Because of its top position, the image capture device can capture images of both the front and rear passenger areas more comprehensively, reducing blind spots caused by objects obstructing the view.

[0049] In the example, a preset image segmentation algorithm can be used to crop out the first and second images from the initial image sequence. For example, cropping can be performed according to pre-defined region rules, which can be based on the understanding of the vehicle cabin layout. For instance, before cropping, the processor 102 divides the initial image into different regions, determining the approximate range of the front and rear areas based on the position of the seats in the vehicle cabin and the known size ratio. Alternatively, semantic segmentation techniques using deep learning can be used to identify the front and rear areas of the vehicle cabin in the initial image, and then accurately crop out the first image belonging to the first occupant and the second image belonging to the second occupant. In the example, the trained model can be a graph-based model, such as a deep learning-based neural network model, which can be trained on a large-scale dataset. When the processor 102 inputs the cropped first and second images, as well as the initial image sequence, into the trained model, the model will automatically correct the frontal facial images of the first and second occupants to the central region of each image in the corresponding image sequence based on the learned features and correction rules. Therefore, although the image capture device located on the inner top of the cabin captures an image of the occupants from above, the corrected first and second image sequences can display natural frontal facial images of the first and second occupants in the central area that conform to the interior environment, thereby further enhancing the user experience. In the example, the model can also handle some special cases, such as when part of the face is occluded. By analyzing the features of the unoccluded part and the overall image information, and by utilizing the feature reasoning ability learned during training, it can correct the frontal facial image to the center area of ​​the image as completely as possible.

[0050] Figure 2 This is a block diagram illustrating a video calling device 200 for a vehicle according to another exemplary embodiment of the present disclosure. Figure 2 As shown, the video call device 200 for a vehicle includes an image capture device 201, a processor 202, a front display 203, and a rear display 204, which are similar to the image capture device 101, processor 102, front display, and rear display described above, respectively, and will not be described again here.

[0051] According to some embodiments, continue to refer to Figure 2The video call device 200 for the vehicle may further include an audio capture device 205 and a rear-seat audio player 206. The audio capture device 205 may be configured to capture first audio emitted by a first occupant in the front area of ​​the vehicle cabin. The processor 202 may be configured to: perform noise reduction processing on the first audio; and control the rear-seat audio player 206 to play the noise-reduced first audio while the rear-seat display 204 displays a corrected first image sequence.

[0052] In the example, the audio capture device 205 may include a microphone, for example, a microphone array with high sensitivity and directionality. The microphone array may be mounted, for example, in front of the vehicle's center console, close to the first occupant (such as the driver and front passenger), to accurately capture first audio from the front passenger area. For example, a circular microphone array may be used to collect sound information from the front passenger area from more directions, improving the ability to capture the first occupant's voice and reducing interference from ambient noise.

[0053] In the example, processor 202 can employ a signal processing-based noise reduction algorithm, such as an adaptive filtering algorithm. When processing the first audio signal, the audio signal is first converted to the frequency domain, and an adaptive filter is used to update and adjust it in real time based on the statistical characteristics of the noise. For example, a minimum mean square adaptive filter can be used. By continuously adjusting the filter coefficients, the mean square value of the output error signal (the desired signal minus the filter output signal) is minimized, thereby effectively suppressing noise in the first audio signal. Simultaneously, to improve the noise reduction effect, deep learning methods can be combined. A large amount of noisy and clean audio data is used to train the neural network, enabling it to learn the mapping relationship for recovering clean audio from noisy frequencies. In practical applications, the first audio signal can be input into the trained neural network to obtain the noise-reduced audio signal.

[0054] Furthermore, to ensure a good user experience, the processor 202 can ensure the synchronization of audio and video. For example, timestamp information can be added when processing and transmitting the corrected first image sequence and the noise-reduced first audio. The timestamp information ensures that the audio played by the rear audio player 206 and the video displayed on the rear monitor 204 are synchronized in time. For example, millisecond-accurate timestamps are added to each frame of the corrected first image sequence and each segment of the noise-reduced first audio. The playback order and timing alignment of the audio and video are controlled based on these timestamps, avoiding audio-video desynchronization. This allows infants and toddlers in the rear seats to receive more comprehensive and realistic comfort.

[0055] According to some embodiments, continue to refer to Figure 2The audio capture device 205 can be further configured to capture a second audio emitted by a second occupant in the rear seat area of ​​the vehicle cabin. The processor 202 can be further configured to: perform voiceprint recognition on the second audio; and, in response to recognizing that the second audio includes an infant's cry, control the rear-seat display 204 to begin displaying a corrected first image sequence, and control the front-seat display 203 to begin displaying a corrected second image sequence.

[0056] In the example, the audio capture device 205 may also include an audio capture device located in the rear passenger area. For example, an additional microphone unit may be installed in the center of the rear passenger area or near the rear seats to form a more comprehensive audio capture network, ensuring that secondary audio from a second occupant in the rear passenger area can be clearly captured.

[0057] Processor 202 may include a voiceprint recognition module, which can be built based on deep learning technology. For example, it can collect a large number of different infant crying samples, as well as other sound samples (such as normal conversations from back-seat passengers, ambient noise, etc.), to build a voiceprint database. These samples can include infant cries of different genders and ages, covering different features such as volume, pitch, duration, and rhythm, thereby ensuring that the voiceprint recognition system has broad adaptability. During training, different sound samples can be input into the network to extract acoustic features of the sound, such as Mel-frequency cepstral coefficients (MFCC), allowing the network to learn to distinguish the characteristic patterns of infant cries from other sounds.

[0058] During voiceprint recognition, the processor 202 first performs frame segmentation on the acquired second audio, that is, divides the continuous audio stream into multiple short frames to meet the input requirements of the voiceprint recognition algorithm. Then, it extracts acoustic features from each frame of audio and inputs them into the trained voiceprint recognition model for voiceprint recognition.

[0059] When the processor 202 recognizes that the second audio includes the sound of an infant crying, it can immediately control the front display 203 to start displaying the corrected second image sequence.

[0060] According to another aspect of this disclosure, a vehicle is provided that includes the aforementioned video call device 100 or 200 for the vehicle.

[0061] According to another aspect of this disclosure, a video call method for vehicles is provided. Figure 3 This is a flowchart illustrating a video call method 300 for a vehicle according to an exemplary embodiment of the present disclosure.

[0062] like Figure 3 As shown, the video calling method 300 for vehicles includes: Step 301: Capture the initial image sequence, which includes facial images of the occupants inside the vehicle cabin; Step 302: Generate a corrected image sequence based on the initial image sequence, wherein the occupant's facial image is located in the central region of each image in the corrected image sequence; and Step 303: Display the corrected image sequence to other occupants in the cabin.

[0063] It should be understood that Figure 3 The various steps of the video calling method 300 for vehicles shown can be combined with... Figure 1 The various devices or modules described correspond to those in the video calling device 100 for a vehicle. Therefore, the features and advantages described above for the video calling device 100 for a vehicle also apply to the video calling method 300 for a vehicle. For the sake of brevity, certain operations, features, and advantages will not be repeated here.

[0064] Furthermore, in some embodiments, the video call method 300 for vehicles can also be related to the above-mentioned... Figure 2 The various devices or modules in the video calling device 200 for vehicles described herein correspond to steps, features, and advantages.

[0065] According to another aspect of this disclosure, a computer device is provided, comprising: a memory, a processor, and a computer program stored on the memory, wherein the processor is configured to execute the computer program to implement the video calling method 300 for a vehicle according to this disclosure.

[0066] refer to Figure 4 Computer device 4000 may include elements connected to or communicating with bus 4002 (possibly via one or more interfaces). For example, computer device 4000 may include bus 4002, one or more processors 4004, one or more input devices 4006, and one or more output devices 4008. The one or more processors 4004 may be any type of processor and may include, but is not limited to, one or more general-purpose processors and / or one or more dedicated processors (e.g., special-purpose processing chips). Processor 4004 may process instructions that execute within computer device 4000, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interface). In other embodiments, multiple processors and / or multiple buses may be used with multiple memories, if desired. Similarly, multiple computer devices may be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 4 Take a processor 4004 as an example.

[0067] Input device 4006 can be any type of device capable of inputting information to computer device 4000. Input device 4006 can receive input numeric or character information, as well as generate key signal input related to user settings and / or function control of the computer device for event extraction, and can include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output device 4008 can be any type of device capable of presenting information, and can include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer.

[0068] Computer device 4000 may also include or be connected to a non-transitory storage device 4010. The non-transitory storage device may be any storage device that is non-transitory and capable of storing data, and may include, but is not limited to, disk drives, optical storage devices, solid-state storage, floppy disks, flexible disks, hard disks, magnetic tapes or any other magnetic media, optical discs or any other optical media, ROM (read-only memory), RAM (random access memory), cache memory and / or any other memory chip or cartridge, and / or any other medium from which a computer can read data, instructions, and / or code. The non-transitory storage device 4010 may be detachable from an interface. The non-transitory storage device 4010 may have data / programs (including instructions) / code / modules for implementing the methods and steps described above.

[0069] Computer device 4000 may also include communication device 4012. Communication device 4012 may be any type of device or system enabling communication with external devices and / or with a network, and may include, but is not limited to, modems, network interface cards, infrared communication devices, wireless communication devices, and / or chipsets such as Bluetooth. TM Devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.

[0070] Computer device 4000 may also include working memory 4014, which may be any type of working memory that can store programs (including instructions) and / or data useful to the operation of processor 4004, and may include, but is not limited to, random access memory and / or read-only memory devices.

[0071] The software elements (programs) may reside in the working memory 4014, including but not limited to the operating system 4016, one or more application programs 4018, drivers, and / or other data and code. Instructions for performing the methods and steps described above may be included in one or more application programs 4018, and the methods described above can be implemented by the processor 4004 reading and executing the instructions of one or more application programs 4018. The executable code or source code of the instructions for the software elements (programs) may also be downloaded from a remote location.

[0072] It should also be understood that various modifications can be made depending on specific requirements. For example, custom hardware can also be used, and / or specific elements can be implemented using hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. For example, some or all of the disclosed methods and apparatus can be implemented by programming hardware (e.g., programmable logic circuits including field-programmable gate arrays (FPGAs) and / or programmable logic arrays (PLAs)) using logic and algorithms according to this disclosure in assembly language or hardware programming languages ​​(such as Verilog, VHDL, C++).

[0073] It should also be understood that the aforementioned methods can be implemented using a server-client model. For example, a client can receive user input data and send it to the server. Alternatively, a client can receive user input data, perform some of the processing described in the aforementioned methods, and send the resulting data to the server. The server can receive data from the client, execute the aforementioned methods or another part thereof, and return the execution result to the client. The client can receive the execution result from the server and, for example, present it to the user via an output device. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is created by computer programs running on corresponding computer devices that have a client-server relationship with each other. The server can be a server in a distributed system or a server incorporating blockchain technology. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product within the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0074] It should also be understood that the components of computer device 4000 can be distributed across a network. For example, some processing can be performed using one processor, while other processing can be performed simultaneously by another processor located far away from that processor. Other components of computer device 4000 can also be distributed similarly. In this way, computer device 4000 can be interpreted as a distributed computing system that performs processing in multiple locations.

[0075] Some embodiments of this disclosure also provide a computer-readable storage medium storing instructions that, when executed individually or jointly by one or more processors of a computer device, cause the computer device to perform the video calling method 300 for a vehicle described above.

[0076] Computer-readable media can be tangible media that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device. Machine-readable media can be machine-readable signal media or machine-readable storage media. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0077] Some embodiments of this disclosure also provide a computer program product including instructions that, when executed individually or jointly by one or more processors of a computer device, cause the computer device to perform the video calling method 300 for a vehicle described above.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and not to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure, and they should all be covered within the scope of the claims and specification of this disclosure. In particular, as long as there is no structural conflict, the various technical features mentioned in the various embodiments can be combined in any way. This disclosure is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A video call device for a vehicle, comprising: An image capture device is configured to capture an initial image sequence, the initial image sequence including facial images of occupants inside the vehicle cabin; The processor is communicatively connected to the image capture device and configured to generate a corrected image sequence based on the initial image sequence, wherein the occupant's facial image is located in the central region of each image in the corrected image sequence; as well as The display is communicatively connected to the processor and configured to display the corrected image sequence to other occupants within the cabin.

2. The video call device according to claim 1, wherein, The initial image sequence includes a first image of the first occupant in the front row of the vehicle cabin and a second image of the second occupant in the rear row of the vehicle cabin. The process of generating a corrected image sequence based on the initial image sequence includes: generating a corrected first image sequence and a corrected second image sequence based on a first image and a second image in the initial image sequence, respectively, wherein the facial image of the first occupant is located in the central region of each image in the corrected first image sequence, and the facial image of the second occupant is located in the central region of each image in the corrected second image sequence. And wherein the display includes: A front-row display is configured to show the corrected second image sequence to a first occupant in the front area of ​​the vehicle cabin; and The rear-seat display is configured to show the corrected first image sequence to a second occupant in the rear area of ​​the vehicle cabin.

3. The video call apparatus for a vehicle according to claim 2, wherein The processor is further configured to: Perform facial recognition on the initial image sequence; and In response to the recognition that the facial image of the second occupant in any frame of the initial image sequence has infant characteristics, the rear-row display is controlled to start displaying the corrected first image sequence, and the front-row display is controlled to start displaying the corrected second image sequence.

4. The video call apparatus for a vehicle according to claim 2, wherein The processor is further configured to: Perform image recognition on the initial image sequence; and In response to the detection that the position change or posture change of the second occupant in any two consecutive frames of the initial image sequence is greater than an amplitude threshold, the rear display is controlled to start displaying the corrected first image sequence, and the front display is controlled to start displaying the corrected second image sequence.

5. The video call apparatus for a vehicle according to any one of claims 2 to 4, wherein The processor is further configured to: In response to receiving a signal indicating that the vehicle's speed is greater than a speed threshold or a signal indicating that the vehicle is in complex road conditions, the front display is controlled to pause the display of the corrected second image sequence.

6. The video call apparatus for a vehicle according to any one of claims 2 to 4, wherein The image capture device is located on the inner top of the vehicle cabin. Furthermore, generating a corrected first image sequence and a corrected second image sequence based on the first and second images in the initial image sequence respectively includes: The first image and the second image are cropped from the initial image sequence, respectively; and The cropped first and second images, along with the initial image sequence, are input into a trained model to obtain a corrected first image sequence and a corrected second image sequence output by the model. The frontal facial image of the first occupant is located in the central region of each image in the corrected first image sequence, and the frontal facial image of the second occupant is located in the central region of each image in the corrected second image sequence.

7. The video call device for a vehicle according to any one of claims 2 to 4, further comprising an audio capture device and a rear-seat audio player, wherein, The audio capture device is configured to capture a first audio signal emitted by a first occupant in the front row of the vehicle cabin. Furthermore, the processor is configured as follows: Noise reduction processing is performed on the first audio; and While the rear display shows the corrected first image sequence, the rear audio player is controlled to play the noise-reduced first audio.

8. The video call device for a vehicle according to claim 7, wherein The audio capture device is further configured to capture a second audio signal emitted by a second occupant in the rear passenger area of ​​the vehicle cabin. Furthermore, the processor is configured as follows: Voiceprint recognition is performed on the second audio; and In response to the recognition that the second audio includes an infant crying sound, the rear display is controlled to start displaying the corrected first image sequence, and the front display is controlled to start displaying the corrected second image sequence.

9. A vehicle comprising a video call device for a vehicle according to any one of claims 1-8.

10. A video call method for a vehicle, comprising: Capture an initial image sequence, which includes facial images of occupants inside the vehicle cabin; A corrected image sequence is generated based on the initial image sequence, wherein the occupant's facial image is located in the central region of each image in the corrected image sequence; as well as The corrected image sequence is displayed to other occupants inside the vehicle.

11. A computer device, comprising: Memory, processor, and computer program stored on said memory, The processor is configured to execute the computer program to implement the control method of claim 10.

12. A non-transitory computer readable storage medium having stored thereon a computer program, wherein, When the computer program is executed by the processor, it implements the video call method for a vehicle as described in claim 10.

13. A computer program product comprising a computer program, wherein, When the computer program is executed by the processor, it implements the video call method for a vehicle as described in claim 10.