Control method and device of electronic rearview mirror system and electronic rearview mirror system

By introducing camera and gesture recognition technology into the electronic rearview mirror system, the driver's gesture movements and automatically adjusts the display content, the safety hazards caused by drivers' manual adjustment of the display screen in the prior art are solved, and intelligent interaction between the driver and the CMS monitor and a better driving experience are achieved.

CN120096454APending Publication Date: 2025-06-06BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237808.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing electronic rearview mirror system requires the driver to manually adjust the display screen during driving, which causes the driver to distract himself and poses safety risks. When the screen size and viewing angle are not in the optimal position, it affects the driver's judgment of road conditions.

Method used

By introducing a first camera and a DMS camera in the CMS system, the driver's gesture movement is monitored, and the collected user video stream is analyzed and processed through the CMS controller, the gestures are identified and corresponding operation instructions are executed, and automatic control of the content displayed in CMS is realized.

Benefits of technology

It realizes intelligent interaction between the driver and the CMS monitor, and adjusts the monitor and field of view through gestures, improves the driver's ability to judge road conditions, reduces the safety risks of manual operation, and improves the driving experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120096454A_ABST
    Figure CN120096454A_ABST
Patent Text Reader

Abstract

The invention discloses a control method and device of an electronic rearview mirror system and the electronic rearview mirror system. The method comprises the steps that a CMS controller obtains a user video stream collected by a first camera; analyzing and processing the user video stream, and determining whether a gesture image exists in the user video stream; if yes, triggering gesture classification on the gesture image in the user video stream to obtain a recognition gesture; determining an operation instruction corresponding to the recognition gesture based on a preset mapping relationship between the recognition gesture and the operation instruction; the operation instruction is executed to control CMS display content, the display content is obtained based on a view field image collected by a second camera, and the second camera is arranged in a rearview mirror outside the cabin. According to the scheme, the gesture action of the driver is monitored through the DMS, the gesture video stream is transmitted to the CMS monitor in real time, intelligent interaction between display of the CMS monitor and the driver is achieved, and the monitor can be rapidly and conveniently adjusted by adjusting the display of the visible field of view of the monitor through gestures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of automotive electronic technology, and more specifically, to a control method and device for an electronic rearview mirror system and an electronic rearview mirror system. Background Art

[0002] The electronic rearview mirror is also called camera-monitor system, or CMS, which includes a camera module and a monitor. Its working process is that the camera collects the original RAW (an image format) image, which is processed by the ISP image processing module to generate a YUV image, which is converted into a digital signal and transmitted to the system controller. The controller is deployed in the monitor and communicates with other control units through CAN. After processing the video signal, it is displayed in real time on the CMS display (i.e., monitor). It has the advantages of being waterproof, windproof, and anti-reflective, while reducing wind resistance and blind spots in the field of vision, and has a better monitoring effect than optical rearview mirrors and human eyes.

[0003] At the same time, CMS can also improve the intelligent driving experience. During driving, the image outside the car becomes an indispensable source of information. CMS can not only meet the functions of an optical rearview mirror, but also expand the viewing angle and reduce the blind spot of turning. Combined with AI technology, it can increase safety warning functions and intelligent interaction, etc., expanding the operability and utilization space of CMS.

[0004] The CMS monitor is installed near the left and right doors in the cabin. In order to correspond to the principle of optical imaging and ensure the driver's sensory experience, the monitor generally crops the image collected by the camera to a suitable display size in advance for display, resulting in a loss of field of view. The driver can also manually adjust the appropriate screen ratio and field of view angle, but during driving, the driver's manual adjustment of the display screen according to the environment will distract attention and pose a driving hazard. If the screen size and viewing angle are not in the optimal position, it will greatly affect the driver's judgment of the road conditions and cause safety hazards. Summary of the invention

[0005] In view of this, this application provides the following technical solutions:

[0006] The first aspect of the present application provides a control method for a CMS system, comprising:

[0007] The CMS controller obtains a user video stream captured by a first camera, where the first camera is a camera capable of capturing an image of the driver's upper body area;

[0008] Analyze and process the user video stream to determine whether there is a gesture image therein;

[0009] If yes, triggering gesture classification of the gesture image in the user video stream to obtain a recognized gesture;

[0010] Based on a preset mapping relationship between the recognition gesture and the operation instruction, determining the operation instruction corresponding to the recognition gesture;

[0011] The operation instruction is executed to realize the control of the CMS display content, wherein the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.

[0012] In a possible implementation, it also includes:

[0013] obtaining a gear position signal and / or a steering signal of the vehicle;

[0014] Based on the gear signal and / or the turn signal, the control is switched to a field of view mode corresponding to the gear signal and / or the turn signal according to a preset logic control, and different field of view modes correspond to different display contents.

[0015] In a possible implementation, based on the gear signal and / or the turn signal, switching to a field of view mode corresponding to the gear signal and / or the turn signal according to a preset logic control includes:

[0016] If the gear position signal indicates that the vehicle enters the reverse gear, the vehicle is controlled to enter the reverse field of view, and in the reverse field of view, the display content of the CMS display corresponds to the image content of the rear wheel part of the vehicle in the field of view image;

[0017] If the turn signal indicates that the vehicle is about to turn, control the vehicle to enter the turning field of view, and in the turning field of view, the display content of the CMS display corresponds to the image content of the side part of the vehicle in the field of view image;

[0018] If the gear position signal indicates that the reverse gear has been exited or the turn signal no longer exists, the control enters the normal field of view.

[0019] In a possible implementation, it also includes:

[0020] Pre-configure custom field of view;

[0021] If the operation instruction corresponds to a switch custom field of view instruction, the control enters the custom field of view.

[0022] In a possible implementation, before executing the operation instruction to control the CMS display content, the method further includes:

[0023] Determining whether the operation instruction conflicts with the current vehicle speed or current state;

[0024] If there is a conflict, the operation is prohibited from being performed.

[0025] In a possible implementation, executing the operation instruction to control the CMS display content includes at least one of the following:

[0026] Execute the operation instruction to adjust the display brightness of the CMS display;

[0027] Execute the operation instruction to adjust the field size of the CMS display content;

[0028] The operation instruction is executed to adjust the viewing angle of the CMS display content.

[0029] In a possible implementation, analyzing and processing the user video stream to determine whether there is a gesture image therein includes at least one of the following:

[0030] Performing frame processing on the user video stream to obtain a frame image data set, and filtering out gesture images containing gesture actions from the frame image data set;

[0031] The user video stream is detected by sliding a detection window to identify gesture images in the user video stream.

[0032] In a possible implementation, triggering gesture classification of the gesture image in the user video stream to obtain a recognized gesture includes:

[0033] The classifier model is used to classify gestures using gesture images in the user video stream; the gesture classification network of the classifier model consists of a data input layer, a convolution calculation layer, a ReLU excitation layer, a pooling layer, and a fully connected layer.

[0034] A second aspect of the present application provides a control device for a CMS system, comprising:

[0035] A video acquisition module, used to obtain a user video stream captured by a first camera, where the first camera is a camera capable of capturing an image of the upper body area of ​​the user;

[0036] A gesture recognition module, used to analyze and process the user video stream to determine whether there is a gesture image therein;

[0037] A gesture classification module, used for classifying gestures on the gesture images in the user video stream to obtain recognized gestures when the recognition result of the gesture recognition module is yes;

[0038] An instruction determination module, used to determine the operation instruction corresponding to the recognition gesture based on a preset mapping relationship between the recognition gesture and the operation instruction;

[0039] The instruction execution module is used to execute the operation instruction to realize the control of the CMS display content, wherein the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.

[0040] A third aspect of the present application provides a CMS system, including:

[0041] The CMS camera is installed in the rearview mirror outside the vehicle cabin and is used to collect and obtain field of view images;

[0042] The DMS camera is installed on the steering column of the steering wheel in the vehicle cabin and is used to collect images of the driver's upper body area;

[0043] A CMS monitor comprises at least one controller and two displays, wherein the two displays are respectively arranged on a dashboard surface near a main driver's door and a co-driver's door in a vehicle cabin, wherein the controller is used to obtain a user video stream captured by a first camera, wherein the first camera is a camera capable of capturing images of the driver's upper body area; the user video stream is analyzed and processed to determine whether there are gesture images therein; if so, gesture classification of the gesture images in the user video stream is triggered to obtain a recognized gesture; based on a preset mapping relationship between recognized gestures and operation instructions, an operation instruction corresponding to the recognized gesture is determined; and the operation instruction is executed to control the CMS display content, wherein the display content is obtained based on a field of view image captured by a second camera, wherein the second camera is arranged in a rearview mirror outside the cabin.

[0044] It can be known from the above technical scheme that the embodiment of the present application discloses a control method, device and electronic rearview mirror system of an electronic rearview mirror system, the method comprising: a CMS controller obtains a user video stream captured by a first camera, the first camera being a camera capable of capturing images of the upper body area of ​​the driver; analyzing and processing the user video stream to determine whether there is a gesture image therein; if so, triggering gesture classification of the gesture image in the user video stream to obtain a recognized gesture; determining the operation instruction corresponding to the recognized gesture based on a preset mapping relationship between the recognized gesture and the operation instruction; executing the operation instruction to realize control of the CMS display content, the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin. The above scheme monitors the driver's gesture movements through the DMS, transmits the gesture video stream to the CMS monitor in real time, realizes intelligent interaction between the CMS monitor display and the driver, and adjusts the monitor and field of view display through gestures, so that the monitor can be adjusted quickly and conveniently. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1 A flow chart of a control method of a CMS system disclosed in an embodiment of the present application;

[0047] Figure 2 This is a schematic diagram of the actual vehicle layout position of the CMS system disclosed in the embodiment of the present application;

[0048] Figure 3 A logic block diagram of a CMS interactive display method disclosed in an embodiment of the present application;

[0049] Figure 4 A schematic diagram of the interactive process of adjusting the field of view of a CMS monitor disclosed in an embodiment of the present application;

[0050] Figure 5 This is a flow chart of the implementation of the control method of the CMS system disclosed in the embodiment of the present application;

[0051] Figure 6 This is a network structure diagram of the gesture classification module of the CMS system disclosed in the embodiment of this application:

[0052] Figure 7 A schematic diagram of the structure of a control device of a CMS system disclosed in an embodiment of the present application;

[0053] Figure 8 A schematic diagram of the structure of a CMS system disclosed in an embodiment of the present application;

[0054] Fig. 9 A schematic diagram of the specific composition structure of the CMS camera and CMS monitor disclosed in the embodiment of the present application;

[0055] Fig.10 This is a schematic diagram of the structure of the DMS camera disclosed in the embodiment of the present application. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0057] Figure 1 The present invention discloses a control method flow chart of a CMS system. The CMS system is an electronic rearview mirror system installed in a vehicle, including a CMS camera, a CMS monitor and a DMS (Driver Monitor System) camera. Figure 2 This is a schematic diagram of the actual vehicle layout of the CMS system disclosed in the embodiment of this application. Figure 1 and Figure 2 As shown, the control method of the CMS system may include:

[0058] Step 101: The CMS controller obtains a user video stream captured by a first camera, where the first camera is a camera capable of capturing images of the upper body area of ​​the driver.

[0059] The CMS controller is the main control module in the aforementioned CMS monitor, which is connected to the CMS camera and the DMS camera respectively, and can obtain images / videos collected by the two cameras. In this embodiment, the first camera is the DMS camera.

[0060] In the implementation, when the vehicle is started, the first camera collects the driver's video in real time and sends the collected user video stream to the CMS controller. When the driver makes a gesture, the corresponding image / video is collected by the first camera and sent to the CMS controller, which performs corresponding recognition processing.

[0061] Step 102: Analyze and process the user video stream to determine whether there is a gesture image therein.

[0062] After receiving the user video stream, the CMS controller will perform real-time detection to detect whether there is a gesture image in the user video stream, that is, an image containing the driver's hand. The first camera can be set on the steering column of the steering wheel, facing the driver, so that the driver's gesture can be recognized.

[0063] Step 103: If yes, trigger gesture classification of the gesture image in the user video stream to obtain the recognized gesture.

[0064] If a gesture image is included, it indicates that the user may be inputting a gesture and wants to control the display content of the CMS display through gestures. Then, a single trigger is used to classify the gesture image to obtain a recognized gesture. At this time, it is necessary to identify and classify the gesture image containing the driver's hand to determine which gesture the user inputs.

[0065] Step 104: Based on a preset mapping relationship between the recognition gesture and the operation instruction, determine the operation instruction corresponding to the recognition gesture.

[0066] Different gestures correspond to different operating instructions, and the mapping relationship between the recognition gesture and the operating instruction is preset in the CMS system. Therefore, after the CMS controller obtains the recognition gesture, it can determine the operating instruction corresponding to the recognition gesture based on the preset mapping relationship between the recognition gesture and the operating instruction. The operating instructions may include but are not limited to: zooming in on the screen, reducing the screen, increasing the display brightness, decreasing the display brightness, adjusting the viewing angle, switching the field of view mode, etc. The field of view mode of the CMS system may include but is not limited to conventional field of view, reversing field of view, steering field of view, and custom field of view. The display content and characteristics of each field of view mode will be described in detail in the embodiments below, and will not be introduced in detail here.

[0067] Step 105: Execute the operation instruction to control the CMS display content, where the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.

[0068] After the operation instruction corresponding to the recognized gesture is determined, the operation instruction can be executed to realize the control of the content displayed on the CMS display.

[0069] The control method of the CMS system described in this embodiment monitors the driver's gestures through the driver video captured by the first camera. When the user inputs a set gesture, it can trigger the CMS controller to execute the corresponding operation instruction to realize the interaction between the driver and the CMS monitor. This solution does not require the user to manually operate and control based on watching the display. Instead, the CMS monitor display content can be controlled by making predetermined gestures. Compared with the traditional manual operation on the display or physical buttons, it is more convenient and has higher driving safety.

[0070] In order to facilitate a better understanding of the contents of the embodiments of the present application, the implementation logic of the solution of the present application is first introduced. Figure 3 This is a logic block diagram of a CMS interactive display method disclosed in an embodiment of the present application. Figure 3 As shown in the figure, the control scheme of the CMS system is mainly based on the image acquisition module and the main control module, wherein the image acquisition module includes the DMS camera (first camera) and the CMS camera (second camera). The DMS is arranged on the column and can recognize the driver's gestures, and save the video data to the gesture detection module in the main control module for gesture detection and recognition; the CMS camera is used to collect images for display on the CMS display. The main control module includes a gesture detection module and a gesture classification module.

[0071] Determine whether the execution subject of the gesture image corresponds to the gesture detection module. Figure 3As shown, the gesture detection module includes two implementations of gesture detection. One is to perform frame processing on the video data collected by the DMS camera, and then perform static gesture recognition from the framed images. Since the video data and images collected by the DMS camera contain gesture parts and other parts, in order to avoid the waste of computing power caused by detecting data that does not contain gesture actions, the video data input by the DMS camera is framed and processed to obtain a data set, and the static gesture images containing gesture actions are screened out as the input of the gesture classification module, and the gesture classification algorithm of the gesture classification module is activated for further classification. Another implementation of gesture detection is to complete it through a dynamic gesture detector, and the video data collected by the DMS camera is passed to the dynamic gesture detector, and the video stream is detected through a sliding detection window. When a dynamic gesture is recognized, this video data is output to the gesture classification module, which is consistent with the input of the static gesture image, and then the gesture classification algorithm in the gesture classification module is activated once.

[0072] Therefore, the analyzing and processing of the user video stream to determine whether there are gesture images therein may include at least one of the following: performing frame processing on the user video stream to obtain a frame image data set, and filtering out gesture images containing gesture actions from the frame image data set; detecting the user video stream through a sliding detection window to identify gesture images in the user video stream.

[0073] In this application, the gesture classification algorithm can be implemented based on the CNN architecture (C3D and ResNext) classifier model. The gesture classification network of the classifier model consists of a data input layer, a convolution calculation layer, a ReLU excitation layer, a pooling layer, and a fully connected layer, and finally outputs the gesture recognition result, that is, the recognized gesture. Match gestures with operation instructions to adjust the brightness of the CMS monitor and the field of view of the picture, and finally display the field of view mode according to the usage scenario, including the normal field of view, the turning field of view, the reversing field of view, the custom field of view, etc., which can realize the real-time display of the on-board control instructions and the driving status and adjust the field of view of the CMS monitor, and realize the intelligent interactive display between the CMS monitor and the user.

[0074] That is, the execution of the operation instruction to control the CMS display content may include at least one of the following: executing the operation instruction to adjust the display brightness of the CMS display; executing the operation instruction to adjust the field of view size of the CMS display content; executing the operation instruction to adjust the field of view angle of the CMS display content. Through the control method of the CMS system disclosed in the embodiment of the present application, the display content of the CMS monitor can be controlled, and the control here includes but is not limited to switching of scene modes, display brightness adjustment, display screen ratio adjustment, etc.

[0075] Specifically, Figure 4 The figure shows the interactive process diagram of the CMS monitor field of view adjustment. Figure 4 As shown, the CMS monitor can switch different CMS field of view modes according to the VCU (Vehicle Control Unit) signal of the driving status. During normal driving, the CMS monitor is in the normal field of view mode. The normal field of view is displayed on the monitor after being cropped at the original resolution according to the preset value. The cropping size meets the regulatory requirements and the magnification meets the regulatory requirements. The vehicle body does not block the driver's field of view, which meets the driver's real-life perception needs and avoids the distance illusion caused by excessive changes in magnification. At the same time, the field of view adjustment function can also be activated in the normal field of view mode. By inputting gesture commands to the DMS camera, the field of view can be adjusted up and down, left and right, and zoomed in and out, thereby realizing the field of view adjustment. When the main control module receives the VCU turning signal, the CMS monitor switches from the normal field of view mode to the turning field of view mode, crops the field of view according to the turning preset value, enlarges the cropped image according to a certain magnification factor, and displays it on the CMS monitor, enlarging the turning blind spot to provide a more accurate temporary field of view. The field of view adjustment function can also be started in this mode, and the field of view can be adjusted up and down, left and right, and zoomed in and out through gestures to achieve zooming in on road details and adjusting the image of the turning area. After the adjustment, the field of view adjustment function can be exited through gestures.

[0076] When the CMS monitor receives the VCU R gear and turn signal signal, the control switches from the normal field of view mode to the reversing field of view mode, and automatically switches the field of view according to the VCU signal of the driving status. The field of view adjustment function can be switched in and out through gesture control (in the implementation, a gesture is made to enable the gesture adjustment function first, and then other gestures are made to control the field of view, to avoid the field of view jump and scene mismatch that may be caused by direct gesture control of the field of view, affecting safe driving); the field of view is cropped according to the preset value, mainly displaying the rear wheels and road conditions. In the implementation, two pictures can be displayed on the CMS monitor at the same time, such as the reversing field of view and the overall field of view, and the reversing field of view can be located below the overall field of view; the turning field of view mode can also be called at the same time and displayed on the monitor at the same time, so as to facilitate the observation of more detailed road conditions; the field of view adjustment function can also be started in this mode, and the field of view can be adjusted up and down, left and right, and zoomed in and out through gestures to complete the adjustment of the field of view so as to observe more detailed information. When the driver exits the R gear or the turn signal is corrected, the reversing field of view mode is exited and the normal field of view is returned.

[0077] Therefore, the control method of the CMS system may also include: obtaining the gear signal and / or turn signal of the vehicle; based on the gear signal and / or turn signal, switching to the field of view mode corresponding to the gear signal and / or turn signal according to preset logic control, different field of view modes correspond to different display contents.

[0078] Specifically, based on the gear signal and / or the turn signal, according to the preset logic control, switching to the field of view mode corresponding to the gear signal and / or the turn signal may include: if the gear signal indicates that the vehicle has entered the reverse gear, control entering the reverse field of view, in which the display content of the CMS display corresponds to the image content of the rear wheel part of the vehicle in the field of view image; if the turn signal indicates that the vehicle is about to turn, control entering the turning field of view, in which the display content of the CMS display corresponds to the image content of the side part of the vehicle in the field of view image. If the gear signal indicates that the reverse gear has been exited or the turn signal no longer exists, control entering the normal field of view.

[0079] In other implementations, the method may further include: pre-configuring a custom field of view; and if the operation instruction corresponds to an instruction to switch the custom field of view, controlling entry into the custom field of view.

[0080] For example, when parking, you can customize the field of view. The maximum field of view is displayed first, and the user can adjust the field of view through gestures to the field of view that meets the user's wishes, including the range and size of the field of view display area. Furthermore, on this basis, the sizes of the preset cropping windows of the other three modes can be set according to the driver's needs.

[0081] Manually adjusting the field of view includes the functions of adjusting the field of view upward, downward, leftward, rightward, and zooming in and out. Adjusting upward is to adjust the step length by translating upward through the field of view cropping window, corresponding to the gesture action of one finger on one hand. Adjusting downward is to adjust the step length by translating downward through the field of view cropping window, corresponding to the gesture action of two fingers on one hand. Adjusting the field of view leftward and rightward is to adjust the step length by translating leftward and rightward through the field of view cropping window, corresponding to the gesture actions of three fingers and four fingers on one hand, respectively. Enlarging the field of view is to enlarge the field of view to the surroundings through the gesture action of five fingers on one hand. Reducing the field of view is to reduce the field of view to the center through the gesture action of making a fist with one hand. The field of view is adjusted through the above gestures. In summary, the present invention automatically switches the field of view according to the VCU signal of the driving state, and the algorithm automatically crops according to the set field of view thresholds and displays them on the CMS monitor. It can also control the entry and exit of the field of view adjustment function based on the gesture command collected by the DMS camera. Entering the field of view adjustment interface is to adjust the field of view through gestures.

[0082] Based on the above content, executing the operation instruction to realize the control of the CMS display content includes at least one of the following: executing the operation instruction to realize the adjustment of the display brightness of the CMS display; executing the operation instruction to realize the adjustment of the field of view size of the CMS display content; executing the operation instruction to realize the adjustment of the field of view angle of the CMS display content.

[0083] In one implementation, when a user triggers a certain viewing mode through a gesture, the CMS monitor will also make a corresponding decision to determine whether to respond to the user's trigger gesture. That is, before executing the operation instruction to control the CMS display content, it also includes: determining whether the operation instruction conflicts with the current vehicle speed or current state; if there is a conflict, prohibiting the execution of the operation.

[0084] It is understandable that during parking or driving, the user may want to switch the field of view mode of the CMS monitor through gesture triggering. For example, when the vehicle speed is less than 20km / h, the user switches to the steering field of view or the reversing field of view through gesture triggering, and the CMS monitor may respond to the user's gesture triggering; if the vehicle speed is higher than 20km / h, in order to ensure driving safety, the CMS monitor may not respond to the user's gesture triggering. This design is to avoid some field of view switching that does not conform to the actual situation, which may prevent the user from watching the video content that he really needs to know in time, thereby affecting the safe driving of the vehicle.

[0085] The control method of the CMS system described in this embodiment can match different field of view display modes in real time according to the VCU signal of the actual driving status, such as the reversing field of view mode, the turning field of view mode and the normal field of view mode, and perform field of view cropping according to different states to display key areas and blind spot information, thereby expanding the driver's field of view, improving the safety function of assisted driving, realizing human-computer interaction, and providing the driver with a more comfortable experience during driving.

[0086] In a specific implementation, the implementation process of the control method of the CMS system is as follows: Figure 5 As shown, including:

[0087] Step 501: Connect the hardware link of the DMS camera and the CMS camera, collect dynamic gesture video through the DMS camera, and transmit it to the CMS controller. The specific implementation method is as follows:

[0088] Connect the camera in the CMS rearview mirror to the CMS monitor through the Fakra harness, ensure that the DMS camera circuit is connected, and power it on. The DMS camera is fixed on the column to monitor the driver's driving status and gestures in real time. When the gesture adjustment function of the CMS monitor field of view is started, the DMS starts to collect dynamic gesture videos and transmit them to the main control module of the CMS controller for processing.

[0089] Step 502: The static gesture image is obtained by performing frame processing on the video data, and the dynamic gesture video is input to the dynamic gesture detector to detect the gesture. The specific implementation method is as follows:

[0090] The main control module of the controller receives dynamic gesture videos from CAN, and there are two types of data input to the classifier. One is to pre-process the video first, take frames of the video data to obtain static gesture images, and use them as 2D data to trigger the gesture classifier for feature extraction and classification. The second is to directly pass the dynamic gesture video data to the gesture detector for analysis, and use a sliding window with a frame number of 1 to detect the video stream frame by frame for dynamic gesture recognition. If a gesture action is detected and the gesture classifier is activated once, if no gesture action is detected, the video stream will continue to be analyzed frame by frame until the video ends.

[0091] Step 503: The dynamic gesture detector detects the gesture and activates the gesture classification module to perform gesture classification and output the recognition result. Figure 6 , the specific implementation method is as follows:

[0092] When S2 recognizes a gesture, it activates the gesture classification module once, inputs the gesture image data into the gesture classification algorithm based on the CNN architecture, and combines the C3D convolutional network and the ResNext classification network to complete gesture feature extraction and gesture classification. The 3D convolutional network uses the complete video frame data as input, and mainly learns time and space features. Compared with the 2D convolutional network that can only learn spatial information, the 3D convolutional network can retain the timing information of the video and the motion information between multiple consecutive frames. By simultaneously superimposing multiple consecutive frames to form a cube and then convolving it with the 3D kernel, the feature map of the convolutional layer can be connected to multiple consecutive frames of the previous layer, thereby capturing the motion information. Secondly, the ResNext classification network is used to classify the motion information. Its network structure is as follows Figure 7 As shown in the figure, the ideas of VGG network, ResNet network and Inception structure are integrated to form a split-convert-merge strategy. 128 channels are input and reduced to multiple 8-channel low-dimensional tensors and 1x1 convolutions. There are 16 branches in total. Each branch transforms the data through 3x3 convolution, and then converts the number of channels from 8 to 128 through 1x1 convolution. Then, the 16 128-dimensional tensors are summed, and the input is added to the final result using the cross-connection layer. Finally, the results are connected to achieve the final classification effect, and the classification result is finally output to the main control module.

[0093] Step 504: Input the recognition result into the main control module, match the gesture with the operation instruction, and adjust the brightness of the CMS monitor and the field of view of the CMS captured image. The specific implementation method is as follows:

[0094] The main control module receives the recognition results from the classification network, matches the gestures with specific operation instructions, and can adjust the brightness of the CMS monitor screen and the CMS field of view display through gesture instructions, adjust the appropriate picture display, customize the adjustment according to user needs, and intelligently switch scenes according to driving status.

[0095] Step 505: Display the final field of view on the CMS monitor. The specific implementation method is as follows:

[0096] The CMS monitor can automatically switch between the reversing field of view mode, turning field of view mode and regular field of view mode according to the reversing, turning and normal driving conditions. In the above three modes, the gesture adjustment field of view function can be activated and gestures can be input for adjustment. In addition, there is an initial custom field of view mode. The display areas of the three field of view modes during driving are preset in advance according to the driver's needs. The final adjustment results will be displayed on the monitor.

[0097] The control method of the CMS system described in the embodiment of the present application monitors the driver's gesture movements through the DMS, transmits the gesture video stream to the CMS monitor in real time, realizes the intelligent interaction between the CMS monitor display and the driver, and adjusts the monitor and field of view display through gestures, so that the monitor can be adjusted quickly and conveniently.

[0098] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0099] The method is described in detail in the embodiments disclosed in the above-mentioned application. The method of the application can be implemented by various forms of devices. Therefore, the application also discloses a device, and a specific embodiment is given below for detailed description.

[0100] Figure 7 This is a schematic diagram of the structure of a control device of a CMS system disclosed in an embodiment of the present application. Figure 7 As shown, the control device 70 of the CMS system may include:

[0101] The video acquisition module 701 is used to obtain a user video stream captured by a first camera, where the first camera is a camera capable of capturing images of the user's upper body area.

[0102] The gesture recognition module 702 is used to analyze and process the user video stream to determine whether there is a gesture image contained therein.

[0103] The gesture classification module 703 is used to classify the gesture image in the user video stream to obtain a recognized gesture when the recognition result of the gesture recognition module is yes.

[0104] The instruction determination module 704 is used to determine the operation instruction corresponding to the recognized gesture based on a preset mapping relationship between the recognized gesture and the operation instruction.

[0105] The instruction execution module 705 is used to execute the operation instruction to realize the control of the CMS display content, wherein the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.

[0106] The control device of the CMS system described in this embodiment monitors the driver's gestures through the driver video collected by the first camera. When the user inputs a set gesture, it can trigger the CMS controller to execute the corresponding operation instructions to realize the interaction between the driver and the CMS monitor. This solution does not require the user to manually operate and control based on watching the display, but can realize the control of the CMS monitor display content by making predetermined gestures. Compared with the traditional manual operation on the display or physical buttons, it is more convenient and safer to drive.

[0107] Furthermore, the present application discloses a CMS system, the structural diagram of which is shown in FIG. Figure 8 As shown. Combined Figure 8 , the CMS system 80 may include:

[0108] The CMS camera 801 is arranged in the rearview mirror outside the vehicle cabin and is used to acquire a field of view image;

[0109] The DMS camera 802 is disposed on the steering column of the steering wheel in the vehicle compartment and is used to acquire an image of the upper body area of ​​the driver;

[0110] The CMS monitor 803 includes at least one controller and two displays, and the two displays are respectively arranged on the dashboard surface near the main driver's door and the co-driver's door in the vehicle cabin. The controller is used to obtain a user video stream captured by a first camera, and the first camera is a camera that can capture images of the driver's upper body area; the user video stream is analyzed and processed to determine whether there are gesture images therein; if so, gesture classification of the gesture images in the user video stream is triggered to obtain a recognized gesture; based on a preset mapping relationship between the recognized gesture and the operation instruction, the operation instruction corresponding to the recognized gesture is determined; and the operation instruction is executed to control the CMS display content, and the display content is obtained based on the field of view image captured by the second camera, and the second camera is arranged in the rearview mirror outside the cabin.

[0111] Among them, the specific composition structure of CMS camera and CMS monitor is as follows Fig. 9 As shown, the camera is embedded in the CMS electronic rearview mirror and installed outside the cabin to collect real-time RAW images of road conditions. The image data is transmitted to the CMS monitor through the Fakra harness. The monitor is installed on the IP station in the cabin near the A-pillar (A-pillar, which is the connecting pillar connecting the roof and the front cabin on the left and right front sides) to facilitate the driver to observe the road conditions in real time. The ISP module and gesture detection algorithm described in this article are integrated in the controller in the monitor. The black level correction BLC, lens shading correction LSC, bad pixel correction, demosaic, 3A, CCM and Gamma are completed through the ISP Pipeline. The image is output at the data output interface of the ISP module. The gesture detection algorithm can realize the monitor screen brightness and field of view adjustment functions, and finally display the processed field of view on the monitor. The DMS camera is installed on the column, and its structure is as follows Fig.10 As shown in the figure, it consists of Window Cover, Led PACB, Lens, Lens holder, sealing ring, SensorPCBA, Back cover and screws. It monitors the driver's status in real time and collects dynamic gesture videos. The dynamic gesture videos are transmitted to the CMS controller through CAN. As the input of the gesture detector, it can recognize and classify dynamic gestures and output the final gesture recognition results.

[0112] The control device of any one of the CMS systems described in the above embodiments includes a processor and a memory. The video acquisition module, gesture recognition module, gesture classification module, instruction determination module, instruction execution module, etc. in the above embodiments are all stored in the memory as program modules, and the processor executes the above program modules stored in the memory to realize corresponding functions.

[0113] The processor includes a kernel, which retrieves the corresponding program module from the memory. One or more kernels can be set, and the processing of the access data can be realized by adjusting the kernel parameters.

[0114] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0115] In an exemplary embodiment, a computer-readable storage medium is also provided, which can be directly loaded into the internal memory of a computer and contains software code. After being loaded and executed by a computer, the computer program can implement the steps shown in any embodiment of the above-mentioned instruction execution module.

[0116] In an exemplary embodiment, a computer program product is also provided, which can be directly loaded into the internal memory of a computer, wherein the computer program contains software codes, and after being loaded and executed by the computer, the computer program can implement the steps shown in any embodiment of the instruction execution module described above.

[0117] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0118] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0119] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0120] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A control method for a CMS system, characterized in that: include: The CMS controller obtains a user video stream captured by a first camera, where the first camera is a camera capable of capturing an image of the driver's upper body area; Analyze and process the user video stream to determine whether there is a gesture image therein; If so, triggering gesture classification of the gesture image in the user video stream to obtain a recognized gesture; Based on a preset mapping relationship between the recognition gesture and the operation instruction, determining the operation instruction corresponding to the recognition gesture; The operation instruction is executed to realize the control of the CMS display content, wherein the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.

2. The control method of the CMS system according to claim 1, characterized in that: Also includes: obtaining a gear position signal and / or a steering signal of the vehicle; Based on the gear signal and / or the turn signal, the control is switched to a field of view mode corresponding to the gear signal and / or the turn signal according to a preset logic control, and different field of view modes correspond to different display contents.

3. The control method of the CMS system according to claim 2, characterized in that: Based on the gear signal and / or the turn signal, switching to a field of view mode corresponding to the gear signal and / or the turn signal according to a preset logic control includes: If the gear position signal indicates that the vehicle enters the reverse gear, the vehicle is controlled to enter the reverse field of view, and in the reverse field of view, the display content of the CMS display corresponds to the image content of the rear wheel part of the vehicle in the field of view image; If the turn signal indicates that the vehicle is about to turn, control the vehicle to enter the turning field of view, and in the turning field of view, the display content of the CMS display corresponds to the image content of the side part of the vehicle in the field of view image; If the gear position signal indicates that the reverse gear has been exited or the turn signal no longer exists, the control enters the normal field of view.

4. The control method of the CMS system according to claim 1, characterized in that: Also includes: Pre-configure custom field of view; If the operation instruction corresponds to a switch custom field of view instruction, the control enters the custom field of view.

5. The control method of the CMS system according to claim 1, characterized in that: Before executing the operation instruction to realize the control of the CMS display content, it also includes: Determining whether the operation instruction conflicts with the current vehicle speed or current state; If there is a conflict, the operation is prohibited from being performed.

6. The control method of the CMS system according to claim 1, characterized in that: The executing the operation instruction to realize the control of the CMS display content includes at least one of the following: Execute the operation instruction to adjust the display brightness of the CMS display; Execute the operation instruction to adjust the field size of the CMS display content; The operation instruction is executed to adjust the viewing angle of the CMS display content.

7. The control method of the CMS system according to claim 1, characterized in that: The analyzing and processing the user video stream to determine whether there is a gesture image therein comprises at least one of the following: Performing frame processing on the user video stream to obtain a frame image data set, and filtering out gesture images containing gesture actions from the frame image data set; The user video stream is detected by sliding a detection window to identify gesture images in the user video stream.

8. The control method of the CMS system according to claim 1, characterized in that: Triggering gesture classification of the gesture image in the user video stream to obtain a recognized gesture, including: The classifier model is used to classify gestures using gesture images in the user video stream; the gesture classification network of the classifier model consists of a data input layer, a convolution calculation layer, a ReLU excitation layer, a pooling layer, and a fully connected layer.

9. A control device for a CMS system, characterized in that: include: A video acquisition module, used to obtain a user video stream captured by a first camera, where the first camera is a camera capable of capturing an image of the upper body area of ​​the user; A gesture recognition module, used to analyze and process the user video stream to determine whether there is a gesture image therein; A gesture classification module, used for classifying gestures on the gesture images in the user video stream to obtain recognized gestures when the recognition result of the gesture recognition module is yes; An instruction determination module, used to determine the operation instruction corresponding to the recognition gesture based on a preset mapping relationship between the recognition gesture and the operation instruction; The instruction execution module is used to execute the operation instruction to realize the control of the CMS display content, wherein the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.

10. A CMS system, characterized in that: include: The CMS camera is installed in the rearview mirror outside the vehicle cabin and is used to collect and obtain field of view images; The DMS camera is installed on the steering column of the steering wheel in the vehicle cabin and is used to collect images of the driver's upper body area; The CMS monitor comprises at least one controller and two displays, wherein the two displays are respectively arranged on the dashboard surface close to the main driver's door and the co-driver's door in the vehicle cabin, wherein the controller is used to obtain a user video stream captured by a first camera, wherein the first camera is a camera capable of capturing an image of the driver's upper body area; analyze and process the user video stream to determine whether there is a gesture image therein; if so, trigger gesture classification of the gesture image in the user video stream to obtain a recognized gesture; Based on a preset mapping relationship between recognition gestures and operation instructions, determine the operation instruction corresponding to the recognition gesture; execute the operation instruction to realize the control of the CMS display content, the display content is obtained based on the field of view image captured by the second camera, and the second camera is set in the rearview mirror outside the cabin.