Human-vehicle interaction method and device, equipment and medium
By collecting images outside the vehicle, compressing and uploading them to the server for enhanced image reconstruction and comparison, the problem of accuracy and low efficiency of user identification outside the vehicle is solved, and efficient suspension control mode activation and user identity confirmation are achieved, improving the user experience.
Patent Information
- Application Number
- CN202510380354.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when a user is located outside the vehicle, the recognition accuracy and efficiency of human-vehicle interaction are low, especially in light changes and crowd gathering environments, resulting in low efficiency in activation of the suspension control mode.
When identifying a user outside the vehicle, the image is collected and the detection frame information is determined, the compression process is performed and uploaded to the server for enhanced image reconstruction, the user's identity is confirmed by using pre-stored user image comparison, and the suspension height change is controlled according to the user's actions.
It improves the accuracy and efficiency of off-car user identification, improves the activation efficiency of suspension control mode, ensures information security, avoids the risks brought by cloud identification, and improves the user experience.
Smart Images

Figure CN120327170A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicle-human interaction, and in particular, to a vehicle-human interaction method, device, equipment, and medium. Background Art
[0002] With the development of AI (Artificial Intelligence) technology, the applications of vehicle-human interaction have become increasingly rich. In related technologies, by collecting user images in the cockpit and recognizing the gestures of the target user in the images to identify the user's intention, so as to meet the user's interaction needs with the vehicle and achieve vehicle-human interaction.
[0003] However, related technologies have not disclosed how to lock a user for interaction when the user is outside the cockpit (i.e., outside the vehicle), lacking a vehicle-human interaction method for the scenario where the user is outside the vehicle. This is mainly because there are many difficulties in vehicle-human interaction in the outside-vehicle scenario. For example, the accuracy of target detection is easily affected by changes in light brightness, and when the user is outside the cockpit, the open site is more likely to cause the user as the detection object to be affected by light, illumination, and light and shadow changes. At the same time, the user outside the cockpit is easily among crowds and gatherings of people, which further increases the difficulty of user recognition. Summary of the Invention
[0004] Based on this, it is necessary to provide a vehicle-human interaction method, device, equipment, and medium for the above technical problems, so as to improve the user experience by enhancing the recognition accuracy and efficiency of the interaction object.
[0005] In a first aspect, an embodiment of the present application provides a vehicle-human interaction method, including:
[0006] In response to receiving a work instruction, determining detection frame information of a first user in a to-be-processed image; wherein, the work instruction includes activating an outside-vehicle control suspension mode, and the detection frame information includes first information of the facial area of the first user;
[0007] Performing compression processing on the first information to obtain second information, and uploading the second information;
[0008] Receiving third information corresponding to the second information; wherein, the third information includes an enhanced image of the facial area of the first user;
[0009] In response to the comparison result of the enhanced image and a pre-stored user image being passed, determining that the first user is the target interaction object of the outside-vehicle control suspension mode;
[0010] Controlling the change of the vehicle suspension height according to the actions of the target interaction object.
[0011] In one embodiment, determining that the first user is the target interaction object for the vehicle exterior control suspension mode in response to the comparison result between the enhanced image and the pre-stored user image being passed includes:
[0012] Determining that the first user is the target interaction object in response to the similarity between the enhanced image and the pre-stored target user image being greater than a preset similarity threshold; wherein the pre-stored list includes the target user image.
[0013] In one embodiment, after the comparison result between the enhanced image and the pre-stored user image is passed, it further includes:
[0014] Determining that the first user is not the target interaction object in response to the similarity between the enhanced image and any of the pre-stored user images being less than or equal to the preset similarity threshold;
[0015] Then determining not to activate the vehicle exterior control suspension mode.
[0016] In one embodiment, determining the detection frame information of the first user in the image to be processed includes:
[0017] Obtaining task description information from the work instruction;
[0018] Matching the target image acquisition rule corresponding to the task description information from a preset plurality of image acquisition rules;
[0019] Collecting the image to be processed based on the target image acquisition rule;
[0020] Using a target detection model to identify the first user in the image to be processed and generating the detection frame information.
[0021] In one embodiment, the task description information includes the physical coordinates of the first user;
[0022] The collecting the image to be processed based on the target image acquisition rule includes:
[0023] In response to the target image acquisition rule being the first rule, for the area where the physical coordinates are located, controlling a first image acquisition device to perform image acquisition according to the first working parameter in the first rule to obtain the image to be processed; or,
[0024] In response to the target image acquisition rule being the second rule, controlling the vehicle lamp to irradiate the area where the physical coordinates are located according to the first lighting parameter in the second rule, and controlling a second image acquisition device to perform image acquisition for the irradiated area according to the second working parameter in the second rule to obtain the image to be processed.
[0025] In one embodiment, determining the detection box information of the first user in the image to be processed includes:
[0026] For the image to be processed, determining a first target detection box through a target detection model;
[0027] In response to the number of the first target detection boxes being greater than 1, determining the target in the first target detection box with the largest area as the first user;
[0028] Based on the first target detection box with the largest area, determining a second target detection box; wherein, the second target detection box includes the face of the first user.
[0029] In one embodiment, determining the detection box information of the first user in the image to be processed includes:
[0030] For the image to be processed, determining a third target detection box through a target detection model;
[0031] In response to the pose of the target in the third target detection box being consistent with the preset pose information, determining the target in the third target detection box as the first user
[0032] In the third target detection box, determining a fourth target detection box; wherein, the fourth target detection box includes the face of the first user.
[0033] In one embodiment, determining the fourth target detection box in the third target detection box includes:
[0034] In the image to be processed, expanding the area of the third target detection box based on a preset rule to obtain a fifth target detection box; wherein, the fifth target detection box contains the third target detection box;
[0035] In the fifth target detection box, determining a sixth target detection box with a number less than a preset quantity threshold; wherein, the sixth target detection box includes the target part of the first user;
[0036] Based on the category information of the target part, selecting the fourth target detection box in the sixth target detection box.
[0037] In one embodiment, compressing the first information to obtain second information includes:
[0038] Inputting the facial region of the first user into a pre-trained encoder to obtain a first vector with a preset length; then the first vector is the second information.
[0039] In one embodiment, the work instruction further includes activating the in-vehicle controlled suspension mode:
[0040] The method further includes: in response to the suspension control mode being the in-vehicle controlled suspension mode, determining the facial region of the first user in the image to be processed;
[0041] Comparing the facial region with the pre-stored user image to determine whether the first user is the target interaction object of the in-vehicle controlled suspension mode.
[0042] In one embodiment, after determining the facial region of the first user in the image to be processed, it further includes:
[0043] Obtaining the interaction mode in the work instruction; wherein, the interaction mode includes a virtual user interaction mode and / or a real user interaction mode;
[0044] The comparing the facial region with the pre-stored user image to determine whether the first user is the target interaction object of the in-vehicle controlled suspension mode includes:
[0045] In response to the similarity between the facial region and the user image being greater than a preset similarity threshold, determining the target interaction object of the in-vehicle controlled suspension mode based on the interaction mode.
[0046] In one embodiment, the determining the target interaction object based on the interaction mode includes: determining the target interaction object of the in-vehicle controlled suspension mode based on the interaction mode;
[0047] Controlling the height of the suspension according to the action of the target interaction object.
[0048] In a second aspect, an embodiment of the present application provides a vehicle-human interaction device, including:
[0049] A first detection module, configured to determine the detection frame information of the first user in the image to be processed in response to receiving a work instruction; wherein, the work instruction includes activating the out-vehicle controlled suspension mode; the detection frame information includes first information of the facial region of the first user;
[0050] A compression module, configured to perform compression processing on the first information to obtain second information and upload the second information;
[0051] An enhancement module, configured to receive third information corresponding to the second information; wherein, the third information includes an enhanced image of the facial region of the first user;
[0052] A comparison module, configured to determine that the first user is the target interaction object for the vehicle exterior control suspension mode in response to the comparison result between the enhanced image and a pre-stored user image being passed.
[0053] A control module, configured to control the change of the vehicle suspension height according to the actions of the target interaction object.
[0054] In one embodiment, the comparison module is specifically configured to determine that the first user is the target interaction object in response to the similarity between the enhanced image and the pre-stored target user image being greater than a preset similarity threshold; wherein the pre-stored list includes the target user image.
[0055] In one embodiment, the comparison module is further configured to determine that the first user is not the target interaction object in response to the similarity between the enhanced image and any of the pre-stored user images being less than or equal to the preset similarity threshold.
[0056] In one embodiment, the first detection module is specifically configured to obtain task description information from the work instruction; match the target image acquisition rule corresponding to the task description information from a preset plurality of image acquisition rules; acquire the image to be processed based on the target image acquisition rule; and identify the first user in the image to be processed by using a target detection model and generate the detection frame information.
[0057] In one embodiment, the task description information includes the physical coordinates of the first user; the first detection module is specifically configured to, in response to the target image acquisition rule being the first rule, control a first image acquisition device to perform image acquisition according to the first working parameter in the first rule for the area where the physical coordinates are located to obtain the image to be processed; or, in response to the target image acquisition rule being the second rule, control the vehicle lamp to irradiate the area where the physical coordinates are located according to the first illumination parameter in the second rule, and control a second image acquisition device to perform image acquisition for the irradiated area according to the second working parameter in the second rule to obtain the image to be processed.
[0058] In one embodiment, the first detection module is further configured to determine a first target detection frame for the image to be processed; in response to the number of the first target detection frames being greater than 1, determine that the target in the first target detection frame with the largest area is the first user; and determine a second target detection frame based on the first target detection frame with the largest area; wherein the second target detection frame includes the face of the first user.
[0059] In one embodiment, the first detection module is further configured to, for the image to be processed, determine a third target detection box through a target detection model; in response to the posture of the target in the third target detection box being consistent with the preset posture information, determine that the target in the third target detection box is the first user; and determine a fourth target detection box in the third target detection box; wherein, the fourth target detection box includes the face of the first user.
[0060] In one embodiment, the first detection module is further configured to, in the image to be processed, expand the area of the third target detection box based on a preset rule to obtain a fifth target detection box; wherein, the fifth target detection box contains the third target detection box; determine a sixth target detection box with a number less than a preset quantity threshold in the fifth target detection box; wherein, the sixth target detection box includes the target part of the first user; and select the fourth target detection box from the sixth target detection boxes based on the category information of the target part.
[0061] In one embodiment, the compression module is specifically configured to input the face area of the first user into a pre-trained encoder to obtain a first vector with a preset length; and the first vector is the second information.
[0062] In one embodiment, the work instruction further includes activating the in-vehicle control suspension mode: the device further includes a second detection module, and the second detection module is specifically configured to, in response to the suspension control mode being the in-vehicle control suspension mode, determine the face area of the first user in the image to be processed; compare the face area with the pre-stored user image to determine whether the first user is the target interaction object of the in-vehicle control suspension mode.
[0063] In one embodiment, the second detection module is further configured to obtain the interaction mode in the work instruction; wherein, the interaction mode includes a virtual user interaction mode and / or a real user interaction mode; in response to the similarity between the face area and the user image being greater than a preset similarity threshold, determine the target interaction object of the in-vehicle control suspension mode based on the interaction mode.
[0064] Fourthly, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps of the method according to the first aspect and any of its embodiments are implemented.
[0065] Fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the first aspect and any of its embodiments are implemented.
[0066] Sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, which when executed by a processor implements the method described in the first aspect and any of its embodiments.
[0067] After receiving a work instruction, the above-mentioned vehicle-human interaction method, device, equipment and medium perform preliminary processing on the image to be processed to identify the facial area of the first user in the image to be processed, perform compression processing, and upload the second information obtained by compression to obtain an enhanced image corresponding to the facial area of the first user, thereby improving the resolution of the facial area of the first user; accordingly, comparing with the user images in the pre-stored list effectively improves the accuracy and efficiency of identifying the first user, thereby improving the activation efficiency of the vehicle exterior control suspension mode and effectively improving the user experience.
[0068] Because this vehicle-human interaction method has low requirements for vehicle-end computing capabilities, storage space and other performances, this method is applicable to a variety of vehicle models and has the advantage of a wide application range; and it can effectively avoid the problem of efficiency decline caused by locally loading a large number of deep learning models and their parameters.
[0069] That is to say, this vehicle-human interaction method realizes efficient information transmission by uploading the second information obtained by compression, and by comparing the enhanced image with the user images in the pre-stored list, it can effectively improve the accuracy and avoid directly comparing the image to be processed with the pre-stored list. When the light is too strong or too weak, or the pixel of the image acquisition device is insufficient and other interference factors cause the facial area of the first user in the image to be processed to be unclear, resulting in a decrease in the user recognition accuracy, so that the user repeatedly sends interaction requests or work instructions to the vehicle but cannot obtain the vehicle control permission and cannot activate the corresponding suspension control mode, thereby leading to a decrease in the suspension control mode activation efficiency. At the same time, it also effectively avoids the information security risk caused by pre-storing the pre-stored list in the cloud. Description of the Drawings
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0071] Figure 1 It is an application environment diagram of the vehicle-human interaction method in an embodiment;
[0072] Figure 2 is a schematic flowchart of a human-vehicle interaction method in an embodiment;
[0073] Figure 3 is a schematic flowchart of the step of determining the detection box information of the first user in the to-be-processed image in an embodiment;
[0074] Figure 4 is a schematic flowchart of the step of determining the detection box information of the first user in the to-be-processed image in another embodiment;
[0075] Figure 5 is a schematic diagram of the human-vehicle interaction method when the user interacts with the vehicle in an embodiment;
[0076] Figure 6 is a structural block diagram of a human-vehicle interaction device in an embodiment;
[0077] Figure 7 is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0078] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are only some of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the protection scope of this application. Without conflict, the embodiments in this application and the features in the embodiments can be combined arbitrarily with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0079] The terms "first" and "second" in the description and claims of this application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. "Multiple" in this application may represent at least two, for example, it may be two, three, or more, and the embodiments of this application do not make limitations.
[0080] The human-vehicle interaction method provided by this application can be applied to, for example Figure 1In the application environment shown, the terminal 102 communicates with the server 104 via a network. After receiving a work instruction, the terminal 102 processes the to-be-processed image collected by the image acquisition device to determine the detection box information of the first user in the to-be-processed image. Moreover, after compressing the first information of the facial region of the first user in the detection box information, the second information obtained by compression is uploaded to the server 104 via the network, so as to process the second information by virtue of the high computing power characteristics of the server 104, realize the high-definition restoration of the face of the first user in the second information, and obtain the third information. In this way, when the terminal 102 receives the third information returned by the server 104, it can identify the identity of the first user locally, thereby realizing the efficient identification of the user identity while ensuring information security and avoiding the risk of leakage of the user's personal information (facial information).
[0081] Among them, the terminal 102 may include, but is not limited to, a vehicle-mounted controller, a vehicle-mounted terminal, a driving computer, etc.
[0082] And, the vehicle-mounted controller may include, but is not limited to, a VCU (Vehicle Control Unit, vehicle whole controller), an MCU (Microcontroller Unit, microcontroller unit), an ECU (Electronic Control Unit, electronic control unit), etc.
[0083] The above-mentioned server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0084] In one embodiment, as Figure 2 shown, a vehicle-human interaction method is provided, and this method can be applied to the Figure 1 terminal in. Taking the vehicle-mounted controller as an example of the terminal for illustration below, this method includes the following steps:
[0085] Step 201, in response to receiving a work instruction, determine the detection box information of the first user in the to-be-processed image.
[0086] Among them, the to-be-processed image includes the above-mentioned first user, and the detection box information includes the first information of the facial region of the first user.
[0087] Specifically, the above-mentioned detection box information may further include the hand region information and / or the whole body region information of the first user.
[0088] The above-mentioned work instruction may indicate to identify the target (person) in the to-be-processed image. Then, the work instruction may include the physical coordinates and / or pose information of the first user.
[0089] Exemplarily, if the work instruction includes the physical coordinates of the first user, the image to be processed can be collected accordingly, and the physical coordinates can be converted into image coordinates to identify the first user in the image to be processed.
[0090] Exemplarily, if the work instruction includes the pose information of the first user, the user pose matching the pose information can be identified in the image to be processed accordingly, and the user with the user pose can be determined as the first user, and the detection box information of the first user can be determined.
[0091] The above work instruction may include activating the out-of-vehicle control suspension mode; and / or, activating the in-vehicle control suspension mode.
[0092] The above out-of-vehicle control suspension mode can be used to indicate that the vehicle controls the suspension movement according to the actions of the target interaction object located outside the vehicle.
[0093] Then, based on the image to be processed, the first information of the first user and its facial area therein can be determined, so as to facilitate the execution of step 204: determining whether the first user is the target interaction object.
[0094] Step 202, perform compression processing on the first information to obtain second information, and upload the second information.
[0095] Specifically, the processing method of the compression processing may include but is not limited to Huffman coding, feature extraction model, and the encoder processes the area where the face of the first user is located.
[0096] Preferably, a pre-trained encoder is used to process the area where the face is located for easy uploading; then the encoded result is the second information.
[0097] In a possible implementation manner, the image to be processed can be cropped based on the first information first to obtain the facial area of the first user.
[0098] Then, the facial area of the first user can be input into a pre-trained encoder to obtain a first vector with a preset length.
[0099] Then the first vector is the second information. Wherein, the facial area corresponds to the facial detection box of the first user.
[0100] The above encoder may include at least one of Convolutional Encoder, Self-Attention Encoder, Residual Encoder, GAN Encoder, and swintransformer (Shifted Window Transformer).
[0101] The above-mentioned second information can be uploaded to the server or the cloud, so that the server or the cloud performs super-resolution reconstruction processing on the second information to obtain third information.
[0102] Due to the limited computing power and storage space of the vehicle-mounted controller, in order to improve the activation efficiency of the suspension control mode and thus enhance the user experience, in one embodiment, it can be achieved by improving the accuracy and efficiency of identifying the target interaction object. The above-mentioned second information can be uploaded to the server or the cloud, so that after receiving the second information, the server or the cloud, based on the advantages of its storage space and computing power, performs super-resolution reconstruction processing on the second information to obtain third information.
[0103] That is, the server or the cloud can pre-store the trained deep learning model. This deep learning model is used to process the second information and output the third information.
[0104] To further enhance information security, in a possible implementation manner, a preset encryption algorithm can be used to encrypt the second information and then upload it to the server or the cloud. After receiving the encrypted information obtained by encrypting the second information, the server or the cloud decrypts the encrypted information based on the preset encryption algorithm and then performs super-resolution reconstruction processing to obtain the third information.
[0105] Step 203: Receive the third information corresponding to the second information.
[0106] Among them, the third information includes the enhanced image of the facial area of the first user.
[0107] Specifically, the above-mentioned third information can carry the same or corresponding identifier as the second information, so as to facilitate determining that this third information is the information corresponding to the second information.
[0108] The above-mentioned enhanced image is a super-resolution image corresponding to the face of the first user.
[0109] Step 204: In response to the comparison result between the enhanced image and the pre-stored user image being passed, determine that the first user is the target interaction object for controlling the suspension mode outside the vehicle.
[0110] Specifically, it can be determined that the first user is the target interaction object according to the similarity between the enhanced image and the pre-stored user images.
[0111] The pre-stored user images are all facial images of users who have been authenticated or authorized. And this facial image meets the preset resolution.
[0112] Among them, the clarity and resolution of the pre-stored user images are each greater than or equal to the clarity and resolution of the enhanced image.
[0113] In a possible implementation, the first user can be determined as the target interaction object by similarity. Specifically, in response to the similarity between the enhanced image and the pre-stored target user image being greater than the preset similarity threshold, the first user is determined as the target interaction object.
[0114] Among them, the target user image is included in the aforementioned user list. That is, the target user image is one of the user images in the aforementioned user list.
[0115] Step 205, control the change of the vehicle suspension height according to the action of the target interaction object.
[0116] Specifically, since the out-of-vehicle control suspension mode can be used to indicate that the vehicle controls the suspension movement according to the actions of the target interaction object located outside the vehicle, after the work instruction is executed to activate the out-of-vehicle control suspension mode, the change of the suspension height can be controlled according to the actions of the target interaction object.
[0117] In the above human-vehicle interaction method, after determining the detection box information of the first user in the image to be processed, by compressing the facial area of the first user in the detection box information and uploading it to the cloud, an enhanced image can be obtained efficiently and accurately, and then the identity of the first user can be recognized efficiently and accurately, effectively improving the activation efficiency of the out-of-vehicle control suspension mode. At the same time, it also realizes the protection of user privacy and avoids the information security risks caused by identifying the user's identity in the cloud.
[0118] In one embodiment, after the comparison result between the enhanced image and the pre-stored user image in the aforementioned step 204 is not passed, it may further include: determining that the first user is not the target interaction object; then determining not to execute the work instruction: not activating the out-of-vehicle control suspension mode.
[0119] Specifically, in response to the similarity between the enhanced image and any pre-stored user image being less than or equal to the preset similarity threshold, it is determined that the first user is not the target interaction object. Then it is determined not to activate the out-of-vehicle control suspension mode.
[0120] Further, the determination of the detection box information of the first user in the aforementioned image to be processed will be described in detail below. Please refer to Figure 3 .
[0121] Step 301, obtain the task description information from the work instruction.
[0122] The task description information can be determined, for example, when the user generates a need to interact with the vehicle-mounted controller, according to the work instruction sent by the user terminal on the user side to the vehicle-mounted controller, so as to determine the corresponding target image acquisition rule according to the task description information, and call the image acquisition device to perform image acquisition using the target image acquisition rule.
[0123] When a user has a need for vehicle - person interaction before entering the cockpit, the task description information can be determined through their user terminal, and then a work instruction containing the task description information is sent to the vehicle - mounted controller through the user terminal.
[0124] The above - mentioned task description information may include the physical coordinates of the first user. The determination method of the physical coordinates may include, but is not limited to, being obtained by positioning the user terminal.
[0125] The task description information at least includes scene information, so that the vehicle - mounted controller can determine the image acquisition rule according to the scene information in the task description information. The scene information can indicate the scene where the user to be recognized is located.
[0126] Exemplarily, the scene information in the above - mentioned task description information may include, but is not limited to, at least one of a day mode, a night mode, and a lighting mode.
[0127] Furthermore, the above - mentioned day mode may further include, but is not limited to, at least one of a sunny day mode, a cloudy day mode, and a foggy day mode. And the night mode may include, but is not limited to, an evening mode and / or a dark night mode.
[0128] Then the scene information in the task description information may include, but is not limited to, at least one of a sunny day mode, a cloudy day mode, a foggy day mode, an evening mode, or a dark night mode.
[0129] Step 302: Match the target image acquisition rule corresponding to the task description information from a plurality of preset image acquisition rules.
[0130] Specifically, the above - mentioned task description information can correspond one - to - one with the target image acquisition rule.
[0131] In a possible implementation manner, the target image acquisition rule matching the task description information can be determined based on a preset correspondence.
[0132] The preset correspondence may include scene information and the target image acquisition rule corresponding one - to - one with the scene information.
[0133] Exemplarily, if the scene information in the task description information is the day mode, then among a plurality of preset image acquisition rules, the first rule can be matched for this day mode. The first rule includes the first working parameters of the image acquisition device. The first working parameters may include, but are not limited to, shooting parameters such as the aperture corresponding to the scene information.
[0134] Exemplarily, if the scene information in the task description information is the night mode, then among the preset multiple image acquisition rules, the second rule can be matched for this night mode. The second rule includes the first working parameters of the image acquisition device and the first lighting parameters of the vehicle lights. The first working parameters may include, but are not limited to, shooting parameters such as the aperture corresponding to the scene information.
[0135] The first lighting parameters of the vehicle lights may include, but are not limited to, at least one of the lighting time, lighting duration, light intensity, and illuminance of each of the multiple vehicle lights.
[0136] Exemplarily, if the scene information in the task description information is the light 1 mode, then among the aforementioned preset multiple image acquisition rules, the third rule can be matched for this light mode.
[0137] The third rule includes the third working parameters of the image acquisition device and the second working parameters of the vehicle lights. Since the scene information is the light mode, the requirements for the light brightness and color change in the second working mode here are higher than those in the first working mode of the vehicle lights.
[0138] The above image acquisition device can be installed at multiple preset positions on the vehicle body to facilitate the acquisition of to-be-processed images in multiple different directions.
[0139] The above image acquisition device may include, but is not limited to, a camera, a video camera, or a pan-tilt head that can acquire images at preset time intervals.
[0140] Step 303: Acquire the to-be-processed image based on the target image acquisition rule.
[0141] Specifically, in response to the target image acquisition rule being the first rule, for the area where the physical coordinates of the first user are located, control the first image acquisition device to acquire images according to the first working parameters in the first rule, and obtain the to-be-processed image.
[0142] Or, in response to the target image acquisition rule being the second rule, control the vehicle lights to irradiate the area where the physical coordinates of the first user are located, and for this irradiated area (i.e., the aforementioned area where the physical coordinates are located), control the second image acquisition device to acquire images according to the second working parameters in the second rule, and obtain the to-be-processed image.
[0143] Then, the irradiated area obtained by the above vehicle light irradiation not only plays a role in guiding the user to interact, but also can distinguish the first user to be recognized from other people in the same area through the irradiated area, promoting the improvement of the accuracy of identifying the first user in step 304.
[0144] The above first image acquisition device and the second image acquisition device can be the same or different.
[0145] The determination method of the area where the physical coordinates of the first user are located may include, but is not limited to: constructing a rectangular frame with a preset length and width centered on the physical coordinates of the first user as the area where the physical coordinates are located.
[0146] In one embodiment, the above first image acquisition device or second image acquisition device can be determined by the following method:
[0147] First, according to the physical coordinates of the first user in the task description information and the positioning information of the vehicle-mounted controller, determine the relative position relationship between the first user and the vehicle.
[0148] Then, according to this relative position relationship, determine among the image acquisition devices installed on the side of the vehicle body closest to the first user.
[0149] Step 304, use the target detection model to identify the first user in the image to be processed and generate detection box information.
[0150] Specifically, the structure of the target detection model may include, but is not limited to, at least one of FCOS (Fully Convolutional One-Stage Object Detection), R-CNN (Region-based Convolutional Neural Networks), YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), and Deformable ConvNets.
[0151] The target detection model is a model obtained by pre-training with an accuracy not lower than the preset accuracy threshold.
[0152] Furthermore, to improve the accuracy of the first user, in a possible implementation manner, the above target detection model includes multiple detection algorithms, and each detection algorithm corresponds to the scene information one by one.
[0153] As mentioned above, the scene information corresponds to the image acquisition rule. Therefore, the above target detection model includes a first detection model corresponding to the first rule and a second detection model corresponding to the second rule.
[0154] To avoid the problem that it is difficult to distinguish the first user from pedestrians due to a large number of pedestrians around the first user and / or the distance between the pedestrians and the first user being too close, in one embodiment, the detection box information of the first user in the image to be processed is determined by the following method: First, for the image to be processed, determine the first target detection box through the target detection model.
[0155] In response to the number of the first target detection boxes being 1, it is determined that the target in the first target detection box is the first user. Then, the detection box information of the first user can be determined accordingly. The detection box information of the first user may include the position parameters of the first target detection box. The position parameters of the first target detection box may include, but are not limited to, the coordinates of the vertices that are on the diagonal of the target detection box and are the vertices of the first target detection box.
[0156] Alternatively, in response to the number of the first target detection boxes being greater than 1, it is determined that the target in the first target detection box with the largest area is the first user. This is because the detection box with the largest area means that the target simultaneously satisfies the conditions of being close enough to the image acquisition device and having the front facing the image acquisition device; in this way, it is realized to determine that the target is the first user with the intention of interacting with the vehicle.
[0157] Finally, based on the first target detection box with the largest area mentioned above, a second target detection box containing the face of the first user can be determined. Specifically, target detection can be continued in the first target detection box to efficiently and accurately identify the face detection box, which is the face of the first user.
[0158] Then, the detection box information of the aforementioned first user further includes the position parameters of the second target detection box.
[0159] To further improve the accuracy of identifying the first user, in one embodiment, the above work instruction may include preset pose information, such as raising a hand with the palm facing the vehicle.
[0160] The detection box information of the first user in the image to be processed can be determined in the following manner. Please refer to Figure 4 :
[0161] Step 401, for the image to be processed, determine a third target detection box through a target detection model.
[0162] The target detection model is a pre-trained target detection model.
[0163] The target detection model detects the target in the image to be processed through the target detection box.
[0164] Step 402, in response to the pose of the target in the third target detection box being consistent with the aforementioned preset pose information, determine that the target in the third target detection box is the first user.
[0165] Step 403, in the third target detection box, determine a fourth target detection box.
[0166] Among them, the fourth target detection box includes the face of the first user. The third target detection box includes the fourth target detection box.
[0167] Specifically, first, in the image to be processed, the area of the third target detection box can be enlarged based on a preset rule to obtain a fifth target detection box. The fifth target detection box contains the third target detection box. The preset rule can be, for example, making the area of the fifth target detection box 1.5 times the area of the third target detection box.
[0168] Then, in the fifth target detection box, determine a sixth target detection box whose number is less than a preset quantity threshold. The sixth target detection box includes the target part of the first user. The detection box information of the sixth target detection box includes the category information of the target part. The detection box information of the sixth target detection box may also include the confidence level of the category information of the target part.
[0169] The above preset quantity threshold corresponds one-to-one with the category information of the target part. For example, if the category information is hand, the preset quantity threshold is 3. For another example, if the category information is face, the preset quantity threshold is 2. In this way, the detection box can be restricted by the preset threshold to improve the detection performance of the target part.
[0170] Finally, based on the category information of the target part, select a fourth target detection box in the sixth target detection box. The category information of the fourth target detection box is pre-set.
[0171] Furthermore, in response to the posture of the target in the third target detection box being inconsistent with the foregoing preset posture information, re-acquire the image to be processed, and re-execute steps 401-403 for this image to be processed. Optionally, the foregoing preset posture information can also be directly obtained when executing the foregoing step 402, rather than being read through the foregoing work instruction.
[0172] Furthermore, the above work instruction may further include activating the in-vehicle control suspension mode. Then, when the suspension control mode is the in-vehicle control suspension mode, it can be determined that the first user is in the vehicle, and then the image to be processed collected by the in-vehicle image acquisition device can be obtained. In this case, due to the limited space in the vehicle, the number of people is relatively small, and most of the area in the image collected by the image acquisition device is the face area, and it has the characteristics of high clarity and resolution.
[0173] Thus, in order to activate the in-vehicle control suspension mode, in one embodiment, the human-vehicle interaction method may further include:
[0174] First, in response to the suspension control mode being the in-vehicle control suspension mode, determine the face area of the first user in the image to be processed.
[0175] Then, the facial region can be compared with the pre-stored user images. When the similarity between the facial region and any user image in the pre-stored list is greater than the preset similarity threshold, it is determined that the first user is the target interaction object for controlling the suspension mode in the vehicle. Then it is determined that the in-vehicle control suspension mode can be activated.
[0176] To further improve the user experience, in one embodiment, after determining the facial region of the first user in the image to be processed, the interaction mode in the work instruction can be obtained. This interaction mode can correspond to the in-vehicle control suspension mode.
[0177] Among them, the interaction mode includes a virtual user interaction mode and / or a real user interaction mode.
[0178] Then when performing the facial comparison to determine whether the first user is the target interaction object: in response to the similarity between the facial region and any pre-stored user image being greater than the aforementioned preset similarity threshold, based on the interaction mode, the target interaction object for the in-vehicle control suspension mode is determined.
[0179] In one embodiment, based on the interaction mode, determining the target interaction object for the in-vehicle control suspension mode includes: based on the interaction mode, determining the target interaction object, activating the in-vehicle control suspension mode, and controlling the height of the suspension according to the actions of the target interaction object.
[0180] Exemplarily, if the interaction mode is a virtual user interaction mode, the target interaction object is a virtual user. This virtual user can be displayed on the vehicle's central control screen, and the actions of the target interaction object (such as gestures) are shown by real-time collecting the user's posture information or voice commands, so as to control the suspension height by corresponding the actions of the virtual user one by one with the actions of the target object.
[0181] Or, in response to the similarity between the facial region and any pre-stored user image being less than or equal to the aforementioned preset similarity threshold, the current user in the vehicle is not the target interaction object, and it can be determined not to activate the in-vehicle control suspension mode.
[0182] The above steps 201 to 204 are particularly applicable to the scenario where the first user is outside the vehicle, has an interaction intention, and interacts with the vehicle. The following combines Figure 5 with steps 201 to 204 for further illustration:
[0183] When the user gradually approaches the vehicle, the user uses their handheld terminal to send an unlocking message to the vehicle through the network or Bluetooth, etc., to unlock the vehicle.
[0184] If the user has the intention of interacting outside the vehicle, that is, the user can send a work instruction to the vehicle before entering the vehicle. Optionally, the work instruction can be generated by the handheld terminal according to the information input by the user after the user inputs the corresponding information to the handheld terminal, and then sent to the vehicle terminal. Alternatively, after the vehicle terminal is awakened, the vehicle terminal issues a voice prompt, such as "Please speak out the scene information", etc., and the user speaks out the voice password for the vehicle terminal to collect; after the vehicle terminal collects the user's voice password, a work instruction is generated based on the voice password.
[0185] The work instruction includes information such as the physical coordinates of the user and the scene information. After receiving the work instruction, the vehicle can match the corresponding image acquisition rules for the scene information in the work instruction. For example, the scene information includes daytime (mode), nighttime (mode), and lighting mode. Then the following image acquisition rules can be matched for the foregoing scene information: the first rule corresponding to daytime (mode), the second rule corresponding to nighttime (mode), or the third rule corresponding to the lighting mode. The vehicle can also determine the image acquisition device for collecting the image to be processed through the physical coordinates of the work instruction, and the image acquisition device is mounted outside the vehicle body. And the working parameters can be determined for the image acquisition device in combination with the physical coordinates, scene information, etc., to ensure that the image to be processed containing the user is collected. Then the vehicle configures the working parameters for the image acquisition device according to the foregoing information, and performs image acquisition through the image acquisition device to obtain the image to be processed.
[0186] Perform object detection on the image to be processed to determine the target detection frame containing the user. The determination method includes, but is not limited to, matching the target posture in the target detection frame with the preset posture information in the work instruction. If the match is successful, it is determined that the target in the target detection frame is the user. Then continue to determine the face of the user based on the target detection frame for encoding and uploading, so as to perform face super-resolution image reconstruction in the cloud to obtain the enhanced image. Finally, the vehicle can use the pre-stored and authenticated user image for comparison to determine whether to execute the work instruction, that is, whether to activate the corresponding suspension control mode.
[0187] If the comparison is successful, it is determined that the user is authenticated, and the suspension control mode outside the vehicle can be activated, so as to control the height change of the vehicle suspension according to the actions of the target interaction object. If the comparison fails, it is determined that the user is not authenticated, and the vehicle does not execute the work instruction: the suspension control mode outside the vehicle is not activated temporarily.
[0188] After the user unlocks the vehicle, if the user's interaction intention is in-vehicle interaction, the user can, after entering the vehicle, select the corresponding physical button and / or virtual button, cooperate with the in-vehicle instruction, and face the in-vehicle image acquisition device, so that the in-vehicle image acquisition device acquires the user's high-definition face image. Then, by directly comparing the face area in the high-definition face image with the pre-stored user image (i.e., the user's face image), it can be determined whether the user is the target interaction image, so as to determine whether to activate the in-vehicle controlled suspension mode.
[0189] Based on the same inventive concept, an embodiment of the present application further provides a vehicle-human interaction method, which can be applied to the cloud or a server and is used to process second information. The second information is the first information of the face of the user to be recognized compressed by a local terminal (vehicle-mounted controller) and then uploaded. Then, the method may include the following steps:
[0190] First, receive the second information. The second information may include the address information of the local terminal.
[0191] Then, perform enhancement processing on the second information through a pre-trained deep learning model to obtain third information. The third information includes an enhanced image of the face of the user to be recognized (i.e., the first user). Specifically, the pre-trained deep learning model may be matched with the model or encoder for compression processing at the local end (vehicle-mounted controller) to efficiently generate a high-definition and accurate super-resolution image as the enhanced image. Exemplarily, the vehicle-mounted controller generates the second information through a pre-trained encoder. Correspondingly, the deep learning model here is an embedding layer and a decoder that are matched with the encoder. Then, the encoder, the embedding layer, and the decoder can be pre-trained together to improve the accuracy of the enhanced image. The pre-training method can be implemented, for example, through Distributed Federated Learning (DFL).
[0192] Next, based on the address information in the foregoing second information, send the third information, so that the vehicle-mounted controller compares the enhanced image in the third information with the pre-stored user image locally to determine the identity of the user to be recognized, thereby effectively improving information security.
[0193] It should be understood that although Figures 2 - 4 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figures 2 - 4At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed and completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0194] Based on the same inventive concept, as Figure 6 shown, an embodiment of the present application provides a vehicle-human interaction device, including: a first detection module 601, a compression module 602, an enhancement module 603, and a comparison module 604, where:
[0195] The first detection module 601 is configured to determine detection frame information of a first user in a to-be-processed image in response to receiving a work instruction. Wherein, the work instruction includes activating an external vehicle control suspension mode, and the detection frame information includes first information of a facial area of the first user.
[0196] The first detection module 601 is specifically configured to: obtain task description information from the work instruction; match the task description information with a target image acquisition rule corresponding to the task description information from a plurality of preset image acquisition rules; collect the to-be-processed image based on the target image acquisition rule; and use a target detection model to identify the first user in the to-be-processed image and generate the detection frame information.
[0197] The task description information includes physical coordinates of the first user; the first detection module 601 is specifically configured to:
[0198] In response to the target image acquisition rule being a first rule, for the area where the physical coordinates are located, control a first image acquisition device to perform image acquisition according to first working parameters in the first rule to obtain the to-be-processed image; or, in response to the target image acquisition rule being a second rule, control a vehicle lamp to irradiate the area where the physical coordinates are located according to first lighting parameters in the second rule, and control a second image acquisition device to perform image acquisition on the irradiated area according to second working parameters in the second rule to obtain the to-be-processed image.
[0199] The first detection module 601 is further configured to:
[0200] For the to-be-processed image, determine a first target detection frame through a target detection model; in response to the number of the first target detection frames being greater than 1, determine that the target in the first target detection frame with the largest area is the first user; and based on the first target detection frame with the largest area, determine a second target detection frame; wherein, the second target detection frame includes the face of the first user.
[0201] The work instruction includes preset pose information; the first detection module 601 is further configured to:
[0202] For the to-be-processed image, determine a third target detection frame through a target detection model; in response to the pose of the target in the third target detection frame being consistent with the preset pose information, determine the target in the third target detection frame as the first user; in the third target detection frame, determine a fourth target detection frame; wherein, the fourth target detection frame includes the face of the first user.
[0203] The first detection module 601 is further configured to:
[0204] In the to-be-processed image, expand the area of the third target detection frame based on a preset rule to obtain a fifth target detection frame; wherein, the fifth target detection frame contains the third target detection frame; in the fifth target detection frame, determine sixth target detection frames with a number less than a preset quantity threshold; wherein, the sixth target detection frames include the target parts of the first user; based on the category information of the target parts, select the fourth target detection frame from the sixth target detection frames.
[0205] The compression module 602 is configured to perform compression processing on the first information to obtain second information, and upload the second information.
[0206] Specifically, the compression module 602 is configured to input the face area of the first user into a pre-trained encoder to obtain a first vector with a preset length; then the first vector is the second information.
[0207] The enhancement module 603 is configured to receive third information corresponding to the second information. Wherein, the third information includes an enhanced image of the face area of the first user.
[0208] The comparison module 604 is configured to, in response to the comparison result between the enhanced image and a pre-stored user image being passed, determine the first user as the target interaction object for the vehicle exterior control suspension mode.
[0209] The control module 605 is configured to control the change of the vehicle suspension height according to the actions of the target interaction object.
[0210] Specifically, the comparison module 604 is configured to, in response to the similarity between the enhanced image and the pre-stored target user image being greater than a preset similarity threshold, determine the first user as the target interaction object; wherein, the pre-stored list includes the target user image.
[0211] For the specific limitations of the vehicle - human interaction device, reference can be made to the limitations of the vehicle - human interaction method in the above text, which will not be elaborated here. Each module in the above - mentioned vehicle - human interaction device can be implemented in whole or in part by software, hardware, or a combination thereof. The above - mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above - mentioned modules.
[0212] Based on the same inventive concept, please refer to Figure 7 , an embodiment of the present application also provides an electronic device. In one embodiment, as shown in the figure, the computer device may include a memory 701, a communication module 703, and one or more processors 702.
[0213] The memory 701 is used to store the computer program executed by the processor 702. The memory 701 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system; the data storage area may store various operation instruction sets, etc.
[0214] The memory 701 may be a volatile memory, such as a random - access memory (RAM); the memory 701 may also be a non - volatile memory, such as a read - only memory, a flash memory, a hard disk drive (HDD), or a solid - state drive (SSD); or the memory 701 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 701 may be a combination of the above - mentioned memories.
[0215] The processor 702 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 702 is used to implement the above - mentioned method for determining the vehicle - human interaction when calling the computer program stored in the memory 701.
[0216] The communication module 703 is used to communicate with terminal devices, site devices, or other network devices.
[0217] In the embodiments of the present application, the specific connection medium between the above - mentioned memory 701, communication module 703, and processor 702 is not limited. In the embodiments of the present application Figure 7 it is shown that the memory 701 and the processor 702 are connected through a bus 704, and the bus 704 is in Figure 7is described in thick lines, and the connection manners between other components are only for illustrative purposes and are not limiting. The bus 704 can be divided into an address bus, a data bus, a control bus, etc. For the sake of description, Figure 7 is only described by a thick line in the figure, but it does not describe that there is only one bus or one type of bus.
[0218] The memory 701 stores a computer storage medium, and the computer storage medium stores computer-executable instructions for implementing the method for determining the vehicle-person interaction method according to the embodiments of the present application. The processor 702 is configured to execute the vehicle-person interaction method according to the above embodiments.
[0219] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0220] Based on the same inventive concept, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0221] In response to receiving a work instruction, determining detection box information of a first user in a to-be-processed image; wherein, the work instruction includes activating an out-of-vehicle control suspension mode, and the detection box information includes first information of the facial area of the first user;
[0222] Performing compression processing on the first information to obtain second information, and uploading the second information;
[0223] Receiving third information corresponding to the second information; wherein, the third information includes an enhanced image of the facial area of the first user;
[0224] In response to the result of comparing the enhanced image with a pre-stored user image being passed, determining that the first user is the target interaction object for the out-of-vehicle control suspension mode;
[0225] Controlling the change of the vehicle suspension height according to the actions of the target interaction object.
[0226] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0227] In response to the similarity between the enhanced image and the pre-stored target user image being greater than a preset similarity threshold, determining that the first user is the target interaction object; wherein, the pre-stored list includes the target user image.
[0228] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0229] Obtain task description information from the work instruction; match the target image acquisition rule corresponding to the task description information from a plurality of preset image acquisition rules;
[0230] Collect the image to be processed based on the target image acquisition rule; use a target detection model to identify the first user in the image to be processed and generate the detection box information.
[0231] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0232] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the vehicle-person interaction method described in any one of the above.
[0233] Among them, the program code for executing the computer program product of the present application can be written in any combination of one or more programming languages. The program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.
[0234] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0235] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0236] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0237] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of user operation steps are performed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0238] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for vehicle-human interaction, characterized in that, Including: In response to receiving a work instruction, determining detection box information of a first user in a to-be-processed image; wherein, the work instruction includes activating an out-of-vehicle controlled suspension mode, and the detection box information includes first information of the facial area of the first user; Performing compression processing on the first information to obtain second information, and uploading the second information; Receiving third information corresponding to the second information; wherein, the third information includes an enhanced image of the facial area of the first user; In response to the comparison result between the enhanced image and a pre-stored user image being passed, determining that the first user is the target interaction object for the out-of-vehicle controlled suspension mode; Controlling the change of the vehicle suspension height according to the actions of the target interaction object.
2. The method according to claim 1, characterized in that, The determining the detection box information of the first user in the to-be-processed image includes: Obtaining task description information from the work instruction; Matching the target image acquisition rule corresponding to the task description information from a plurality of preset image acquisition rules; Collecting the to-be-processed image based on the target image acquisition rule; Using a target detection model to identify the first user in the to-be-processed image and generating the detection box information.
3. The method according to claim 2, wherein The task description information includes the physical coordinates of the first user; The collecting the to-be-processed image based on the target image acquisition rule includes: In response to the target image acquisition rule being the first rule, for the area where the physical coordinates are located, controlling a first image acquisition device to perform image acquisition according to the first working parameter in the first rule to obtain the to-be-processed image; or, In response to the target image acquisition rule being the second rule, controlling the vehicle lights to irradiate the area where the physical coordinates are located according to the first illumination parameter in the second rule, and controlling a second image acquisition device to perform image acquisition for the irradiated area according to the second working parameter in the second rule to obtain the to-be-processed image.
4. The method according to claim 1, wherein The determining the detection box information of the first user in the to-be-processed image includes: For the to-be-processed image, determining a first target detection box through a target detection model; In response to the number of the first target detection boxes being greater than 1, determining the target in the first target detection box with the largest area as the first user; Based on the first target detection box with the largest area, determining a second target detection box; wherein, the second target detection box includes the face of the first user.
5. The method according to any one of claims 1 to 4, characterized in that, The determining the detection box information of the first user in the to-be-processed image includes: For the to-be-processed image, determining a third target detection box through a target detection model; In response to the posture of the target in the third target detection box being consistent with the preset posture information, determining the target in the third target detection box as the first user; In the third target detection box, determining a fourth target detection box; wherein, the fourth target detection box includes the face of the first user.
6. The method according to claim 5, characterized in that, The determining the fourth target detection box in the third target detection box includes: In the to-be-processed image, expanding the area of the third target detection box based on a preset rule to obtain a fifth target detection box; wherein, the fifth target detection box contains the third target detection box; In the fifth target detection box, determine a sixth target detection box with a number less than a preset quantity threshold; wherein, the sixth target detection box includes the target part of the first user. Based on the category information of the target part, select the fourth target detection box from the sixth target detection boxes.
7. The method according to any one of claims 1 to 4, characterized in that, The compressing the first information to obtain second information includes: Input the facial region of the first user into a pre-trained encoder to obtain a first vector of a preset length; then the first vector is the second information.
8. The method according to claim 1, wherein The work instruction further includes activating the in-vehicle controlled suspension mode: The method further includes: in response to the suspension control mode being the in-vehicle controlled suspension mode, determining the facial region of the first user in the to-be-processed image. Compare the facial region with the pre-stored user image to determine whether the first user is the target interaction object of the in-vehicle controlled suspension mode.
9. The method according to claim 8, wherein After determining the facial region of the first user in the to-be-processed image, it further includes: Obtain the interaction mode in the work instruction; wherein, the interaction mode includes a virtual user interaction mode and / or a real user interaction mode. The comparing the facial region with the pre-stored user image to determine whether the first user is the target interaction object of the in-vehicle controlled suspension mode includes: In response to the similarity between the facial region and the user image being greater than a preset similarity threshold, determine the target interaction object of the in-vehicle controlled suspension mode based on the interaction mode.
10. A human-vehicle interaction device, characterized in that, including: A first detection module, configured to determine the detection box information of the first user in the to-be-processed image in response to receiving a work instruction; wherein, the work instruction includes activating the out-of-vehicle controlled suspension mode, and the detection box information includes first information of the facial region of the first user. A compression module, configured to compress the first information to obtain second information and upload the second information. An enhancement module, configured to receive third information corresponding to the second information; wherein, the third information includes an enhanced image of the facial region of the first user. A comparison module, configured to determine that the first user is the target interaction object of the out-of-vehicle controlled suspension mode in response to the comparison result between the enhanced image and the pre-stored user image being passed. A control module, configured to control the change of the vehicle suspension height according to the actions of the target interaction object.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 9.