Interaction method and device for unmanned vehicle to give way to pedestrian, electronic equipment and medium

By acquiring images and speeds of pedestrians, determining yielding scenarios, and adjusting interactive prompts, the problem of unclear intentions in interactions between autonomous vehicles and pedestrians is solved, improving pedestrian safety and interaction effectiveness.

CN121268901APending Publication Date: 2026-01-06CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511528880.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

The interactive prompts when existing autonomous vehicles detect pedestrians at intersections are unclear, resulting in poor pedestrian safety and failing to effectively ensure pedestrian safety.

Method used

By acquiring images and speeds of pedestrians, the system identifies scenarios for yielding to pedestrians, including thank-you and stopping scenarios. Based on the scenario and speed, it determines interactive prompts and the autonomous vehicle's movement status, enabling clear interaction intentions and real-time adjustments.

Benefits of technology

It improves pedestrian safety, ensures clear interactive prompts, responds promptly to changes in pedestrian movement, and enhances the interaction between autonomous vehicles and pedestrians.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121268901A_ABST
    Figure CN121268901A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an interaction method and device for an unmanned vehicle to give way to pedestrians, electronic equipment and a medium. Obtaining a motion image and a motion speed of a pedestrian; according to the motion image and the motion speed, determining a pedestrian commission scene; the pedestrian commission scene comprises an appreciation commission scene and a parking commission scene; and according to the pedestrian commission scene and the movement speed, determining an interaction prompt of pedestrian commission and a movement state of the unmanned vehicle. According to the embodiment of the invention, the interaction intention is clarified, and the pedestrian passing safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to unmanned vehicle technology, and more particularly to an interaction method, device, electronic device, and medium for an unmanned vehicle to yield to pedestrians. Background Technology

[0002] As intelligent driving technology evolves to higher levels, the safety of interactions between vehicles and traffic participants (pedestrians and other vehicles) has become a core technological bottleneck.

[0003] In existing technologies, when an autonomous vehicle detects a pedestrian at an intersection, it usually only issues a "caution and avoid" warning.

[0004] However, the "Caution: Avoidance" prompt is not clearly intended, has a poor interaction with pedestrians' courtesy, and cannot guarantee pedestrian safety. Summary of the Invention

[0005] This application provides an interaction method, device, electronic device, and medium for unmanned vehicles to yield to pedestrians, so as to clarify the interaction intention and improve pedestrian safety.

[0006] In a first aspect, embodiments of this application provide an interaction method for an autonomous vehicle yielding to pedestrians, the interaction method including:

[0007] Acquire motion images and speeds of pedestrians;

[0008] Based on moving images and speed of movement, determine the scenarios for yielding to pedestrians; scenarios for yielding to pedestrians include scenarios of thanking others for yielding and scenarios of stopping to yield.

[0009] Based on the pedestrian yielding scenario and the speed of movement, determine the interactive prompts for yielding to pedestrians and the movement status of the autonomous vehicle.

[0010] Secondly, embodiments of this application also provide an interactive device for an autonomous vehicle to yield to pedestrians, the interactive device for an autonomous vehicle to yield to pedestrians includes:

[0011] The motion data acquisition module is used to acquire motion images and speeds of pedestrians;

[0012] The pedestrian yielding scenario determination module is used to determine pedestrian yielding scenarios based on moving images and movement speed; pedestrian yielding scenarios include thank-you yielding scenarios and stopping yielding scenarios;

[0013] The interaction prompt determination module is used to determine the interaction prompts for yielding to pedestrians and the movement status of the autonomous vehicle based on the pedestrian yielding scenario and movement speed.

[0014] Thirdly, embodiments of this application also provide an electronic device, which includes:

[0015] One or more processors;

[0016] Storage device for storing one or more programs;

[0017] When one or more programs are executed by one or more processors, the one or more processors implement any of the interactive methods provided in the embodiments of this application for an autonomous vehicle to yield to pedestrians.

[0018] Fourthly, embodiments of this application also provide a storage medium including computer-executable instructions, which, when executed by a computer processor, are used to perform any of the interactive methods for an autonomous vehicle to yield to pedestrians as provided in embodiments of this application.

[0019] Fifthly, embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements any of the interactive methods for an autonomous vehicle to yield to pedestrians as provided in embodiments of this application.

[0020] This application acquires pedestrian motion images and speed to obtain the pedestrian's real-time status, providing a data foundation for determining pedestrian yielding scenarios. Based on the motion images and speed, it identifies pedestrian yielding scenarios, including thank-you yielding and stop-to-yield scenarios, accurately determining these two different scenarios to ensure the clarity of interactive prompts. Based on the pedestrian yielding scenario and speed, it determines the pedestrian yielding interactive prompts and the autonomous vehicle's motion state. Based on the speed, it further clarifies the interactive intent and determines the autonomous vehicle's motion state, enabling real-time adjustments to the interactive prompts and the autonomous vehicle's motion state based on the speed, responding promptly to changes in pedestrian movement and improving pedestrian safety. Therefore, the technical solution of this application solves the problem that the "pay attention and avoid" prompts lack clear intent, have poor interaction with pedestrians, and fail to guarantee pedestrian safety, achieving the effect of clear interactive intent and improved pedestrian safety. Attached Figure Description

[0021] Figure 1 This is a flowchart of an interaction method for an unmanned vehicle to yield to pedestrians, as described in Embodiment 1 of this application;

[0022] Figure 2 This is a flowchart of an interaction method for an unmanned vehicle to yield to pedestrians, as described in Embodiment 2 of this application;

[0023] Figure 3 This is a flowchart of an interaction method for an unmanned vehicle to yield to pedestrians, as described in Embodiment 3 of this application;

[0024] Figure 4 This is a schematic diagram of the structure of an interactive device for an unmanned vehicle to yield to pedestrians, as shown in Embodiment 4 of this application.

[0025] Figure 5 This is a schematic diagram of the structure of an electronic device according to Embodiment 5 of this application. Detailed Implementation

[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first" and "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] Example 1

[0029] Figure 1 This is a flowchart of an interaction method for an unmanned vehicle to yield to pedestrians, provided in Embodiment 1 of this application. This embodiment is applicable to situations where an unmanned vehicle interacts with a pedestrian when it detects a pedestrian at an intersection. The method can be executed by an interaction device for the unmanned vehicle to yield to pedestrians. This device can be implemented in software and / or hardware and is specifically configured in the unmanned vehicle.

[0030] See Figure 1 The interaction method shown for an autonomous vehicle yielding to pedestrians includes the following steps:

[0031] S110. Obtain the motion image and speed of the pedestrian.

[0032] The moving images can be images of pedestrians captured by an image acquisition device, used to determine pedestrian yielding scenarios. For example, moving images can be captured by a monocular camera (8 megapixels resolution), which can be mounted on the rearview mirror of the windshield, to capture moving images of pedestrians in front of and to the side of the autonomous vehicle as it passes through an intersection.

[0033] Movement speed can be the speed of pedestrian movement collected by devices such as speed radar, used to determine the scenario of yielding to pedestrians. For example, movement speed can be obtained by 4D millimeter-wave radar (detection range 0.5-300 meters). The 4D millimeter-wave radar can be installed in the front bumper of the vehicle and transmit pedestrian speed, distance, and azimuth data through the controller local area network bus (the recognition rate for pedestrian targets is ≥95%).

[0034] S120. Based on the moving image and the speed of movement, determine the pedestrian yielding scenario; pedestrian yielding scenarios include thank-you yielding scenarios and stopping yielding scenarios.

[0035] Moving images can be used to identify pedestrian movements and thus determine their crossing intentions. For example, moving images can be used to identify pedestrians' stillness, walking, waving, and hesitation, as well as their direction of movement, and combined with their speed, to determine the scenario for yielding to pedestrians.

[0036] The pedestrian yielding scenario can refer to an autonomous vehicle yielding to pedestrians at an intersection for those with different crossing intentions and speeds. For example, pedestrian yielding scenarios can include thank-you yielding scenarios and stop-to-yield scenarios. A thank-you yielding scenario involves the autonomous vehicle expressing gratitude to the pedestrian for yielding, in which the vehicle proceeds through the intersection and thanks the pedestrian for yielding. A stop-to-yield scenario involves the autonomous vehicle stopping and prompting the pedestrian to proceed.

[0037] For example, if a pedestrian's direction of movement is perpendicular to the direction of movement of the autonomous vehicle, and the pedestrian waves to signal the autonomous vehicle to pass, and the pedestrian's movement speed is slow or stationary, then the pedestrian yielding scenario can be determined as a thank-you yielding scenario; if the pedestrian is walking or hesitating, and the movement speed is fast, then the pedestrian yielding scenario can be determined as a stop yielding scenario.

[0038] S130. Determine the pedestrian yielding interaction prompt and the autonomous vehicle's motion status based on at least one of the pedestrian yielding scenario and the vehicle's speed.

[0039] Interactive prompts can be messages indicating to pedestrians that they are crossing or expressing gratitude. These prompts facilitate intentional interaction with pedestrians, preventing situations where autonomous vehicles and pedestrians simultaneously cross or wait at intersections, thus improving intersection efficiency and safety, and enhancing the clarity of the prompts. For example, interactive prompts can include at least one form, such as voice or text, without specific limitation. In the same pedestrian yielding scenario, different interactive prompts can be determined based on the pedestrian's different movement speeds, providing corresponding interactive prompts tailored to the user's movement state, improving the intelligence of the interaction, making it easier for pedestrians to understand the prompts, and enhancing the user-friendliness of the interaction.

[0040] In the "Thank you for yielding" scenario, the interaction prompt for yielding to pedestrians can be identified as a "thank you" message, and the autonomous vehicle's movement status can be determined as "passing." In the "stop and yield" scenario, the interaction prompt for yielding to pedestrians can be identified as a "pedestrian crossing prompt," and different pedestrian crossing prompts can be determined based on the vehicle's speed to guide the pedestrian to cross, while the autonomous vehicle's movement status is determined as "stopped." For example, when the pedestrian's speed is slow, the interaction prompt could be, "Please cross, I will stop"; when the pedestrian's speed is fast, the interaction prompt could be, "Please cross safely, please be careful, I will wait."

[0041] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant regions.

[0042] As intelligent driving technology evolves from L2 (a technical term, a level of autonomous driving) to L3 (a technical term, a level of autonomous driving) and higher, the safety of interactions between vehicles and traffic participants (pedestrians and other vehicles) has become a core technological bottleneck. A large number of intelligent driving interaction conflict accidents stem from "misjudgment of intent," meaning that the vehicle has performed a yielding maneuver, but the pedestrian or other vehicle, not having clearly received the signal, takes a synchronized action (such as a vehicle suddenly starting when a pedestrian hesitates to cross the street, or two vehicles simultaneously decelerating and then simultaneously accelerating).

[0043] In terms of conveying interactive intent, existing technologies mainly rely on two types of methods: passive signal transmission and active signal transmission. Passive signal transmission indirectly expresses intent through the vehicle's own movement (such as deceleration, stopping, or flashing turn signals). For example, automatically decelerating when a pedestrian is detected and alerting surrounding road users by illuminating the brake lights. Active signal transmission uses an external display screen to statically display text such as "Yielding," but lacks dynamic adjustment capabilities and offers poor text interaction.

[0044] The technical solution of this embodiment acquires the pedestrian's motion image and speed to obtain the pedestrian's real-time status, providing a data foundation for determining the pedestrian yielding scenario. Based on the motion image and speed, the pedestrian yielding scenario is determined; this scenario includes a thank-you yielding scenario and a stop-to-yield scenario, accurately identifying these two different scenarios to ensure the clarity of the interactive prompts. Based on the pedestrian yielding scenario and speed, the interactive prompts for yielding and the autonomous vehicle's motion status are determined. Based on the speed, the interactive intent and the autonomous vehicle's motion status are further clarified, enabling real-time adjustments to the interactive prompts and the autonomous vehicle's motion status according to the speed, responding promptly to changes in pedestrian movement and improving pedestrian safety. Therefore, the technical solution of this application solves the problem that the "pay attention and avoid" prompt has an unclear intent, poor interaction with pedestrians, and fails to guarantee pedestrian safety, achieving the effect of clear interactive intent and improved pedestrian safety.

[0045] Example 2

[0046] Figure 2 This is a flowchart of an interaction method for an unmanned vehicle to yield to pedestrians, provided in Embodiment 2 of this application. The technical solution of this embodiment is further refined based on the above technical solution.

[0047] Furthermore, the phrase "determine the pedestrian yielding interaction prompts and autonomous vehicle movement status based on the pedestrian yielding scenario and movement speed" is further refined into: "If the pedestrian yielding scenario is a stop yielding scenario, then determine the avoidance prompt words including the autonomous vehicle's intention to stop based on the movement speed, and determine the autonomous vehicle's movement status as stopped," so as to determine the avoidance prompt words and autonomous vehicle movement status in the stop yielding scenario.

[0048] See Figure 2 The interactive method shown includes:

[0049] S210. Obtain the motion image and speed of the pedestrian.

[0050] S220. Based on the moving image and the speed of movement, determine the pedestrian yielding scenario; pedestrian yielding scenarios include thank-you yielding scenarios and stopping yielding scenarios.

[0051] S230. If the pedestrian yielding scenario is a parking yielding scenario, then determine the yielding prompt word including the autonomous vehicle's intention to stop based on the movement speed, and determine that the autonomous vehicle's movement state is stopped.

[0052] Escape prompts can be words used in parking and yielding scenarios to indicate that the autonomous vehicle is stopping to yield to pedestrians, based on its speed. These prompts clarify the autonomous vehicle's stopping intention. If the pedestrian yielding scenario is a parking and yielding scenario, the autonomous vehicle's movement state is determined to be stopped, yielding to pedestrians and allowing them to cross first. In parking and yielding scenarios, based on the pedestrian's speed, corresponding escape prompts indicating the autonomous vehicle's stopping intention are determined, improving the intelligence and semantic clarity of the prompts and enhancing the interactive experience.

[0053] Optionally, if the pedestrian yielding scenario is a parking yielding scenario, then the scenario label is determined based on the movement speed, and the yielding prompt words including the autonomous vehicle's intention to stop are determined based on the scenario label.

[0054] For example, if a pedestrian's movement speed is low, the scene label is determined to be the pedestrian in a hesitant state, and the avoidance prompt could be: "Please cross the street with confidence, I will stop and wait." For example, a low movement speed range could be 0.1-0.2 m / s (a technical term, a unit of speed, meters per second). For example, if a pedestrian's movement speed is moderate, the scene label is determined to be the pedestrian in a slow, tentative crossing state, and the avoidance prompt could be: "Please quicken your pace and cross, I will stop and wait." For example, a moderate movement speed range could be 0.2-0.5 m / s. For example, if a pedestrian's movement speed is high, the scene label is determined to be the pedestrian in a fast crossing state, and the avoidance prompt could be: "Please cross carefully, I have stopped." For example, a high movement speed range could be greater than 0.5 m / s. Avoidance prompts can be communicated to pedestrians through text and voice. Since pedestrians may not be able to easily see the text displayed on the autonomous vehicle while crossing, voice communication improves the real-time nature and effectiveness of the interaction, ensuring that users receive avoidance prompts promptly.

[0055] In an optional embodiment, after determining the pedestrian yielding interaction prompt and the autonomous vehicle's movement status based on the pedestrian yielding scenario and movement speed, the method further includes: if the interaction prompt is a voice prompt, determining the volume of the voice playback based on the distance between the pedestrian and the autonomous vehicle.

[0056] When the interactive prompts are voice prompts, the volume of the voice playback needs to be determined based on the distance between the pedestrian and the autonomous vehicle, taking into account the attenuation of sound during propagation. This ensures that the pedestrian can clearly hear the interactive prompts without feeling uncomfortable due to excessive volume.

[0057] Specifically, the voice prompts are emitted through a speaker installed inside the front grille of the autonomous vehicle, featuring an omnidirectional sound design to ensure pedestrians within a 5-10m range can clearly receive the voice message. The speaker output receives the audio signal of the interactive prompts and the volume control signal (the base volume can be 60 decibels) through an audio power amplifier; the volume is dynamically adjusted according to the pedestrian distance to ensure that pedestrians in different positions can accurately receive the courtesy information from the interactive prompts; and there is at least a 3-second interval between two voice prompts. For example, the dynamic volume adjustment based on pedestrian distance could be that the volume increases by 3 decibels for every 5 meters of distance, up to a maximum of 70 decibels.

[0058] If the interactive prompt is a voice prompt, the volume of the voice playback is determined based on the distance between the pedestrian and the autonomous vehicle to ensure that the pedestrian can hear the interactive prompt clearly, ensuring the effectiveness of the voice prompt, and ensuring that the voice prompt is user-friendly without causing discomfort to the pedestrian due to excessive volume.

[0059] Optionally, if the pedestrian yielding scenario is a stop yielding scenario, then after determining the avoidance prompt word including the autonomous vehicle's intention to stop based on the movement speed, and after determining that the autonomous vehicle's movement state is stopped, it also includes: detecting the lateral distance between the pedestrian and the autonomous vehicle, and after the lateral distance is greater than the preset starting distance, determining that the autonomous vehicle's movement state is to pass.

[0060] The preset starting distance can be the minimum lateral distance between the pedestrian and the autonomous vehicle when the vehicle starts moving after yielding to pedestrians. This is used to determine the moment when the autonomous vehicle's movement state changes from stationary to proceeding. For example, the preset starting distance can be 1.5 meters. When the lateral distance exceeds the preset starting distance, the autonomous vehicle's movement state is determined to proceed, avoiding collisions with pedestrians due to premature starting.

[0061] The following example, using a crosswalk without traffic lights, illustrates the interaction process of an autonomous vehicle yielding to pedestrians in a stop-and-yield scenario:

[0062] The scene is characterized as follows: zebra crossings on secondary urban roads without traffic lights, frequent pedestrian crossings during morning and evening rush hours (such as intersections around vegetable markets), and pedestrians are mainly elderly people and shoppers, who move slowly and are prone to hesitation.

[0063] Interaction flow:

[0064] When the autonomous vehicle approaches the intersection (15 meters away from the intersection), the monocular camera and millimeter-wave radar identify three pedestrians at the zebra crossing at a speed of 0.15 m / s. The autonomous vehicle determines that the pedestrian yielding scenario is a stop yielding scenario, the scenario label is "pedestrians hesitate at the intersection", the yielding prompt is "please cross the street with peace of mind, I will stop and wait", and the autonomous vehicle's movement status is determined to be stopped.

[0065] The vehicle's external speakers are automatically activated, playing the voice message: "Please cross the street with confidence, I will wait for you," at a volume of 69 decibels.

[0066] One pedestrian was detected to begin walking (at a speed of 0.3 m / s). The scene was identified as "pedestrian slowly and tentatively crossing the street". The avoidance prompt was "Please hurry across, I am waiting". The voice message was played: "Please hurry across, I am waiting". The vehicle remained stationary.

[0067] After all three pedestrians have crossed (the minimum lateral distance between the pedestrians and the driverless car is greater than or equal to 1.5 meters), the driverless car plans its starting path, determines its movement state as "passing", and slowly passes through the intersection.

[0068] In an optional embodiment, the interaction prompt for yielding to pedestrians and the movement status of the autonomous vehicle are determined based on the pedestrian yielding scenario and the movement speed. The method further includes: if the pedestrian yielding scenario is a thank-you scenario, then a thank-you prompt word including the autonomous vehicle's intention to pass is determined based on the movement speed, and the movement status of the autonomous vehicle is determined to be "passing".

[0069] The "thank you" prompt can be a word of gratitude for pedestrians waiting for the autonomous vehicle to pass, determined based on the vehicle's speed in a "thank you for yielding" scenario. This clarifies the autonomous vehicle's intention to proceed. If the pedestrian yielding scenario is a "thank you for yielding" scenario, then the autonomous vehicle's movement state is determined to be "progress." In a "thank you for yielding" scenario, the corresponding yield prompt is determined based on the pedestrian's speed, improving the intelligence and semantic clarity of the "thank you" prompt and enhancing the interactive experience.

[0070] Optionally, if the pedestrian yielding scenario is a thank-you scenario, then the scenario label is determined based on the movement speed, and the avoidance prompt word is determined based on the scenario label.

[0071] For example, when the pedestrian's speed is 0, the scene label is set to "Waiting for the autonomous vehicle to pass" and the thank-you message is set to "Thank you for yielding, I will pass quickly," clearly indicating the pedestrian's intention to pass. When the pedestrian's speed is less than 0.1 m / s and decelerates, the scene label is set to "Yielding to the autonomous vehicle to pass" and the thank-you message is set to "Thank you for yielding, I will pass slowly," clearly indicating the pedestrian's intention to pass, while also indicating that the autonomous vehicle will pass slowly without urging the pedestrian to slow down, thus improving the user-friendliness of the interaction.

[0072] If the scenario of yielding to pedestrians is a thank-you scenario, then the system determines the thank-you prompt message based on the vehicle's speed to clearly indicate the autonomous vehicle's intention to pass, and sets the vehicle's movement state to "passing". When a pedestrian yields to the autonomous vehicle, the system expresses gratitude to the pedestrian, clearly indicating the vehicle's intention to pass. This improves the efficiency of autonomous vehicle passage, avoids excessive yielding that causes traffic congestion, and enhances the intelligence of the autonomous vehicle.

[0073] In existing technologies, the expression of intent is ambiguous, easily leading to secondary conflicts: the fixed voice prompt "Please give way" does not clearly distinguish between "vehicles yielding" and "requesting the other party to give way," which can easily cause misunderstandings in complex scenarios such as intersections. For example, pedestrians may misjudge that a vehicle is about to pass and stop, creating a stalemate of "vehicles waiting - pedestrians hesitating." The main reason for secondary conflicts is the lack of context-specific wording design, the failure to dynamically adjust the expression based on the relative position of pedestrians, low traffic efficiency, and the tendency to create traffic congestion.

[0074] The technical solution of this embodiment determines the avoidance prompt word including the autonomous vehicle's intention to stop based on the movement speed if the pedestrian yielding scenario is a parking yielding scenario, and determines that the autonomous vehicle's movement state is stopped. In the parking yielding scenario, the parking intention of the autonomous vehicle is clearly defined, improving the traffic efficiency of the intersection. By obtaining the movement speed, the corresponding avoidance prompt word is determined in real time, and timely responses are made to changes in pedestrian speed, improving the intelligence and flexibility of yielding.

[0075] Example 3

[0076] Figure 3 This is a flowchart of an interaction method for an unmanned vehicle to yield to pedestrians, provided in Embodiment 3 of this application. The technical solution of this embodiment is further refined based on the above technical solution.

[0077] Furthermore, the process of "determining the pedestrian yielding scenario based on the moving image and the speed of movement" is further refined into: "identifying the pedestrian's orientation based on the moving image and determining the pedestrian's intention to pass based on the pedestrian's orientation; if the intention to pass is to pass, then identifying the pedestrian's actions based on the moving image; and determining the pedestrian yielding scenario based on the pedestrian's actions and the speed of movement," in order to accurately determine the pedestrian yielding scenario.

[0078] See Figure 3 The interactive method shown includes:

[0079] S310. Acquire the motion image and speed of the pedestrian.

[0080] S320. Based on the moving image, identify the pedestrian's orientation and determine the pedestrian's intention to pass based on the pedestrian's orientation.

[0081] Pedestrian orientation can be the direction in which a pedestrian's body is facing, used to determine the intention to cross. The intention to cross can be the pedestrian's intention to cross a sidewalk or intersection perpendicular to the direction of travel of the autonomous vehicle. For example, the intention to cross can include both having the intention to cross and not having the intention to cross.

[0082] For example, a deep learning network can be used to identify the direction a pedestrian is facing based on motion images. When the pedestrian's direction is perpendicular to the autonomous vehicle's direction of movement, the intention to pass is determined to be "yes"; when the pedestrian's direction is parallel to the autonomous vehicle's direction of movement, the intention to pass is determined to be "no".

[0083] S330. If the intention to pass is to pass, then the pedestrian's actions are identified based on the motion image.

[0084] Pedestrian actions can include standing still, walking, waving, and hesitating, which are used to determine scenarios for yielding to pedestrians. For example, pedestrian actions can be identified from motion images using a deep learning model.

[0085] To address the needs of pedestrian body orientation and action recognition, a recognition scheme with 15 key points (covering key parts such as head, torso, and limbs) can be designed, which simultaneously realizes key point localization, body orientation judgment, and action recognition through a multi-branch network.

[0086] The key point system design is as follows:

[0087] Based on the characteristics of human kinematics, 15 representative key points were selected and divided into three categories:

[0088] Head regions (3): vertex of the head, tip of the nose, and midpoint of the chin;

[0089] Trunk areas (5): neck point, left shoulder, right shoulder, left hip, and right hip;

[0090] Limb regions (7): left elbow, right elbow, left wrist, right wrist, left knee, right knee, and left ankle (prioritize identifying key points on one side of the lower limb to reduce computation, and infer the other side through symmetry).

[0091] The detailed design of the network structure is as follows: input layer design and preprocessing, backbone network design, multi-scale feature fusion module design, key point detection branch design, body orientation recognition branch design, action recognition branch design, and output post-processing design.

[0092] The input layer design and preprocessing are as follows:

[0093] Input specifications: The motion image is a 3-channel RGB (technical term, an image format) image with a resolution of 512×512 pixels. If the motion image does not meet the size requirements, the longer side will be scaled to 512 pixels while maintaining the original image ratio, and the remaining pixels will be padded with 0.

[0094] The preprocessing steps are as follows:

[0095] Pixel value normalization: Converts the range [0,255] to [-1,1];

[0096] Geometric transformations: random scaling (0.8-1.2 times) and random rotation (±15 degrees);

[0097] Light enhancement: Randomly adjusts brightness (±30%) and contrast (±20%).

[0098] Target cropping: The image is cropped using pedestrian detection boxes to reduce background interference.

[0099] The backbone network can be designed using EfficientNet-B0 (a technical term, referring to a network model). Choosing the lightweight EfficientNet-B0 as the basis for feature extraction offers advantages such as fewer parameters (approximately 5.3 megabytes) and strong feature representation capabilities. The backbone network design includes the following:

[0100] Main structure: It consists of 16 "Inverted Bottleneck" modules, including: depthwise separable convolutions, squeeze activation modules, sliding window convolutions, and output features. Using depthwise separable convolutions reduces computational complexity. Squeeze activation modules enhance the weights of key feature channels. Sliding window convolutions expand the receptive field.

[0101] The output features include four feature maps at different scales (C1-C4), with the following scales for each feature map:

[0102] C1: 128×128×40 (step size 4);

[0103] C2: 64×64×112 (step size 8);

[0104] C3: 32×32×320 (step size 16);

[0105] C4: 16×16×1280 (step size 32).

[0106] The multi-scale feature fusion module design can employ an improved Feature Pyramid Network (FPN) structure to fuse multi-scale features. The Feature Pyramid Network is a feature pyramid network used to solve the multi-scale object detection problem in computer vision, and includes the following structural features:

[0107] Top-down fusion: C4 is reduced to 256 channels through 1×1 convolution, upsampled to 32×32 and fused with C3; the fusion result is then upsampled to 64×64 and fused with C2, finally obtaining a 64×64×256 fused feature map F;

[0108] Feature enhancement: An attention mechanism is added after fusing feature map F to enhance the features of human body contour and key point regions.

[0109] The key point detection branch is designed as follows:

[0110] Network structure:

[0111] Input layer: fused feature map F (64×64×256);

[0112] Convolutional layer: 3 3×3 convolutional blocks (dimensions from 256 to 128 to 64), each convolution is followed by a batch normalization layer and a Swish activation function;

[0113] Output layer: 15 heatmaps (64×64×15) and 15 offset maps (64×64×30);

[0114] Heatmap output: Each key point corresponds to a heatmap, and the peak position indicates the coordinates of the key point. The higher the value, the greater the confidence level.

[0115] Offset map output: Each keypoint contains offsets in the x and y directions, used to correct coordinate quantization errors.

[0116] Loss functions include heatmap loss function and offset loss function.

[0117] For example, the heatmap loss can use the Focal Loss function, and the offset loss can use the L1 loss function.

[0118] The body orientation recognition branch design is as follows:

[0119] Input features: A 256-dimensional vector obtained by global average pooling of the fused feature map F and the relative coordinates of 15 key points (normalized);

[0120] The network structure is as follows:

[0121] Fully connected layer 1: 256 dimensions reduced to 128 dimensions, including 30 coordinate features;

[0122] Fully connected layer 2: 128 dimensions reduced to 64 dimensions, with BatchNorm layer and Dropout layer;

[0123] Output layer: probability distribution of 4 orientations (0°, 90°, 180° and 270°);

[0124] Loss function: Cross-entropy loss with class weights, which solves the class imbalance problem.

[0125] The action recognition branch design is as follows:

[0126] Input features: fused feature maps of 3 consecutive frames (temporal features extracted through 3D convolution);

[0127] Key point motion vectors (displacement changes of 15 key points);

[0128] The network structure is as follows:

[0129] 3D convolutional layer: 3×3×3 convolutional kernel, extracting temporal-spatial features;

[0130] LSTM layer (technical term, a network structure): 2-layer bidirectional LSTM, capturing action temporal dependencies;

[0131] Output layer: Probability distribution of 5 types of actions (stationary, walking, waving, hand-waving, and hesitating);

[0132] Loss function: Cross-entropy loss.

[0133] The output post-processing design is as follows:

[0134] Key point coordinate calculation: Non-maximum suppression (NMS) is applied to the heatmap to extract peak points; precise coordinates are calculated by combining the offset.

[0135] Actual coordinates = (heatmap peak coordinates × step size) + offset;

[0136] Coordinate transformation: Converting image coordinates to world coordinates (based on camera intrinsics and distance information);

[0137] Body orientation calculation:

[0138] Based on the fusion of shoulder-hip line vector and head orientation vector, the classification results are smoothed (sliding window averaging).

[0139] Action recognition result optimization: Combine key point motion speed and acceleration features to filter out misjudgments, and adopt a temporal voting mechanism. The temporal voting mechanism can confirm the action after at least 2 out of 3 consecutive frames are consistent.

[0140] S340. Determine the scenario for yielding to pedestrians based on their movements and speed.

[0141] When a pedestrian is stationary, waving, or moving at a low speed, it can be assumed that the pedestrian is signaling the driverless vehicle to go first, and the pedestrian yielding scenario is defined as a thank-you yielding scenario. When a pedestrian is walking or hesitating, and moving at a high speed, it can be assumed that the pedestrian intends to go first, and the pedestrian yielding scenario is defined as a stop yielding scenario.

[0142] In one optional embodiment, pedestrian actions include: remaining still, waving, waving, walking, and hesitating. Accordingly, based on the pedestrian actions and movement speed, a pedestrian yielding scenario is determined, including: if the pedestrian action is remaining still, waving, or waving, and the movement speed is less than the avoidance threshold, then the pedestrian yielding scenario is determined to be a thank-you yielding scenario; if the pedestrian action is walking or hesitating, and the movement speed is not less than the avoidance threshold, then the pedestrian yielding scenario is determined to be a stop yielding scenario.

[0143] The yielding threshold can be a speed threshold for determining the pedestrian yielding scenario. The yielding threshold can be determined by professional technicians based on experience or experimentation; this application does not impose specific limitations on it. Pedestrian actions include: remaining still, waving, waving, walking, and hesitating. Waving or waving can be considered a signal to the autonomous vehicle to pass first. Therefore, if the pedestrian's action is still, waving, or waving, and the speed is less than the yielding threshold, the pedestrian yielding scenario is determined to be a thank-you yielding scenario. If the pedestrian's action is walking or hesitating, and the speed is not less than the yielding threshold, the pedestrian can be considered to be proceeding. In this case, to ensure pedestrian safety, the pedestrian yielding scenario is determined to be a stop yielding scenario.

[0144] If a pedestrian's actions are stationary, waving, or swinging, and their speed is less than the avoidance threshold, then the pedestrian yielding scenario is determined to be a thank-you yielding scenario, responding to the pedestrian's gesture of yielding and improving traffic efficiency; if a pedestrian's actions are walking or hesitating, and their speed is not less than the avoidance threshold, then the pedestrian yielding scenario is determined to be a stop yielding scenario, promptly yielding to pedestrians when they attempt to cross, ensuring their safety.

[0145] S350 determines the pedestrian yielding interaction prompts and the autonomous vehicle's movement status based on the pedestrian yielding scenario and movement speed.

[0146] The technical solution of this embodiment identifies the pedestrian's direction based on the motion image and determines the pedestrian's intention to pass based on the pedestrian's direction. If the intention to pass is positive, the pedestrian's actions are identified based on the motion image. By accurately identifying the pedestrian's actions, the accuracy of subsequent pedestrian yielding scenarios is improved. The pedestrian yielding scenario is determined based on the pedestrian's actions and speed. The current pedestrian yielding scenario can be determined in real time based on the pedestrian's actions and speed, and the pedestrian yielding scenario can be adjusted in a timely manner according to changes in the pedestrian. This ensures that the autonomous vehicle can determine the corresponding pedestrian yielding scenario in real time according to changes in the pedestrian, improving the autonomous vehicle's timely response to changes in the pedestrian.

[0147] Example 4

[0148] Figure 4The diagram shown is a schematic representation of an interactive device for an autonomous vehicle yielding to pedestrians, provided in Embodiment 4 of this application. This embodiment is applicable to situations where an autonomous vehicle interacts with a pedestrian when it detects one at an intersection. The specific structure of this interactive device for yielding to pedestrians is as follows:

[0149] The motion data acquisition module 410 is used to acquire motion images and motion speed of pedestrians;

[0150] The pedestrian yielding scenario determination module 420 is used to determine pedestrian yielding scenarios based on moving images and movement speed; pedestrian yielding scenarios include thank-you yielding scenarios and stopping yielding scenarios;

[0151] The interaction prompt determination module 430 is used to determine the interaction prompt for yielding to pedestrians and the movement status of the autonomous vehicle based on the pedestrian yielding scenario and movement speed.

[0152] The technical solution of this embodiment acquires the pedestrian's motion image and speed to obtain the pedestrian's real-time status, providing a data foundation for determining the pedestrian yielding scenario. Based on the motion image and speed, the pedestrian yielding scenario is determined; this scenario includes a thank-you yielding scenario and a stop-to-yield scenario, accurately identifying these two different scenarios to ensure the clarity of the interactive prompts. Based on the pedestrian yielding scenario and speed, the interactive prompts for yielding and the autonomous vehicle's motion status are determined. Based on the speed, the interactive intent and the autonomous vehicle's motion status are further clarified, enabling real-time adjustments to the interactive prompts and the autonomous vehicle's motion status according to the speed, responding promptly to changes in pedestrian movement and improving pedestrian safety. Therefore, the technical solution of this application solves the problem that the "pay attention and avoid" prompt has an unclear intent, poor interaction with pedestrians, and fails to guarantee pedestrian safety, achieving the effect of clear interactive intent and improved pedestrian safety.

[0153] Optionally, the interactive prompt confirmation module 430 includes:

[0154] The yielding prompt word determination unit is used to determine the yielding prompt word, including the autonomous vehicle's intention to stop, based on the vehicle's movement speed if the yielding to pedestrians scenario is a stopping yielding scenario, and to determine that the autonomous vehicle's movement state is stopped.

[0155] Optionally, the interactive prompt confirmation module 430 also includes:

[0156] The thank-you prompt word determination unit is used to determine the thank-you prompt word, including the autonomous vehicle's intention to pass, based on the vehicle's speed if the pedestrian yielding scenario is a thank-you yielding scenario, and to determine the autonomous vehicle's movement state as passing.

[0157] Optional, the pedestrian yielding scenario determination module 420 includes:

[0158] The intent determination unit is used to identify the direction of a pedestrian based on a moving image, and to determine the pedestrian's intention to pass based on the direction of the pedestrian.

[0159] The pedestrian action recognition unit is used to recognize pedestrian actions based on motion images if the pedestrian's intention to pass is to pass.

[0160] The pedestrian yielding scenario determination unit is used to determine the pedestrian yielding scenario based on the pedestrian's actions and movement speed.

[0161] Optional pedestrian actions include: remaining still, waving, waving, walking, and hesitating;

[0162] Correspondingly, the pedestrian yielding scenario determination unit includes:

[0163] The "Thank You for Yielding" scenario determination subunit is used to determine the pedestrian yielding scenario as a "Thank You for Yielding" scenario if the pedestrian's action is stationary, waving, or waving, and the movement speed is less than the avoidance threshold.

[0164] The parking yielding scenario determination subunit is used to determine the pedestrian yielding scenario as a parking yielding scenario if the pedestrian's action is walking or hesitating, and the movement speed is not less than the avoidance threshold.

[0165] Optional interactive devices for autonomous vehicles yielding to pedestrians also include:

[0166] The volume determination module is used to determine the volume of the voice prompt based on the distance between the pedestrian and the autonomous vehicle if the interactive prompt is a voice prompt.

[0167] The interactive device for unmanned vehicles yielding to pedestrians provided in this application embodiment can execute the interactive method for unmanned vehicles yielding to pedestrians provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the interactive method for unmanned vehicles yielding to pedestrians.

[0168] According to embodiments of the present invention, the present invention also provides an electronic device, a readable storage medium, and a computer program product.

[0169] Example 5

[0170] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this application, as shown below. Figure 5 As shown, the electronic device includes a processor 510, a memory 520, an input device 530, and an output device 540; the number of processors 510 in the electronic device can be one or more. Figure 5 Taking a processor 510 as an example; the processor 510, memory 520, input device 530, and output device 540 in the electronic device can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0171] The memory 520, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the interaction method of the unmanned vehicle yielding to pedestrians in the embodiments of this application (e.g., motion data acquisition module 410, pedestrian yielding scenario determination module 420, and interaction prompt determination module 430). The processor 510 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 520, thereby implementing the aforementioned interaction method of the unmanned vehicle yielding to pedestrians.

[0172] The memory 520 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 520 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 520 may further include memory remotely located relative to the processor 510, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0173] Input device 530 can be used to receive input character information and generate key signal inputs related to user settings and function control of the electronic device. Output device 540 may include display devices such as a display screen.

[0174] Example 6

[0175] Embodiment Six of this application also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform an interactive method for an autonomous vehicle to yield to pedestrians. The method includes: acquiring a motion image and motion speed of a pedestrian; determining a yielding scenario based on the motion image and motion speed; the yielding scenario includes a thank-you yielding scenario and a stop yielding scenario; and determining an interactive prompt for yielding to pedestrians and the motion state of the autonomous vehicle based on the yielding scenario and motion speed.

[0176] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the method operations described above, but can also execute related operations in the interactive method of unmanned vehicles yielding to pedestrians provided in any embodiment of this application.

[0177] Based on the above description of the implementation methods, those skilled in the art can clearly understand that this application can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0178] It is worth noting that in the above-mentioned embodiment of the interactive device for autonomous vehicles yielding to pedestrians, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this application.

[0179] Note that the above are merely preferred embodiments and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the appended claims.

Claims

1. An interaction method for an autonomous vehicle yielding to pedestrians, characterized in that, The method comprises: acquiring a motion image and a motion speed of a pedestrian; determining a courtesy pedestrian scenario according to the motion image and the motion speed; the courtesy pedestrian scenario comprises a thank-you courtesy scenario and a stop courtesy scenario; determining an interaction prompt for the courtesy pedestrian and a motion state of the unmanned vehicle according to the courtesy pedestrian scenario and the motion speed.

2. The method of claim 1, wherein, The determining of the interaction prompt for the courtesy pedestrian and the motion state of the unmanned vehicle according to the courtesy pedestrian scenario and the motion speed comprises: if the courtesy pedestrian scenario is the stop courtesy scenario, determining an avoidance prompt word comprising a stop intention of the unmanned vehicle according to the motion speed, and determining the motion state of the unmanned vehicle as stopping.

3. The method of claim 1, wherein, The determining of the interaction prompt for the courtesy pedestrian and the motion state of the unmanned vehicle according to the courtesy pedestrian scenario and the motion speed further comprises: if the courtesy pedestrian scenario is the thank-you courtesy scenario, determining a thank-you prompt word comprising a passing intention of the unmanned vehicle according to the motion speed, and determining the motion state of the unmanned vehicle as passing.

4. The method of claim 1, wherein, The determining of the courtesy pedestrian scenario according to the motion image and the motion speed comprises: recognizing a pedestrian direction according to the motion image, and determining a passing intention according to the pedestrian direction; if the passing intention is a passing intention, recognizing a pedestrian action according to the motion image; determining the courtesy pedestrian scenario according to the pedestrian action and the motion speed.

5. The method of claim 4, wherein, The pedestrian action comprises: standing still, waving a hand, waving, walking and hesitating. Correspondingly, the determining of the courtesy pedestrian scenario according to the pedestrian action and the motion speed comprises: if the pedestrian action is standing still, waving a hand or waving, and the motion speed is less than an avoidance threshold, determining the courtesy pedestrian scenario as the thank-you courtesy scenario; if the pedestrian action is walking or hesitating, and the motion speed is not less than the avoidance threshold, determining the courtesy pedestrian scenario as the stop courtesy scenario.

6. The method of claim 1, wherein, After the determining of the interaction prompt for the courtesy pedestrian and the motion state of the unmanned vehicle according to the courtesy pedestrian scenario and the motion speed, the method further comprises: if the interaction prompt is a voice prompt, determining a volume of voice playing according to a distance between the pedestrian and the unmanned vehicle.

7. An interactive device for an autonomous vehicle to yield to pedestrians, characterized in that, The method comprises: a motion data acquisition module, configured to acquire a motion image and a motion speed of a pedestrian; a courtesy pedestrian scenario determination module, configured to determine a courtesy pedestrian scenario according to the motion image and the motion speed; the courtesy pedestrian scenario comprises a thank-you courtesy scenario and a stop courtesy scenario; an interaction prompt determination module, configured to determine an interaction prompt for the courtesy pedestrian and a motion state of the unmanned vehicle according to the courtesy pedestrian scenario and the motion speed.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the interaction method when the unmanned vehicle is courteous to a pedestrian, as claimed in any one of claims 1-6.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the interaction method when the unmanned vehicle is courteous to a pedestrian, as claimed in any one of claims 1-6.

10. A computer program product, characterised in that, The computer program product comprises a computer program, which, when executed by a processor, implements the interaction method when the unmanned vehicle is courteous to a pedestrian, as claimed in any one of claims 1-6.