Parking method, device, parking equipment and storage medium

CN122808707APending Publication Date: 2026-09-25ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610974034.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]由于阴影与低矮障碍物如路沿、石块、轮胎具有高度相似的边缘纹理和暗区形态,基于可见光图像的感知算法难以有效区分二者,导致自动泊车时频繁产生误报,严重影响泊车的流畅性

Benefits of technology

[0010]在本申请实施例中,通过控制发光单元工作于目标编码组合模式,向车辆的周围环境发射探测光,并控制采集单元对周围环境进行图像采集,得到图像序列,通过控制发光单元发射具有时空编码特征的探测光,使得投射到周围环境的光不再是均匀的环境光,而是携带特定时空信息的主动光,并对图像序列进行时空解码,得到周围环境对应的目标物理特征,通过目标物理特征从物理反射特性上区分出阴影区域与实体物体,进一步地根据目标物理特征生成泊车感知信息时滤除阴影区域的干扰,从而有效避免由阴影区域产生的误报,提高泊车的流畅性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122808707A_ABST
    Figure CN122808707A_ABST
Patent Text Reader

Abstract

The application discloses a parking method and device, a parking equipment and a storage medium, and belongs to the technical field of vehicles. The method comprises the following steps: controlling a light emitting unit to work in a target coding combination mode, emitting detection light to the surrounding environment of a vehicle, and controlling an acquisition unit to collect images of the surrounding environment to obtain an image sequence; the target coding combination mode encodes light in a time sequence coding and spatial coding mode, and space-time decoding is performed on the image sequence to obtain target physical features corresponding to the surrounding environment; the target physical features support distinguishing shadow areas from solid objects in the surrounding environment; parking perception information of the surrounding environment from which the shadow areas are filtered out is generated according to the target physical features, and the vehicle is controlled to park according to the parking perception information. The application can improve the fluency of parking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of vehicle technology, and in particular relates to a parking method, device, parking equipment and storage medium. Background Technology

[0002] With the rapid development of intelligent driving technology, automatic parking has become one of the core functions of intelligent connected vehicles. At night or in low-light conditions, automatic parking relies primarily on the coordinated operation of visible light cameras and ultrasonic radar to perceive the surrounding environment. However, urban road environments contain complex artificial light sources such as streetlights and vehicle headlights, which create strong contrasts and dynamic shadow areas on surfaces such as the ground and walls.

[0003] Because shadows and low obstacles such as curbs, stones, and tires have highly similar edge textures and dark area shapes, perception algorithms based on visible light images have difficulty effectively distinguishing between them, leading to frequent false alarms during automatic parking and seriously affecting the smoothness of parking. Summary of the Invention

[0004] This application provides a parking method, apparatus, parking equipment, and storage medium that can improve parking efficiency.

[0005] In a first aspect, embodiments of this application provide a parking method, the method comprising: The light-emitting unit is controlled to operate in the target coding combination mode, emitting detection light to the vehicle's surrounding environment, and the acquisition unit is controlled to acquire images of the surrounding environment to obtain an image sequence; the target coding combination mode uses temporal coding and spatial coding to encode the light, thereby giving the detection light spatiotemporal coding characteristics; Spatiotemporal decoding of image sequences yields target physical features corresponding to the surrounding environment; these target physical features help distinguish between shadowed areas and solid objects in the surrounding environment. Based on the target's physical characteristics, parking perception information is generated that filters out the surrounding environment of the shadowed area, and the vehicle is controlled to park based on the parking perception information.

[0006] Secondly, embodiments of this application provide a parking device, the device comprising: The control module is used to control the light-emitting unit to work in the target coding combination mode, emit detection light to the vehicle's surrounding environment, and control the acquisition unit to acquire images of the surrounding environment to obtain an image sequence; the target coding combination mode uses temporal coding and spatial coding to encode the light, so as to endow the detection light with spatiotemporal coding characteristics; The decoding module is used to perform spatiotemporal decoding on image sequences to obtain the target physical features corresponding to the surrounding environment; the target physical features can help distinguish between shadow areas and solid objects in the surrounding environment; The parking module is used to generate parking perception information based on the target's physical characteristics, filtering out the surrounding environment of shadowed areas, and to control the vehicle to park based on the parking perception information.

[0007] Thirdly, embodiments of this application provide a parking device, including: a processor and a memory storing computer program instructions; When the processor executes computer program instructions, it implements parking methods such as the first aspect.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the parking method as described in the first aspect.

[0009] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a vehicle's processor, cause the vehicle to perform the parking method as described in the first aspect.

[0010] In this embodiment, by controlling the light-emitting unit to operate in the target encoding combination mode, a detection light is emitted to the vehicle's surrounding environment. The acquisition unit is controlled to acquire images of the surrounding environment to obtain an image sequence. By controlling the light-emitting unit to emit detection light with spatiotemporal encoding characteristics, the light projected onto the surrounding environment is no longer uniform ambient light, but active light carrying specific spatiotemporal information. The image sequence is spatiotemporally decoded to obtain the target physical features corresponding to the surrounding environment. The target physical features are used to distinguish shadow areas from physical reflection characteristics. Furthermore, when generating parking perception information based on the target physical features, interference from shadow areas is filtered out, thereby effectively avoiding false alarms caused by shadow areas and improving parking smoothness. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic flowchart of the parking method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the parking device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the parking equipment provided in the embodiments of this application. Detailed Implementation

[0013] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0015] In all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. Additionally, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments obtained.

[0016] At night or in low-light conditions, automatic parking primarily relies on visible light cameras and ultrasonic radar for perception. However, weak or complex artificial light sources, such as streetlights, can create strong contrasts and dynamic shadows on the ground and walls. These shadows visually resemble the edges and textures of low-lying obstacles, causing perception algorithms based on visible light images to frequently generate false alarms for ghost obstacles. This leads to unnecessary emergency braking or parking path interruptions during automatic parking, severely impacting user experience and the smoothness of the function.

[0017] To address the problems of the prior art, embodiments of this application provide a parking method, apparatus, parking equipment, and storage medium. The parking method provided by this application embodiment will be described first below.

[0018] Figure 1 A schematic flowchart of a parking method provided in one embodiment of this application is shown. Figure 1 As shown, the parking method provided in this application embodiment is applied to parking equipment. Specifically, the parking equipment can be a vehicle or an electronic device. The method can be executed by the parking equipment, or by components of the parking equipment, such as its processor, chip, or chip system. It can also be implemented by logic modules or software that implement all or part of the parking equipment's functions. In practical applications, the parking equipment can be a terminal, server, service platform, cloud, distributed system, Internet of Things (IoT), vehicle network system, etc. Further, the terminal can be a smartphone, tablet, laptop, desktop computer, etc. The method includes the following steps 101-103.

[0019] Step 101: Control the light-emitting unit to work in the target coding combination mode, emit detection light to the vehicle's surrounding environment, and control the acquisition unit to acquire images of the surrounding environment to obtain an image sequence; the target coding combination mode uses temporal coding and spatial coding to encode the light, thereby giving the detection light spatiotemporal coding characteristics.

[0020] In this step, the light-emitting unit can be an active light-emitting device that projects detection light with a specific coding pattern into the vehicle's surrounding environment. The light-emitting unit can be an infrared light-emitting diode array, a structured light projector, or a laser generator, etc.

[0021] The acquisition unit can be used as a device to acquire images of the environment around a vehicle. The acquisition unit captures the image sequence formed by the reflection of the detection light emitted by the light-emitting unit in the environment. The acquisition unit can be an infrared camera, a surround-view camera, etc.

[0022] Preferably, the light-emitting unit can be an infrared light-emitting diode (LED) array, and the acquisition unit can be a surround-view camera.

[0023] It should be noted that the light-emitting unit and the data-collecting unit can be independent of the vehicle or integrated into the vehicle; no specific limitation is made here.

[0024] For example, the light-emitting unit can be a programmable infrared coded emission unit. The light-emitting units can be distributed around the vehicle, usually sharing the same location as the acquisition unit. Each light-emitting unit includes an array of multiple infrared LEDs and its independent driving circuit.

[0025] The infrared LED array can use 850nm or 940nm automotive-specific wavelengths, matched with the infrared sensitive wavelength of the acquisition unit. The luminous power of a single LED can be 5-10mW. The number of luminous units can be configured according to the vehicle size, such as 1-2 units per side for small cars and 2-3 units per side for medium and large vehicles, ensuring coverage of the vehicle's surrounding perception blind spots and eliminating monitoring blind areas.

[0026] To improve the reliability of parking equipment, a primary-backup redundancy design can be adopted, with one backup light-emitting unit configured for each light-emitting unit. When the primary light-emitting unit is found to have no signal feedback, abnormal light-emitting power, or invalid data, it is determined to be a primary unit failure, and the system will immediately and automatically switch to the backup light-emitting unit.

[0027] The acquisition unit can be a vehicle-mounted surround-view camera, whose image sensor is sensitive to the aforementioned infrared band and can be hardware-synchronized with the infrared emitting unit.

[0028] In some implementations, the target coding combination pattern can be a pre-set fixed coding combination pattern or a coding combination pattern determined based on the environmental characteristics of the vehicle's surrounding environment; no specific limitation is made here.

[0029] The target coding combination mode encodes light using temporal coding and spatial coding. Temporal coding allows the projected light spot pattern or brightness to change periodically or non-periodically over time, such as inter-frame phase shift, flickering at a specific frequency, or switching between multiple frames, so that the same location in space has different temporal feature codes at different times. Spatial coding can present the projected light spot in space as a specific pattern, such as pseudo-random speckle pattern, stripe pattern, or grid pattern, so that pixels at different locations in space are assigned different spatial feature codes.

[0030] By superimposing temporal and spatial coding, the probe light emitted by the luminescent unit is endowed with spatiotemporal coding characteristics.

[0031] The acquisition frequency of the acquisition unit can be matched with the encoding period of the light-emitting unit to fully capture the reflection signals of the probe light at different times and spatial locations, thereby obtaining an image sequence containing spatiotemporal coding features.

[0032] Step 102: Perform spatiotemporal decoding on the image sequence to obtain the target physical features corresponding to the surrounding environment; the target physical features can help distinguish between shadow areas and solid objects in the surrounding environment.

[0033] In this step, spatiotemporal decoding may include temporal decoding and spatial decoding of the image sequence. Furthermore, a decoding method corresponding to the encoding method of the target encoding combination mode can be used to perform spatiotemporal decoding of the image sequence. Specific spatiotemporal decoding methods can be found in the relevant descriptions below, and will not be repeated here.

[0034] In some implementations, physical objects such as walls, other vehicles, and pedestrians have physical surfaces that can reflect probe light with spatiotemporal coding features. By spatiotemporally decoding the image sequences collected from these areas, high-confidence target physical features can be obtained. In contrast, shadow areas are dark areas where the light source is blocked, resulting in no direct light reaching them. The grayscale changes in the images they exhibit are caused by ambient diffuse reflection light and do not contain the spatiotemporal coding features injected by the transmitting unit.

[0035] During spatiotemporal decoding, shadowed regions cannot provide decoding results with physical depth and material reflection characteristics consistent with the emitted target encoding combination pattern, i.e., target physical parameters. In some implementations, shadowed regions and solid objects can be clearly distinguished from each other through physical mechanisms using target physical parameters.

[0036] The target physical parameters may include reflectance features and signal-to-noise ratio (SNR) features obtained by temporally decoding the reflected light intensity time sequence of the image sequence. The reflectance features are used to distinguish the material of the object from the shadow, and the SNR features are used to characterize the reliability of the decoding. It may also include depth distortion features obtained by spatially decoding the geometric distortion of the spatially encoded pattern in the image sequence, and the depth distortion features are used to calculate the three-dimensional shape of the object.

[0037] Step 103: Based on the target's physical characteristics, generate parking perception information of the surrounding environment after filtering out the shadow area, and control the vehicle to park based on the parking perception information.

[0038] In this step, parking perception information can be constructed based on the target's physical characteristics that support the distinction between shaded areas and physical objects. This parking perception information can be a list of obstacles filtered out of shaded areas and a map of passable areas in the surrounding environment; it can also be either the obstacle list or the passable area map. There are no specific limitations here, and it can be generated according to the actual situation.

[0039] In some implementations, the category of each region in an image sequence can be identified based on the target's physical characteristics to distinguish between shadowed areas and physical objects. After filtering out the shadowed areas, parking perception information is obtained, which can then be used to control the vehicle for parking.

[0040] In some implementations, when the parking device is a vehicle, the automatic parking system in the vehicle can automatically park based on parking perception information; when the parking device is a server or other terminal device, the parking perception information can be sent to the vehicle, and the automatic parking system in the vehicle can automatically park based on the parking perception information.

[0041] In this embodiment, by controlling the light-emitting unit to operate in the target encoding combination mode, a detection light is emitted to the vehicle's surrounding environment. The acquisition unit is controlled to acquire images of the surrounding environment to obtain an image sequence. By controlling the light-emitting unit to emit detection light with spatiotemporal encoding characteristics, the light projected onto the surrounding environment is no longer uniform ambient light, but active light carrying specific spatiotemporal information. The image sequence is spatiotemporally decoded to obtain the target physical features corresponding to the surrounding environment. The target physical features are used to distinguish shadow areas from physical reflection characteristics. Furthermore, when generating parking perception information based on the target physical features, interference from shadow areas is filtered out, thereby effectively avoiding false alarms caused by shadow areas and improving parking smoothness.

[0042] The complexity and diversity of the surrounding environment may prevent a single, fixed coding combination pattern from optimally adapting to all environmental conditions, thereby affecting the efficiency of detection light emission and image acquisition, as well as the accuracy of decoding. In one embodiment of this application, the target coding combination pattern can be determined through the following steps: Obtain the first environmental feature; the first environmental feature includes various environmental features corresponding to the surrounding environment. In the preset coding pattern library, a target environment feature matching the first environment feature is determined, and the coding combination pattern corresponding to the target environment feature is determined as the initial coding combination pattern; the preset coding pattern library stores at least one environment feature and the coding combination pattern corresponding to the environment feature; Determine the target encoding combination pattern based on the initial encoding combination pattern.

[0043] In this embodiment, the first environmental feature may include various parking-related environmental features such as ambient light intensity, the proportion of shadowed areas, and the parking scenario, without being specifically limited here. The first environmental feature can be acquired through sensors mounted or integrated on the vehicle, or through user input, etc.

[0044] The preset coding library can be a database pre-built based on experimental calibration or historical parking data. The preset coding library stores at least one environmental feature and the corresponding coding combination pattern for that environmental feature. For example, the preset coding library stores an environmental feature with low ambient light intensity and a shadow area ratio greater than 50%. The corresponding coding combination pattern might be a combination of low-frequency anti-overexposure temporal coding and specific anti-reflective pseudo-random spatial coding.

[0045] In some implementations, the coding mode library also includes extreme coding combination modes for extreme environments such as rain and fog. For example, in an extreme coding combination mode, the emission power of the light-emitting unit is set to the highest level, and an image noise reduction algorithm is activated to ensure perception performance.

[0046] Furthermore, the similarity between the acquired first environmental feature and each environmental feature in the preset coding pattern library can be calculated to determine the target environmental feature, and the coding combination pattern corresponding to the target environmental feature in the preset coding pattern library can be determined as the initial coding combination pattern.

[0047] In some implementations, environmental features with a similarity threshold greater than a preset similarity threshold can be identified as target environmental features. In some implementations, there may be multiple target environmental features. One of the multiple target environmental features can be randomly selected as the final target environmental feature, or the target environmental feature with the highest similarity among the multiple target environmental features can be selected as the final target environmental feature.

[0048] Furthermore, the initial coding combination pattern can be directly used as the target coding combination pattern, or the initial coding combination pattern can be fine-tuned based on the first environmental characteristics to obtain the target coding combination pattern; alternatively, the initial coding combination pattern can be fine-tuned based on real-time feedback, such as decoding signal-to-noise ratio and image quality, to obtain the target coding combination pattern. This allows for the adoption of optimal coding strategies for different environmental conditions, ensuring shadow filtering effectiveness while also considering computational efficiency and response speed.

[0049] In some implementations, the coding pattern combination patterns in the coding pattern library can be updated based on the collected valid parking scenarios to ensure that the adaptability of the coding pattern combination patterns in the coding pattern library is continuously improved.

[0050] In this embodiment, the coding combination mode is adaptively selected according to the environmental characteristics of the surrounding environment, which effectively improves the adaptability of the coding combination mode in complex and ever-changing parking environments.

[0051] In one embodiment of this application, the initial encoding combination mode can be adjusted to obtain the target encoding combination mode. Specifically, determining the target encoding combination mode based on the initial encoding combination mode may include the following: Based on the pre-defined relationship between various environmental features and the coding methods in the coding combination mode, the coding methods in the initial coding combination mode are adjusted to obtain the target coding combination mode.

[0052] In this embodiment, the relationship between each environmental feature and each encoding method in the encoding combination mode can be a pre-established set of rules describing the mapping relationship between environmental features and each encoding method in the encoding combination mode. Through the relationship, each encoding method in the initial encoding combination mode can be adjusted according to each environmental feature to obtain the target encoding combination mode.

[0053] In some implementations, the coding methods in the initial coding combination mode are adjusted. This can be done by adjusting the parameters of each coding method, such as adjusting the dot density of spatial coding or the frequency of temporal coding; or by adjusting the weight of each coding method in spatiotemporal decoding, such as increasing the weight coefficient of spatial matching in spatiotemporal decoding, so as to distinguish shadows from entities by relying on the features of actively projected structured light.

[0054] In some implementations, the first environmental feature may include ambient light intensity and shadow area proportion, wherein the ambient light intensity may be associated with temporal coding and the shadow area proportion may be associated with spatial coding.

[0055] Ambient light intensity can be positively correlated with the frequency in timing coding, that is, the higher the ambient light intensity, the higher the frequency in timing coding.

[0056] The proportion of shadow areas can be positively correlated with the weight of spatial coding. That is, the higher the proportion of shadow areas, the greater the weight of spatial coding in the initial coding combination mode. Specifically, the density of spatial coding can be increased, such as switching from sparse speckle to high-density random speckle, or increasing the weight coefficient of spatial matching in spatiotemporal decoding, so as to distinguish shadows from entities by relying on the features of actively projected structured light.

[0057] For example, when the ambient light intensity is <5 lux and the shadow area in the image accounts for >30%, the weight of spatial coding in the initial coding combination mode can be increased to obtain the target coding combination mode; when the ambient light intensity is between 5 and 10 lux and the shadow is small, the frequency of temporal coding in the initial coding combination mode can be increased, that is, high-frequency temporal coding can be used, and the spatial coding in the initial coding combination mode can be left unchanged to obtain the target coding combination mode.

[0058] In this embodiment, the initial coding combination mode can be finely adjusted according to the correlation between preset environmental features and each coding feature in the coding combination mode, thereby generating a more environmentally adaptable target coding combination mode, so that the probe light can better adapt to the complexity and diversity of the parking environment.

[0059] In one embodiment of this application, crosstalk from signals in other areas can be avoided by controlling the operating mode of the light-emitting unit. Specifically, controlling the light-emitting unit to operate in a target encoding combination mode to emit detection light to the vehicle's surrounding environment can include the following: The parking detection cycle is divided into multiple time periods, and each time period is mapped to a different directional area of ​​the surrounding environment. During each time period, the light-emitting unit in the direction area corresponding to the time period operates in the target encoding combination mode, emitting detection light to the vehicle's surrounding environment.

[0060] In this embodiment, the working cycle of parking detection can be the entire time period during which light detection and image acquisition are performed when the vehicle is parking. The specific value of the working cycle can be set according to the actual parking scenario requirements, and is not specifically limited here.

[0061] The working cycle of parking detection can be evenly divided into multiple time periods, or it can be divided according to environmental characteristics and vehicle movement status. For example, the working cycle of parking detection is set to 20ms, which is evenly divided into four 5ms time periods.

[0062] Furthermore, each time period is mapped to a different directional area of ​​the surrounding environment, that is, a correspondence is established between time periods and different directional areas. The different directional areas of the surrounding environment can be centered on the vehicle, dividing the surrounding environment space that the vehicle needs to perceive into multiple different directional areas in the horizontal direction, such as 360 degrees, and / or the vertical direction, such as the left front area, the front area, the right front area, and the right side area.

[0063] For example, multiple directional areas around the vehicle can be pre-defined, such as front left, front right, rear left, and rear right, and one or more specific time periods can be assigned to each area; or, the mapping relationship between time periods and directional areas can be dynamically adjusted according to the vehicle's current driving direction, steering intention, or sensor data, prioritizing the detection of the most relevant directional area at present.

[0064] In some implementations, multiple light-emitting units are mounted or integrated on the vehicle in a surrounding arrangement. During each time period, the light-emitting units in the direction area corresponding to that time period are controlled to operate in a target coding combination mode to emit detection light to the vehicle's surrounding environment. That is, during a specific time period, only the light-emitting units responsible for detecting in a specific direction area are activated and made to emit detection light according to the target coding combination mode.

[0065] In some implementations, during each time period, the light-emitting unit in the direction corresponding to that time period is controlled to operate in the target encoding combination mode. While emitting detection light to the vehicle's surrounding environment, the acquisition unit in the direction corresponding to that time period is triggered to ensure that the exposure time of the acquisition unit is accurately aligned with the effective emission phase of the encoded light, thereby avoiding crosstalk of multiple signals from the source.

[0066] For example, a working cycle is 20ms, which is divided into four 5ms time periods, in the order of left front → right front → left back → right back. In each time period, only the light-emitting unit and the acquisition unit in the corresponding direction are activated.

[0067] In this embodiment, the parking detection cycle is divided into multiple time periods, and each time period is mapped to a different directional area of ​​the surrounding environment. Within each time period, the light-emitting unit corresponding to the directional area of ​​that time period operates in a target encoding combination mode, emitting detection light into the vehicle's surrounding environment. This allows the detection light to be emitted in a time-multiplexed and spatially directional manner, ensuring that at any given moment, only the specific directional area corresponding to the current time period is actively illuminated, rather than the entire environment. This effectively avoids unnecessary consumption of detection light in non-target areas, significantly reducing overall energy consumption. Simultaneously, because the detection light from different directional areas is staggered in time, mutual interference between beams from different directions is greatly reduced, and the signal-to-noise ratio and clarity of the image sequences acquired by the acquisition unit can be improved.

[0068] In one embodiment of this application, spatiotemporal decoding of an image sequence to obtain the target's physical features corresponding to the surrounding environment may include the following: Based on the temporal coding sequence in the target coding combination mode and the brightness change sequence of the image sequence, the reflectivity features and signal-to-noise ratio features of the surrounding environment are determined; the reflectivity features are used to distinguish the material of objects in the surrounding environment, and the signal-to-noise ratio features are used to characterize the reliability of the decoding. Feature point matching is performed on the spatially encoded point array in the image sequence to determine the depth distortion features of the surrounding environment; the depth distortion features are used to characterize the three-dimensional shape of objects in the surrounding environment. The target's physical characteristics include reflectivity, signal-to-noise ratio, and depth distortion.

[0069] In this embodiment, the temporal coding sequence can be part of the target coding combination mode, used to give the probe light a coding pattern in the time dimension. The temporal coding sequence can modulate the switching state of the light-emitting unit through a preset pseudo-random sequence to form a light signal with specific temporal characteristics.

[0070] When the probe light is emitted in a time-coded sequence, the light reflected from the object also carries this time-coded information, which is represented by the change in brightness over time in the image sequence.

[0071] A brightness variation sequence in an image sequence can be a sequence of image pixel brightness changes over time, captured by the acquisition unit within consecutive time frames. The brightness variation sequence can be obtained by sampling and analyzing the brightness value of each pixel in the image sequence over time, resulting in a curve or sequence of brightness changes over time; alternatively, the image sequence can be processed by inter-frame differencing or temporal filtering to highlight brightness fluctuations caused by the probe light, thus forming a brightness variation sequence.

[0072] Furthermore, by comparing the temporal coding sequence with the brightness change sequence of each pixel, the reflectivity and ambient light substrate of each pixel are calculated to determine the reflectivity characteristics and signal-to-noise ratio characteristics of the surrounding environment.

[0073] In some implementations, reflectivity characteristics can be obtained by comparing the temporal coding sequence of the probe light with the image brightness change sequence, thereby quantifying the intensity of the object's reflection of the probe light. Reflectivity characteristics can be inferred by calculating the correlation or amplitude ratio between the received brightness change sequence and the known emission temporal coding sequence; alternatively, reflectivity can be calculated by measuring the brightness response of the same pixel under different temporal coding states and then using the ratio of these response values ​​to the emission light intensity.

[0074] Signal-to-noise ratio (SNR) is a metric used to measure the ratio of the decoded signal strength to the noise strength, thus characterizing the reliability or credibility of the decoding result. A high SNR means the decoding result is less affected by noise and is more reliable. SNR can be calculated by comparing the strength of the encoded signal extracted from the brightness variation sequence with the background noise level; alternatively, it can be obtained by statistically analyzing the signal fluctuation amplitude and average noise level during the decoding process.

[0075] In some implementations, the spatial coding dot matrix can be part of the target coding combination pattern, used to give the probe light a coding pattern in the spatial dimension. The spatial coding dot matrix can be represented as a light spot or pattern with a specific geometric arrangement.

[0076] Furthermore, for spatially encoded dot matrices, feature point matching can be used, combined with known dot projection geometry parameters such as dot spacing and projection angle, to calculate the dot matrix distortion caused by the target depth using the principle of monocular vision geometry, thereby generating a depth residual map that represents the preliminary depth information, i.e., depth distortion features.

[0077] In some implementations, feature point matching can be achieved by identifying and tracing feature points in a spatially coded dot matrix within an image sequence and comparing them with a preset or reference spatially coded dot matrix. By matching the positional changes or deformations of these feature points in the image, the three-dimensional information of the object can be inferred.

[0078] Feature point matching can be achieved by using correlation-based or descriptor-based methods to identify and match corresponding feature points in a spatially coded point array under different image frames or different viewpoints; alternatively, it can be achieved by using projection geometry principles to match the spatially coded point array detected in the image with a known projection pattern, thereby establishing a correspondence between pixels and three-dimensional spatial points.

[0079] Depth distortion features can describe the shape and depth information of an object's surface in three-dimensional space. When a spatially encoded bitmap is projected onto the surface of an object with different depths, the shape, density, or position of the bitmap will be distorted due to perspective effects. By analyzing these distortions, the object's three-dimensional geometry and depth information can be deduced.

[0080] Depth distortion features can be calculated based on the parallax or projection deformation of feature points in an image, combined with the geometric parameters of the dot matrix projection, such as the calibration parameters of the camera and projector, using the principle of triangulation. Alternatively, depth distortion features can be obtained by analyzing the local distortion, density changes, or relative displacement of the pattern formed by the spatially encoded dot matrix on the object's surface, using structured light 3D reconstruction algorithms.

[0081] In some implementations, reflectivity features, signal-to-noise ratio features, and depth distortion features can be represented as feature maps.

[0082] In other implementations, grayscale gradient maps of the image sequence can also be calculated and used as auxiliary edge features of the grayscale gradient map as one of the feature maps subsequently input into a pre-trained neural network.

[0083] In this embodiment, by analyzing the temporal coding sequence and the image brightness change sequence, the reflectivity characteristics of the object and the signal-to-noise ratio characteristics of the decoded data can be obtained to identify objects of different materials and evaluate the reliability of the perceived data. Simultaneously, by performing feature point matching on the spatial coding matrix, the depth distortion characteristics of the object can be accurately obtained, thereby accurately representing the three-dimensional shape of the object. The coupling of these two methods, from a physical perspective, distinguishes between two-dimensional light and shadow interference and three-dimensional entities, significantly enhancing the vehicle's automatic parking system's ability to distinguish between shadow areas and solid objects during parking, avoiding parking errors or unnecessary obstacle avoidance behaviors caused by shadow misjudgment in traditional methods.

[0084] In one embodiment of this application, parking perception information, which filters out shadow areas, is generated based on the target's physical characteristics. This may include the following: The target physical features are input into a pre-trained neural network model, which outputs a dense depth map and a semantic segmentation map of the surrounding environment. The dense depth map is used to represent the three-dimensional depth information of each point in the surrounding environment, and the semantic segmentation map is used to represent the category of objects in the surrounding environment. Based on the dense depth map and semantic segmentation map, parking perception information is generated after filtering out shadow areas.

[0085] In this embodiment, the pre-trained neural network model can be a computational model that has been trained on a large amount of data and possesses specific perceptual capabilities. The neural network model can adopt a deep learning architecture, such as a convolutional neural network or a recurrent neural network, etc., without being specifically limited here.

[0086] In some implementations, the neural network model can adopt an encoder-decoder structure, where the decoder can include two branches: depth estimation and semantic segmentation.

[0087] A large amount of simulation and real-vehicle data generated based on the active coding infrared principle can be used as training data. The training data covers different low-light conditions, shadow types, and obstacles, and has pixel-level annotations. In some implementations, the semantic annotation accuracy of pixel-level annotations needs to be ≥99%, and the depth annotation error needs to be ≤3cm to ensure the accuracy of the neural network model after training.

[0088] Furthermore, the training data is divided into a training set, a validation set, and a test set. The validation set can contain real vehicle and simulation data under different scenarios and different types of interference. In some implementations, the training data can be divided into the training set, validation set, and test set in a 7:2:1 ratio.

[0089] When training a neural network model, a multi-task loss function can be used for joint optimization. The multi-task loss function can be represented by the following equation: L_total=α × L_depth+β × L_seg (1) In the formula, L_total Represents the total loss of the neural network model L_depth This indicates the depth estimation task loss. L_ seg This represents the loss of the semantic segmentation task. α and β This represents the weighting coefficient.

[0090] The weight coefficients α and β can be set according to the actual situation. For example, they can be set between 0.4 and 0.6, and can also be dynamically adjusted according to the depth estimation accuracy and semantic segmentation accuracy of the validation set.

[0091] The output of a neural network model can be a dense depth map with confidence, as well as a pixel-level semantic segmentation map.

[0092] In the dense depth map, depth values ​​can locate the position of obstacles; confidence level can characterize the reliability of depth values ​​in the dense depth map. For example, for low-confidence areas, which may be noise, distant blurry objects, or weak reflections, a more conservative strategy such as deceleration and further confirmation can be adopted. However, for obstacles detected in high-confidence areas, obstacle avoidance can be decisively executed, greatly improving the decision-making safety and intelligence level of the entire automatic parking system.

[0093] The pixel-level semantic segmentation map includes annotations for the ground, obstacles, dynamic / static shadows, and specular reflections.

[0094] Furthermore, the target's physical features, capable of distinguishing shadowed areas from solid objects, are used as input to a pre-trained neural network model. The pre-trained neural network model deeply mines the geometric and semantic information contained within the target's physical features and transforms it into a dense depth map and a semantic segmentation map. The dense depth map provides an accurate 3D geometric structure of the surrounding environment, while the semantic segmentation map performs pixel-level classification of various objects and regions within the surrounding environment.

[0095] Furthermore, based on the dense depth map and semantic segmentation map, parking perception information that accurately filters out shadow areas is generated.

[0096] In this embodiment, physical features capable of distinguishing shadows are input into a pre-trained neural network model to efficiently and accurately extract dense depth maps and semantic segmentation maps from the complex surrounding environment. The depth maps and semantic segmentation maps are then combined to generate parking perception information that filters out shadow areas, greatly improving the accuracy and reliability of parking perception information.

[0097] In one embodiment of this application, parking perception information with shadow areas filtered out is generated based on a dense depth map and a semantic segmentation map, which may include the following: Obtain the reflectance of regions marked as specular reflection in the semantic segmentation map; If the reflectivity is greater than a preset reflectivity threshold, the area of ​​specular reflection is defined as the interference area. The shaded and interfering regions in the semantic segmentation map are identified as the target regions; The depth values ​​of the regions corresponding to the target region in the dense depth map are set to invalid, resulting in a target dense depth map that has filtered out the target region. Based on the target dense depth map, parking perception information is generated after filtering out shadow areas.

[0098] In this embodiment, a set of pixels marked as specular reflection in the semantic segmentation map can be obtained. Based on the aforementioned reflectivity features, the reflectivity of the specular reflection region is obtained. If the reflectivity is greater than a preset reflectivity threshold, the specular reflection region is determined not to be a real object, but rather an interference region that may produce optical artifacts. The preset reflectivity threshold can be set according to actual conditions and is not specifically limited here; for example, it can be set to 0.8-1.0.

[0099] Furthermore, the regions marked as shaded and interference regions in the semantic segmentation map are identified as the target regions. Shaded and interference regions are, in essence, invalid optical regions that cannot provide effective, true 3D depth information. In the dense depth map output by the pre-trained neural network model, the depth value of the pixel position corresponding to the target region is directly set to invalid, for example, by assigning a specific invalid identifier such as NaN, 0 or a maximum value, to indicate that the depth information here is unavailable, thus obtaining a target dense depth map that filters out the target region, thereby eliminating erroneous depth information caused by shadows and strong specular reflections.

[0100] In some implementations, dense depth maps are confidence-based depth maps, which can significantly reduce the detection confidence of obstacles corresponding to target regions. Thus, when generating parking perception information, for potential obstacle targets mainly located in shadow areas, the detection confidence of the target region as a whole can be multiplied by a very small coefficient, such as 0.1 or 0.01, making the target region almost ignored in the final ranking and decision-making process unless supported by other strong evidence.

[0101] Furthermore, parking perception information, after filtering out shadow areas, is generated based on the target dense depth map.

[0102] In some implementations, parking perception information may include a list of obstacles and a map of accessible areas.

[0103] Specifically, the obstacle list can be generated as follows: In the semantic segmentation map, extract all connected regions marked as obstacles. Based on the target dense depth map, calculate the average depth, nearest-neighbor depth, 3D bounding box, and position of each obstacle region. Combine the depth confidence and semantic segmentation confidence of the obstacle region to calculate the overall confidence of the obstacle region. Organize the attributes of all obstacles in the obstacle region, such as type, position, size, velocity, and overall confidence, into a structured list, i.e., the obstacle list.

[0104] A passable area map can be generated as follows: Areas marked as ground in the semantic segmentation map with continuous and reasonable depth values ​​are initially identified as passable areas. Based on the target dense depth map, the slope and undulation height of the ground are calculated. Areas occupied by obstacles, areas with slopes or undulations exceeding a preset safety threshold, and areas lacking depth information or with confidence levels below a preset confidence threshold are excluded. Finally, a two-dimensional grid map or area outline map is generated, clearly marking safe and drivable areas; this is the passable area map. The preset safety threshold and preset confidence threshold can be set according to actual conditions and are not specifically limited here.

[0105] In this embodiment, by identifying and filtering out erroneous depth information caused by shadows and highly reflective mirror areas in the parking environment, the accuracy and reliability of parking perception information are significantly improved, thereby enhancing the safety of the parking process.

[0106] It should be noted that the various embodiments described in this application can be combined with each other or implemented individually without conflict, and this application does not limit this.

[0107] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0108] For ease of understanding, a specific embodiment will be used as an example: 1) Step S1: Environmental perception and coding strategy decision.

[0109] The vehicle's ambient light intensity sensor can acquire real-time ambient light intensity at a frequency of 10Hz, and obtain vehicle speed, gear status, and other information through the vehicle's Controller Area Network (CAN) bus.

[0110] When the vehicle is determined to be in a typical parking scenario, i.e., ambient light intensity ≤ 10 lux and vehicle speed ≤ 10 km / h, it automatically enters the active coding perception mode.

[0111] The optimal coding combination mode is selected from the coding strategy library (i.e., the preset coding pattern library) according to preset rules, that is, by matching the environmental features around the vehicle with the environmental features in the coding strategy library. The coding combination mode can be a collaborative combination of temporal coding for anti-interference and spatial coding for deep sampling.

[0112] The coding strategy library can be dynamically adjusted based on environmental characteristics. For example, when the ambient light intensity is <5 lux and the shadow area accounts for >30% of the image, the spatial coding weight is increased; when the ambient light intensity is between 5-10 lux and the shadows are small, high-frequency temporal coding is prioritized. The coding strategy library also supports online updates based on actual operating data. When the cumulative collection of valid scene data reaches 100,000 frames, parameter optimization and updates can be automatically initiated to ensure continuous improvement in the adaptability of the coding strategy.

[0113] 2) Step S2: Synchronous encoding projection and image acquisition.

[0114] It can generate synchronous control signals and send them to the drive circuits of each infrared emitting unit (i.e., light-emitting unit). The drive circuit converts them into precise pulse width modulation (PWM) or digital signals to control the infrared LED array to operate in time-division and zone according to the selected mode (i.e., target code combination mode) in the 850nm or 940nm vehicle-specific band.

[0115] Specifically, the time-sharing and zone-based working system can be implemented as follows: a complete working cycle is 20ms, divided into four 5ms time periods, in the order of left front → right front → left rear → right rear. Within each time period, only the infrared emitting unit and the surround-view camera (i.e., the acquisition unit) in the corresponding direction are activated. A hardware synchronization circuit is used, sharing a synchronization clock with a synchronization accuracy of ≤1ms. This means the time difference between infrared emission and camera exposure does not exceed 1ms, ensuring that the camera exposure time is precisely aligned with the effective emission phase of the coded light, thus fundamentally avoiding crosstalk from multiple signals.

[0116] 3) Step S3: Encoding feedback decoding and special feature map generation.

[0117] The image sequences acquired from each direction are processed to generate dedicated feature maps: Timing Decoding: Compare the preset emission coding sequence with the brightness change sequence of each pixel, calculate the reflectance of the pixel and the ambient light substrate, and generate a reflectance map (i.e., reflectance feature) and a signal-to-noise ratio map (signal-to-noise ratio feature).

[0118] Spatial decoding: For spatially encoded dot matrices in an image sequence, feature point matching is used, combined with known dot projection geometry parameters such as dot spacing and projection angle, to calculate the dot matrix distortion caused by the target depth using monocular vision geometry principles, and generate a depth residual map (i.e., depth distortion feature) that represents the preliminary depth information.

[0119] Simultaneously, the grayscale gradient map of the original image is calculated as an auxiliary edge feature.

[0120] The above feature maps together constitute the input for subsequent neural network inference.

[0121] 4) Step S4: Integrated neural network inference and training.

[0122] The multi-channel dedicated feature maps generated by S3 are input into the trained integrated neural network model. This neural network model adopts an encoder-decoder structure, with the decoder containing two branches: depth estimation and semantic segmentation. During training, a multi-task loss function is used for joint optimization, which can be referred to in Equation 1 above.

[0123] A large amount of simulation and real vehicle data generated based on the active coding infrared principle can be used as training data. The training data covers different low-light conditions, shadow types, and obstacles, and has pixel-level annotations. The semantic annotation accuracy is ≥99%, and the depth annotation error is ≤3cm. The training data is divided into training set, validation set, and test set in a 7:2:1 ratio. The validation set contains real vehicle and simulation data with different scenarios and different interference types.

[0124] The final output of the neural network model: result Figure 1 : Dense depth map with confidence level.

[0125] result Figure 2 Pixel-level semantic segmentation map, which is labeled with ground, obstacles, dynamic / static shadows, specular reflections, etc.

[0126] 5) Step S5: Result fusion, environment adaptation and output.

[0127] Post-processing involves fusing depth maps and semantic maps: For regions marked as shaded in the semantic segmentation map, their depth values ​​in the dense depth map are invalidated, and the corresponding obstacle detection confidence is significantly reduced.

[0128] For areas marked as specular reflection, in addition to the above operations, a reflectivity threshold verification is added, with a preset threshold of 0.8-1.0. If the reflectivity of the area exceeds the threshold, it is identified as high reflectivity interference and excluded, thereby effectively avoiding missed obstacle detection due to mislabeling.

[0129] For extreme environments such as rain and fog, it can automatically increase the infrared emission power (up to 15mW) and start the image noise reduction algorithm, while calling the extreme environment coding mode in the coding strategy library to ensure perception performance.

[0130] Based on the parking method provided in the above embodiments, this application also provides specific implementations of the parking device. Please refer to the following embodiments.

[0131] See Figure 2 The parking device 300 provided in this application embodiment may include: The control module 301 is used to control the light-emitting unit to work in the target coding combination mode, emit detection light to the vehicle's surrounding environment, and control the acquisition unit to acquire images of the surrounding environment to obtain an image sequence; the target coding combination mode uses temporal coding and spatial coding to encode the light, which is used to give the detection light spatiotemporal coding characteristics. Decoding module 302 is used to perform spatiotemporal decoding on image sequences to obtain target physical features corresponding to the surrounding environment; the target physical features support the distinction between shadow areas and solid objects in the surrounding environment; The parking module 303 is used to generate parking perception information of the surrounding environment after filtering out the shadow area based on the physical characteristics of the target, and to control the vehicle to park based on the parking perception information.

[0132] In one embodiment of this application, the target encoding combination pattern is determined through the following steps: Obtain the first environmental feature; the first environmental feature includes various environmental features corresponding to the surrounding environment. In the preset coding pattern library, a target environment feature matching the first environment feature is determined, and the coding combination pattern corresponding to the target environment feature is determined as the initial coding combination pattern; the preset coding pattern library stores at least one environment feature and the coding combination pattern corresponding to the environment feature; Determine the target encoding combination pattern based on the initial encoding combination pattern.

[0133] In one embodiment of this application, determining the target encoding combination pattern based on the initial encoding combination pattern includes: Based on the pre-defined relationship between various environmental features and the coding methods in the coding combination mode, the coding methods in the initial coding combination mode are adjusted to obtain the target coding combination mode.

[0134] In one embodiment of this application, the control module 301 may include: The sub-module is used to divide the parking detection work cycle into multiple time periods and map each time period to different directional areas of the surrounding environment. The transmitting submodule is used to control the light-emitting units in the corresponding directional area to operate in the target encoding combination mode during each time period, so as to emit detection light to the vehicle's surrounding environment.

[0135] In one embodiment of this application, the decoding module 302 may include: The temporal decoding submodule is used to determine the reflectivity and signal-to-noise ratio features of the surrounding environment based on the temporal coding sequence in the target coding combination mode and the brightness change sequence of the image sequence. The reflectivity features are used to distinguish the material of objects in the surrounding environment, and the signal-to-noise ratio features are used to characterize the reliability of the decoding. The spatial decoding submodule is used to perform feature point matching on the spatially encoded point array in the image sequence to determine the depth distortion features of the surrounding environment; the depth distortion features are used to characterize the three-dimensional shape of objects in the surrounding environment. The target's physical characteristics include reflectivity, signal-to-noise ratio, and depth distortion.

[0136] In one embodiment of this application, the parking module 303 may include: The output submodule is used to input the target physical features into a pre-trained neural network model, and the pre-trained neural network model outputs a dense depth map and a semantic segmentation map of the surrounding environment; wherein, the dense depth map is used to represent the three-dimensional depth information of each point in the surrounding environment, and the semantic segmentation map is used to represent the category of objects in the surrounding environment. The generation submodule is used to generate parking perception information that filters out shadow areas based on the dense depth map and semantic segmentation map.

[0137] In one embodiment of this application, the generation submodule may include: The acquisition unit is used to acquire the reflectance of regions marked as specular reflection in the semantic segmentation map; The first determining unit is used to determine the area of ​​specular reflection as the interference area when the reflectivity is greater than a preset reflectivity threshold. The second determining unit is used to determine the shaded and interference regions in the semantic segmentation map as the target region; The filtering unit is used to invalidate the depth values ​​of the regions corresponding to the target region in the dense depth map, thereby obtaining a target dense depth map with the target region filtered out. The generation unit is used to generate parking perception information that has been filtered out of shadow areas based on the target dense depth map.

[0138] The parking device provided in this application embodiment can realize all the processes implemented in the aforementioned parking method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0139] Figure 3 A schematic diagram of the hardware structure of the parking equipment provided in an embodiment of this application is shown.

[0140] The parking device may include a processor 401 and a memory 402 storing computer program instructions.

[0141] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0142] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.

[0143] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.

[0144] The processor 401 implements any of the parking methods described in the above embodiments by reading and executing computer program instructions stored in the memory 402.

[0145] In some embodiments, the parking device can be a vehicle. In these embodiments, the processor 401 implements the above by reading and executing computer program instructions stored in the memory 402. Figure 1 Any parking method in the method embodiments.

[0146] In one example, the parking device may also include a communication interface 403 and a bus 410. For example, Figure 3 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.

[0147] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0148] Bus 410 includes hardware, software, or both, that couples the components of the parking equipment together. This is an example, not a limitation. The bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0149] Furthermore, in conjunction with the parking methods described in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the parking methods described in the above embodiments.

[0150] This application embodiment may also provide a computer program product, wherein when the instructions in the computer program product are executed by the processor of the parking device, the parking device performs any of the parking methods described in the above embodiments.

[0151] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0152] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM, floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0153] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0154] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0155] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. A parking method, characterized in that, The method includes: The light-emitting unit is controlled to operate in a target encoding combination mode to emit detection light to the vehicle's surrounding environment, and the acquisition unit is controlled to acquire images of the surrounding environment to obtain an image sequence; the target encoding combination mode uses temporal encoding and spatial encoding to encode the light, thereby giving the detection light spatiotemporal encoding characteristics; Spatiotemporal decoding is performed on the image sequence to obtain the target physical features corresponding to the surrounding environment; the target physical features support the distinction between shadow areas and solid objects in the surrounding environment; Based on the target's physical characteristics, parking perception information of the surrounding environment, after filtering out the shadow area, is generated, and the vehicle is controlled to park based on the parking perception information.

2. The method according to claim 1, characterized in that, The target coding combination pattern is determined through the following steps: Obtain a first environmental feature; the first environmental feature includes various environmental features corresponding to the surrounding environment; A target environment feature matching the first environment feature is determined from a preset encoding pattern library, and the encoding combination pattern corresponding to the target environment feature is determined as the initial encoding combination pattern; the preset encoding pattern library stores at least one environment feature and the encoding combination pattern corresponding to the environment feature; The target encoding combination pattern is determined based on the initial encoding combination pattern.

3. The method according to claim 2, characterized in that, Determining the target encoding combination pattern based on the initial encoding combination pattern includes: Based on the preset correlation between the environmental features and the encoding methods in the encoding combination mode, the encoding methods in the initial encoding combination mode are adjusted to obtain the target encoding combination mode.

4. The method according to claim 1, characterized in that, The control light-emitting unit operates in a target encoding combination mode, emitting detection light to the vehicle's surrounding environment, including: The parking detection cycle is divided into multiple time periods, and each time period is mapped to a different directional area of ​​the surrounding environment. During each of the aforementioned time periods, the light-emitting units in the directional area corresponding to the time period are controlled to operate in the target encoding combination mode to emit detection light to the vehicle's surrounding environment.

5. The method according to claim 1, characterized in that, The step of performing spatiotemporal decoding on the image sequence to obtain the target physical features corresponding to the surrounding environment includes: Based on the temporal coding sequence in the target coding combination mode and the brightness change sequence of the image sequence, the reflectivity features and signal-to-noise ratio features of the surrounding environment are determined; the reflectivity features are used to distinguish the material of objects in the surrounding environment, and the signal-to-noise ratio features are used to characterize the reliability of the decoding. Feature point matching is performed on the spatially encoded dot matrix in the image sequence to determine the depth distortion features of the surrounding environment; the depth distortion features are used to characterize the three-dimensional shape of objects in the surrounding environment. The target physical features include the reflectivity feature, the signal-to-noise ratio feature, and the depth distortion feature.

6. The method according to claim 1, characterized in that, The step of generating parking perception information of the surrounding environment, after filtering out the shadow area, based on the target's physical characteristics includes: The target physical features are input into a pre-trained neural network model, which outputs a dense depth map and a semantic segmentation map of the surrounding environment. The dense depth map is used to characterize the three-dimensional depth information of each point in the surrounding environment, and the semantic segmentation map is used to characterize the category of objects in the surrounding environment. The parking perception information, after filtering out the shadow areas, is generated based on the dense depth map and the semantic segmentation map.

7. The method according to claim 6, characterized in that, The step of generating the parking perception information, after filtering out the shadow region, based on the dense depth map and the semantic segmentation map includes: Obtain the reflectance of the regions marked as specular reflection in the semantic segmentation map; If the reflectivity is greater than a preset reflectivity threshold, the area of ​​specular reflection is defined as an interference area. The shaded areas and the interference areas in the semantic segmentation map are identified as the target areas; The depth values ​​of the regions corresponding to the target regions in the dense depth map are set to invalid, resulting in a target dense depth map that has filtered out the target regions; Based on the target density depth map, the parking perception information that filters out the shadow areas is generated.

8. A parking device, characterized in that, The device includes: The control module is used to control the light-emitting unit to operate in the target encoding combination mode, emit detection light to the vehicle's surrounding environment, and control the acquisition unit to acquire images of the surrounding environment to obtain an image sequence; the target encoding combination mode uses temporal encoding and spatial encoding to encode the light, thereby giving the detection light spatiotemporal encoding characteristics; The decoding module is used to perform spatiotemporal decoding on the image sequence to obtain the target physical features corresponding to the surrounding environment; the target physical features support the distinction between shadow areas and solid objects in the surrounding environment; The parking module is used to generate parking perception information of the surrounding environment after filtering out the shadow area based on the target's physical characteristics, and to control the vehicle to park based on the parking perception information.

9. A parking device, characterized in that, The parking equipment includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the parking method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the parking method as described in any one of claims 1-7.