Parking identification method and device based on multi-modal large model
Parking images taken by the user terminal are used to identify parking using multimodal large models, solving the problem of inaccurate identification of fixed cameras in complex environments, and achieving high-accurate parking specification judgments.
Patent Information
- Application Number
- CN202510151267.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-20
AI Technical Summary
The existing parking recognition method is based on fixed cameras, and cannot obtain effective perspective in complex environments such as multi-story parking lots and narrow streets, resulting in inaccurate identification.
The parking recognition method based on a multimodal large model is adopted to identify the parking image taken by the user terminal to determine whether the vehicle is parked in a standardized manner, including the complete display of the vehicle in the image, the number of the vehicle included in the image, and the vehicle is in the return area.
It realizes parking identification that operates stably in complex environments, reduces hardware costs and maintenance costs, and improves judgment accuracy.
Smart Images

Figure CN120182946A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a parking recognition method and device based on a multimodal large model. Background Art
[0002] With the wide application of shared electric bicycles, strict requirements are put forward for the parking positions of electric bicycles. It is required that the vehicles must be parked within the designated areas to maintain urban traffic order and the cleanliness of public spaces.
[0003] Existing methods are based on fixedly installing cameras on electric bicycles, and these cameras are used to monitor the parking positions and postures of the vehicles. The fixedly installed cameras cannot obtain effective perspectives in complex environments such as multi-story parking lots and narrow streets, resulting in inaccurate recognition and an inability to accurately judge whether the vehicles are parked in compliance. Summary of the Invention
[0004] The present invention provides a parking recognition method and device based on a multimodal large model to improve the accuracy of judging whether a vehicle is parked in compliance.
[0005] The present invention provides a parking recognition method based on a multimodal large model, including the following steps: Receiving a parking image of a target vehicle sent by a user terminal; Based on the multimodal large model, pre-judging the parking image to determine that the parking image meets preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area; Based on the parking image, determining the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line; Based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, determining whether the target vehicle is parked in compliance.
[0006] According to the parking recognition method based on a multimodal large model provided by the present invention, the step of pre-judging the parking image based on the multimodal large model to determine that the parking image meets preset conditions includes: Based on predefined prompt words, guiding the multimodal large model to recognize the target vehicle in the parking image, determining that the target vehicle is completely displayed in the parking image, guiding the multimodal large model to recognize the vehicle number in the parking image, determining that the vehicle number is the vehicle number of the target vehicle, and guiding the multimodal large model to recognize the surrounding information of the target vehicle, determining that the target vehicle is in the return area.
[0007] A parking recognition method based on a multimodal large model provided by the present invention, based on the multimodal large model, pre-judging the parking image, further including: In the case where it is determined that the parking image does not meet the preset conditions, based on the conditions not met during the pre-judging process of the multimodal large model, determine the photographing adjustment information; Send the photographing adjustment information to the usage terminal of the target vehicle.
[0008] A parking recognition method based on a multimodal large model provided by the present invention, based on the parking image, determining the placement angle of the target vehicle, including: Based on the parking image, extract the contour information of the target vehicle; Based on the contour information, determine the angle between the target vehicle and the reference axis; Based on the angle between the target vehicle and the reference axis, determine the placement angle of the target vehicle.
[0009] A parking recognition method based on a multimodal large model provided by the present invention, based on the parking image, determining the distance between the target vehicle and other vehicles and the distance between the target vehicle and the parking line, including: Identify the seat of the target vehicle in the parking image, and determine the seat position information of the seat in the parking image; Identify other vehicles in the parking image, and determine the position information of other vehicles in the parking image; Identify the parking line in the parking image, and determine the parking line position information of the parking line in the parking image; Based on the seat position information and the position information of other vehicles, determine the distance between the target vehicle and other vehicles, and based on the seat position information and the parking line position information, determine the distance between the target vehicle and the parking line.
[0010] A parking recognition method based on a multimodal large model provided by the present invention, the identifying the seat of the target vehicle in the parking image and determining the seat position information of the seat in the parking image includes: Perform image segmentation on the parking image to determine the seat image of the target vehicle; Convert the seat image into a point cloud image, and based on the density clustering method, segment the point cloud clusters of the point cloud image; Extract the largest cluster from the segmented point cloud clusters, determine the seat shape of the target vehicle, and based on the seat shape, determine the seat position information of the seat in the parking image.
[0011] The present invention also provides a parking recognition device based on a multimodal large model, including the following modules: A data receiving module, configured to receive a parking image of a target vehicle sent by a user terminal; A pre-judgment module, configured to pre-judge the parking image based on the multimodal large model to determine that the parking image meets preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area; An image judgment module, configured to determine the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line based on the parking image; A standard parking judgment module, configured to determine whether the target vehicle parks in a standard manner based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the parking recognition method based on the multimodal large model as described in any one of the above.
[0013] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the parking recognition method based on the multimodal large model as described in any one of the above.
[0014] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the parking recognition method based on the multimodal large model as described in any one of the above.
[0015] The parking recognition method and device based on the multimodal large model provided by the present invention perform recognition through the parking image captured by the user terminal instead of a fixed camera, without the need to install a camera on the electric vehicle, and can operate stably in complex environments such as multi-story parking lots and narrow streets, without being affected by the viewing angle limitation of the fixed camera, reducing the hardware cost and maintenance expenses. Based on the multimodal large model, pre-judge whether the captured parking image meets the preset conditions. After the pre-judgment passes, further judge the key elements of the vehicle based on the parking image, realizing the automatic judgment process of whether the parking is standard and improving the accuracy of the judgment. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of a parking recognition method based on a multimodal large model provided by the present invention.
[0018] Figure 2 It is a schematic flowchart of the implementation process based on the multimodal large model provided by the present invention.
[0019] Figure 3 It is a schematic flowchart of the seat recognition process provided by the present invention.
[0020] Figure 4 It is a schematic structural diagram of a parking recognition device based on a multimodal large model provided by the present invention.
[0021] Figure 5 It is a schematic structural diagram of an electronic device provided by the present invention. Specific embodiments
[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0023] Figure 1 It is a schematic flowchart of a parking recognition method based on a multimodal large model provided by the present invention. As Figure 1 shown, the method includes the following: Step 110: Receive the parking image of the target vehicle sent by the user terminal; Step 120: Based on the multimodal large model, pre-judge the parking image to determine that the parking image meets the preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area; Step 130: Based on the parking image, determine the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line; Step 140: Based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, determine whether the target vehicle parks regularly.
[0024] The execution subject of the parking recognition method based on the multi-modal large model provided by the present invention can be an electronic device, a component in the electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), or a personal computer (PC), etc. The present invention does not make specific limitations.
[0025] Taking a computer executing the parking recognition method based on the multi-modal large model provided by the present invention as an example, the technical solution of the present invention will be described in detail below.
[0026] In step 110, a parking image of a target vehicle sent by a user terminal is received. The parking image is an image taken based on a use terminal after the vehicle is parked and returned. The taken parking image is used to determine whether the target vehicle parks according to parking regulations.
[0027] In step 120, based on the multi-modal large model, the parking image is pre-judged to determine that the parking image meets preset conditions, and the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area.
[0028] The multi-modal large model is an artificial intelligence model that can process and understand different types of data (such as text, images, sounds, etc.). When processing a parking image, such a model can analyze the image content and possible attached text information (such as license plate number) at the same time.
[0029] Based on the multi-modal large model, judgments are respectively made on the preset conditions including that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area. The preliminary judgment on whether the parking image is compliant is realized.
[0030] Specifically, the multi-modal large model can first identify the vehicle in the parking image, and determine whether the vehicle is completely displayed in the image based on whether the vehicle part exceeds the image boundary and whether the parking image contains the vehicle.
[0031] When it is determined that the target vehicle is completely displayed in the image, the multi-modal large model identifies the vehicle number in the parking image to determine whether the vehicle in the parking image is the target vehicle.
[0032] When it is determined that the parking image contains the number of the target vehicle, the multi-modal large model determines that the vehicle is in the return area based on identifying specific landmarks or signs in the parking lot or according to the position information of the vehicle.
[0033] If it is determined that the parking image meets the preset conditions, the further recognition process of whether the vehicle is parked regularly can continue for the parking image. If the parking image does not meet the conditions, the prediction result will be fed back to the user and improvement suggestions will be provided.
[0034] In step 130, based on the parking image, determine the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line.
[0035] Specifically, computer vision techniques such as convolutional neural networks (CNNs) or object detection algorithms can be used to identify and locate the target vehicle in the image. Through image segmentation techniques such as Mask R-CNN, the precise contour of the vehicle is extracted.
[0036] Utilize the contour information of the vehicle to calculate the angle between the vehicle and the horizontal line of the image, thereby determining the placement angle of the vehicle.
[0037] When the target vehicle is detected in the image, identify the positions of other surrounding vehicles. Use the bounding box coordinates obtained from object detection to calculate the minimum distance between the target vehicle and the bounding boxes of other vehicles. Convert the pixel distance in the image to the distance in the real world to determine the distance between the target vehicle and other vehicles.
[0038] Use image processing techniques such as edge detection or Hough transform to identify the parking line in the image. Calculate the closest distance between the bounding box of the target vehicle and the parking line. According to the parking rules, determine whether the distance between the target vehicle and the parking line is within the allowed range.
[0039] In step 140, based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, determine whether the target vehicle is parked regularly.
[0040] Set the standards for the vehicle placement angle, vehicle spacing, and the distance between the vehicle and the parking line according to the parking rules. Compare the calculated angles and distances with the preset rules to determine whether the vehicle is parked regularly.
[0041] Optionally, the implementation process based on the multi-modal large model can be as Figure 2 shown in the schematic diagram of the implementation process based on the multi-modal large model provided by the present invention.
[0042] The user takes a picture of the vehicle at a fixed angle, and the image should include the license plate, the parking posture of the vehicle, and the ground parking line; The multi-modal large model analyzes the image, identifies the license plate information, determines whether the vehicle is the vehicle ridden by the user, and completes the first judgment; The multi-modal large model further identifies the ground markings to determine whether the vehicle is parked within the specified area. The specified area includes parking markings and the parking of similar vehicles; Determine whether the parking image meets the preset conditions through the multi-modal large model, provide feedback according to the output result of the multi-modal large model, and the multi-modal large model can also be optimized.
[0043] The parking recognition method based on the multi-modal large model provided by the present invention uses the parking image taken by the user terminal instead of a fixed camera for recognition, eliminating the need to install a camera on the electric bicycle. It can operate stably in complex environments such as multi-story parking lots and narrow streets, without being affected by the limited perspective of a fixed camera, reducing hardware costs and maintenance expenses. Based on the multi-modal large model, it pre-judges whether the taken parking image meets the preset conditions. After passing the pre-judgment, it further judges the key elements of the vehicle based on the parking image, realizing the automatic judgment process of whether the parking is standardized and improving the accuracy of the judgment.
[0044] In one embodiment, based on the multi-modal large model, pre-judging the parking image to determine that the parking image meets the preset conditions includes: guiding the multi-modal large model based on predefined prompt words to identify the target vehicle in the parking image, determining that the target vehicle is completely displayed in the parking image, guiding the multi-modal large model to identify the vehicle number in the parking image, determining that the vehicle number is the vehicle number of the target vehicle, and guiding the multi-modal large model to identify the surrounding information of the target vehicle to determine that the target vehicle is in the return area.
[0045] Pre-define a set of prompt words, which are related to the recognition task of the parking image and are used to guide the multi-modal large model to realize the pre-judgment process of the parking image.
[0046] Use the prompt words to guide the multi-modal large model to focus on the target vehicle in the image. Specifically, it includes: guiding the multi-modal large model to identify the target vehicle in the parking image to determine whether the target vehicle is completely displayed in the image. Specifically, the vehicle contour can be identified and it can be judged whether the vehicle is cropped by the image boundary.
[0047] Guide the multi-modal large model to identify the vehicle number in the image. Determine whether the identified vehicle number matches the number of the target vehicle.
[0048] Guide the multi-modal large model to analyze the environment around the vehicle to determine whether the vehicle is in the return area.
[0049] Based on the output of the multimodal large model, determine whether the image meets the preset conditions, that is, whether the vehicle is completely displayed, whether the vehicle number is correct, and whether the vehicle is within the return area.
[0050] In one embodiment, based on the multimodal large model, pre-judging the parking image further includes: in the case where it is determined that the parking image does not meet the preset conditions, determining photo-taking adjustment information based on the conditions not met during the pre-judging process of the multimodal large model; sending the photo-taking adjustment information to the usage terminal of the target vehicle.
[0051] After the multimodal large model analyzes the parking image, it will output a pre-judging result indicating whether the image meets the preset conditions. If the image does not meet the conditions, the model will indicate which specific conditions are not met, such as the vehicle not being completely displayed, the vehicle number being invisible, or the vehicle not being within the return area.
[0052] Analyze the reasons for not meeting the conditions. For example, the vehicle may partially exceed the image boundary, the license plate may be reflective, or the image may be blurred. And based on the analyzed reasons, determine the photo-taking adjustment information and send the photo-taking adjustment information to the usage terminal of the target vehicle, so that the user can re-obtain the adjusted parking image based on the photo-taking adjustment information.
[0053] In one embodiment, based on the parking image, determining the placement angle of the target vehicle includes: based on the parking image, extracting the contour information of the target vehicle; based on the contour information, determining the angle between the target vehicle and the reference axis; based on the angle between the target vehicle and the reference axis, determining the placement angle of the target vehicle.
[0054] The reference axis can be the center line of the parking space, the boundary line of the parking lot, or other pre-set straight lines.
[0055] Preprocess the parking image, including grayscale conversion, denoising, edge enhancement, etc., to improve the clarity of the vehicle contour in the image. Use a deep learning model (such as a convolutional neural network CNN) or traditional image processing techniques (such as Sobel operator, Canny edge detection) to detect the vehicle in the image and locate its position.
[0056] When the vehicle is detected, extract the contour information of the vehicle. This usually involves steps such as contour tracking and contour thinning to obtain the precise boundary of the vehicle's outer shape.
[0057] Based on the vehicle contour information, calculate the main axis of the vehicle. This can be achieved by analyzing the geometric characteristics of the vehicle contour (such as the minimum bounding rectangle or detecting lines using the Hough transform). The angle between the main axis of the vehicle and the reference axis can be determined by calculating the angle between the direction vectors of two line segments, thereby further determining the placement angle of the target vehicle.
[0058] In one embodiment, based on the parking image, determining the distance between the target vehicle and other vehicles and the distance between the target vehicle and the parking line includes: identifying the seat of the target vehicle in the parking image to determine the seat position information of the seat in the parking image; identifying other vehicles in the parking image to determine the position information of other vehicles in the parking image; identifying the parking line in the parking image to determine the position information of the parking line in the parking image; determining the distance between the target vehicle and other vehicles based on the seat position information and the position information of other vehicles, and determining the distance between the target vehicle and the parking line based on the seat position information and the position information of the parking line.
[0059] Specifically, computer vision technology can be used to identify the seat position of the target vehicle in the parking image. The seat is usually a fixed feature and can be used as a reference point for the vehicle position. Once the seat is detected, extract its specific position information in the image, usually in the form of pixel coordinates.
[0060] Use object detection algorithms to identify and locate other vehicles in the image. For each detected other vehicle, extract its position information in the image, including the coordinates of the bounding box.
[0061] Use image processing techniques, such as the Hough transform, to detect the parking line in the image. The parking line is usually a straight line or a curve with certain geometric characteristics. Determine the specific position of the parking line in the image and extract its pixel coordinates.
[0062] Based on the seat position information of the target vehicle and the position information of other vehicles, calculate the distance between the two. This can be achieved by calculating the minimum distance between the center point of the seat and the bounding box of other vehicles. Use the seat position information of the target vehicle and the position information of the parking line to calculate the distance between the vehicle and the parking line.
[0063] In one embodiment, the seat of the target vehicle in the parking image is recognized to determine the seat position information of the seat in the parking image, including: performing image segmentation on the parking image to determine the seat image of the target vehicle; converting the seat image into a point cloud image, and based on the density clustering method, segmenting the point cloud clusters of the point cloud image; extracting the largest cluster from the segmented point cloud clusters to determine the seat shape of the target vehicle, and based on the seat shape, determining the seat position information of the seat in the parking image.
[0064] The process of recognizing the seat can be as Figure 3 shown in the seat recognition process schematic diagram provided by the present invention.
[0065] Rough masks of the seat area can be generated through image segmentation, grayscale conversion, and binarization techniques. Point clouds are constructed based on the masks and smoothed, denoised, and standardized in coordinates to optimize the point cloud data. The density clustering (DBSCAN) method is used to segment the point cloud clusters, and the largest cluster is extracted for subsequent fitting shape calculation.
[0066] Specifically, through segmentation, subsequent image processing is concentrated on the seat area, reducing unnecessary calculations, replacing the original contour recognition model, reusing depth information, and extracting contours.
[0067] The segmented seat image is converted into a point cloud image. This can be achieved by converting the position and color information of each pixel into a point in three-dimensional space.
[0068] The generated point cloud is filtered and denoised to remove outliers and unnecessary details for better recognition of the seat shape.
[0069] Based on the density clustering method (Density-Based Spatial Clustering of Applications with Noise, DBSCAN), clustering analysis is performed on the point cloud image. The DBSCAN algorithm can divide clusters according to the density of points and does not require pre-specifying the number of clusters, which is suitable for segmenting point clouds with irregular shapes.
[0070] According to the parameter settings of the algorithm (such as the neighborhood radius and the minimum number of points), the DBSCAN algorithm segments the point cloud into multiple clusters, and each cluster represents a potential object or a part of an object.
[0071] All the segmented point cloud clusters are analyzed to find the cluster with the largest number of points, which usually represents the main part of the seat. The point cloud data is extracted from the largest cluster, and the shape of the seat is determined by calculating the geometric center or other feature points of the cluster.
[0072] Determine the specific seat position information of the vehicle seat in the parking image based on the characteristic points of the seat shape, such as the center point or the boundary point.
[0073] The parking recognition device based on the multimodal large model provided by the present invention will be described below. The parking recognition device based on the multimodal large model described below can be mutually corresponding and referred to the parking recognition method based on the multimodal large model described above.
[0074] The data receiving module 410 is configured to receive the parking image of the target vehicle sent by the user terminal. The pre-judgment module 420 is configured to pre-judge the parking image based on the multimodal large model to determine that the parking image meets the preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area. The image judgment module 430 is configured to determine the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line based on the parking image. The standard parking judgment module 440 is configured to determine whether the target vehicle parks in a standard manner based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line.
[0075] The parking recognition device based on the multimodal large model provided by the present invention identifies through the parking image taken by the user terminal instead of a fixed camera, without installing a camera on the electric bicycle, and can operate stably in complex environments such as multi-story parking lots and narrow streets, without being affected by the viewing angle limitation of the fixed camera, reducing the hardware cost and maintenance expenses. Based on the multimodal large model, pre-judge whether the taken parking image meets the preset conditions. After the pre-judgment passes, further judge the key elements of the vehicle based on the parking image, realizing the automatic judgment process of whether the parking is standard and improving the accuracy of the judgment.
[0076] In one embodiment, the pre-judgment module 420 is specifically configured to: The pre-judgment of the parking image based on the multimodal large model to determine that the parking image meets the preset conditions includes: Based on the predefined prompt words, guide the multimodal large model to identify the target vehicle in the parking image, determine that the target vehicle is completely displayed in the parking image, and guide the multimodal large model to identify the vehicle number in the parking image, determine that the vehicle number is the vehicle number of the target vehicle, and guide the multimodal large model to identify the surrounding information of the target vehicle to determine that the target vehicle is in the return area.
[0077] In one embodiment, the pre-judgment module 420 is also specifically configured to: Based on the multimodal large model, predicting the parking image further includes: In the case where it is determined that the parking image does not meet the preset conditions, determining photographing adjustment information based on the conditions not met during the prediction process of the multimodal large model; Sending the photographing adjustment information to the usage terminal of the target vehicle.
[0078] In one embodiment, the image judgment module 430 is specifically configured to: Based on the parking image, determining the placement angle of the target vehicle, including: Based on the parking image, extracting the contour information of the target vehicle; Based on the contour information, determining the angle between the target vehicle and the reference axis; Based on the angle between the target vehicle and the reference axis, determining the placement angle of the target vehicle.
[0079] In one embodiment, the image judgment module 430 is further specifically configured to: Based on the parking image, determining the distance between the target vehicle and other vehicles and the distance between the target vehicle and the parking line, including: Identifying the seat of the target vehicle in the parking image to determine the seat position information of the seat in the parking image; Identifying other vehicles in the parking image to determine the position information of other vehicles in the parking image; Identifying the parking line in the parking image to determine the parking line position information of the parking line in the parking image; Based on the seat position information and the position information of other vehicles, determining the distance between the target vehicle and other vehicles, and based on the seat position information and the parking line position information, determining the distance between the target vehicle and the parking line.
[0080] In one embodiment, the image judgment module 430 is further specifically configured to: Identifying the seat of the target vehicle in the parking image to determine the seat position information of the seat in the parking image, including: Performing image segmentation on the parking image to determine the seat image of the target vehicle; Converting the seat image into a point cloud image, and based on the density clustering method, segmenting the point cloud clusters of the point cloud image; Extracting the largest cluster from the segmented point cloud clusters, determining the seat shape of the target vehicle, and based on the seat shape, determining the seat position information of the seat in the parking image.
[0081] Figure 5 An entity structure diagram of an electronic device is exemplified, as Figure 5 shown. The electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a parking recognition method based on a multi-modal large model. The method includes: receiving a parking image of a target vehicle sent by a user terminal; Based on the multi-modal large model, pre-judging the parking image to determine that the parking image meets preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area; Based on the parking image, determining the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line; Based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, determining whether the target vehicle parks in a standard manner.
[0082] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0083] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the parking recognition method based on the multi-modal large model provided by the above-mentioned various methods. The method includes: receiving a parking image of a target vehicle sent by a user terminal; Based on the multimodal large model, predict the parking image and determine that the parking image meets the preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area; Based on the parking image, determine the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line; Based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, determine whether the target vehicle parks in a standard manner.
[0084] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the parking recognition method based on the multimodal large model provided by the above methods. The method includes: receiving a parking image of a target vehicle sent by a user terminal; Based on the multimodal large model, predict the parking image and determine that the parking image meets the preset conditions, where the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in the return area; Based on the parking image, determine the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line; Based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, determine whether the target vehicle parks in a standard manner.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A parking recognition method based on a multimodal large model, characterized in that: include: Receiving a parking image of a target vehicle sent by a user terminal; Based on the multimodal large model, the parking image is pre-judged to determine that the parking image meets preset conditions, wherein the preset conditions include that the vehicle is completely displayed in the image, the image contains the vehicle number, and the vehicle is in a vehicle return area; Based on the parking image, determining the placement angle of the target vehicle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line; Based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line, it is determined whether the target vehicle is parked in a standardized manner.
2. The parking recognition method based on multimodal large model according to claim 1, characterized in that: Pre-judging the parking image based on the multimodal large model to determine that the parking image meets the preset conditions includes: Based on the predefined prompt words, the multimodal large model is guided to identify the target vehicle in the parking image, and determine that the target vehicle is completely displayed in the parking image; the multimodal large model is guided to identify the vehicle number in the parking image, and determine that the vehicle number is the vehicle number of the target vehicle; the multimodal large model is guided to identify the surrounding information of the target vehicle, and determine that the target vehicle is in the vehicle return area.
3. The parking recognition method based on multimodal large model according to claim 1, characterized in that: The prejudgment of the parking image based on the multimodal large model further includes: In the case where it is determined that the parking image does not satisfy the preset condition, determining photographing adjustment information based on the unsatisfied condition in the multimodal large model pre-judgment process; Send the photographing adjustment information to the user terminal of the target vehicle.
4. The parking recognition method based on multimodal large model according to claim 1, characterized in that: Determining a placement angle of the target vehicle based on the parking image includes: Based on the parking image, extracting contour information of the target vehicle; Based on the profile information, determining an angle between the target vehicle and a reference axis; Based on the angle between the target vehicle and a reference axis, a placement angle of the target vehicle is determined.
5. The parking recognition method based on multimodal large model according to claim 1, characterized in that: Determining the distance between the target vehicle and other vehicles and the distance between the target vehicle and a parking line based on the parking image includes: Identifying a seat of a target vehicle in the parking image and determining seat position information of the seat in the parking image; Identify other vehicles in the parking image and determine position information of the other vehicles in the parking image; Identify the parking line in the parking image and determine the parking line position information of the parking line in the parking image; Based on the seat position information and the other vehicle position information, the distance between the target vehicle and the other vehicles is determined, and based on the seat position information and the stop line position information, the distance between the target vehicle and the stop line is determined.
6. The parking recognition method based on multimodal large model according to claim 5, characterized in that: The identifying the seat of the target vehicle in the parking image and determining the seat position information of the seat in the parking image includes: Performing image segmentation on the parking image to determine a seat image of the target vehicle; Converting the seat image into a point cloud image, and segmenting the point cloud image into point cloud clusters based on a density clustering method; A maximum cluster is extracted from the segmented point cloud clusters, a seat shape of the target vehicle is determined, and based on the seat shape, seat position information of the seat in the parking image is determined.
7. A parking recognition device based on a multimodal large model, characterized in that: include: A data receiving module, used for receiving a parking image of a target vehicle sent by a user terminal; A prejudgment module, configured to prejudge the parking image based on a multimodal large model to determine whether the parking image meets preset conditions, wherein the preset conditions include that the vehicle is fully displayed in the image, the image contains the vehicle number, and the vehicle is in a vehicle return area; An image determination module, for determining, based on the parking image, a placement angle of the target vehicle, a distance between the target vehicle and other vehicles, and a distance between the target vehicle and a parking line; The standard parking judgment module is used to determine whether the target vehicle is parked in a standard manner based on the placement angle, the distance between the target vehicle and other vehicles, and the distance between the target vehicle and the parking line.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the parking recognition method based on the multimodal large model is implemented as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the parking recognition method based on a multimodal large model as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the parking recognition method based on a multimodal large model as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Digital video monitoring early warning system and method based on smart city
CN120580649A