Self-service car washing article intelligent positioning and anti-lost system and method
By combining visible light and near-infrared image feature analysis, inter-frame feature correlation and trajectory manifold analysis, a precise monitoring area is delineated and the occupancy confidence is calculated, which solves the problem of misjudgment of car wash items under foam coverage and realizes accurate positioning and reliable management of items in self-service car washes.
Patent Information
- Application Number
- CN202511640429.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-11-11
AI Technical Summary
In self-service car washes, foam dripping or splashing during foam spraying makes it difficult for traditional visual algorithms to identify the outline features of the items being washed, leading to misjudgments of lost items and disrupting the user's car wash process.
The benchmark analysis unit acquires image data of the target car wash items, and combines visible light and near-infrared image feature analysis to generate contour feature matrix and surface texture data cluster. The silent anchor position is determined by the inter-frame feature correlation and trajectory manifold analysis of the item tracking unit, the monitoring unit delineates the precise monitoring area, the item recognition unit calculates the occupancy confidence, and the positioning warning unit makes adaptive decisions to avoid misjudgment under foam coverage.
Accurately identify the location of items for car washing, reduce false alarm rates, ensure a normal car washing process for users, and improve the intelligence and reliability of item management in self-service car wash scenarios.
Smart Images

Figure CN121095889B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of car wash item positioning technology, specifically to a smart positioning and anti-loss system and method for self-service car wash items. Background Technology
[0002] Currently, in self-service car washes, the location and anti-loss management of washing and care items are usually achieved by deploying panoramic monitoring cameras in the car wash area to cover the entire car wash area. Combined with computer vision technology, the system can realize the early warning function for lost items through image recognition and dynamic tracking of washing and care items.
[0003] However, the above-mentioned positioning scheme still has the following defects in practical application: When the user sprays foam on the vehicle, some foam will drip or splash from the car body to the ground. The high-density accumulation of foam may completely cover the small car wash brush placed on the ground. Due to the influence of foam obstruction, traditional visual algorithms have difficulty extracting the outline features and key identification information of the target car wash items from the image of the covered area, resulting in misjudgment of lost items and triggering alarms, interfering with the user's normal car wash process. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a self-service car wash item intelligent positioning and anti-loss system and method, which solves the aforementioned problems.
[0005] The above-mentioned technical objective of the present invention is achieved through the following technical solution:
[0006] A self-service car wash item positioning and anti-loss system includes:
[0007] The benchmark analysis unit is used to acquire image data of the target car wash items in the target area, perform feature analysis on the preprocessed image data, and obtain the benchmark feature vector of the target car wash items. The target area is: the self-service car wash area.
[0008] The item tracking unit is used to acquire video streams of the target area in real time, and continuously track the target car wash items in the video stream based on the reference feature vector. When it is detected that the target car wash item has been released by the user and its position is on the ground in consecutive frames, the silent anchor position of the target car wash item is determined.
[0009] The monitoring unit is used to define a precise monitoring area for the target car wash items, centered on the silent anchor position, and generate a domain image base.
[0010] The item recognition unit continuously compares the video stream with the domain image substrate. When the visual features of the target car wash item in the precise monitoring area are covered by foam, the morphological features of the covering foam are identified and calculated to generate the occupancy confidence score. The occupancy confidence scores of multiple frames are integrated to obtain the occupancy confidence score history sequence.
[0011] The positioning and early warning unit is used to calculate the historical sequence of occupancy confidence to obtain an adaptive decision threshold. The occupancy confidence is compared with the adaptive decision threshold to obtain an early warning instruction.
[0012] Furthermore, feature analysis is performed on the preprocessed image data to obtain the baseline feature vector of the target car wash items, including:
[0013] Image data is divided into visible light images and near-infrared images;
[0014] The near-infrared image data is processed in layers to generate a contour feature matrix;
[0015] Feature extraction is performed on the visible light image of the image data to generate surface texture data clusters;
[0016] Abnormal data points in the contour feature matrix and surface texture data cluster are identified and removed to generate a feature set;
[0017] The contour features and texture features in the feature set are fused to generate a baseline feature vector for tracking target car wash items.
[0018] Furthermore, the target car wash item is continuously tracked in the video stream based on the reference feature vector. When it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined, including:
[0019] Inter-frame feature association is performed on consecutive video frames of the video stream to generate a cross-frame topology sequence containing inter-frame position association information;
[0020] Based on cross-frame topology sequences, the baseline feature vector is dynamically matched with the item features in each frame to generate a trajectory manifold representing the continuous positional changes of the target car wash items.
[0021] The positional change trend in the trajectory manifold and the pixel features around the target car wash item are analyzed to identify the contact features between the user's hand and the target car wash item, and to generate interaction state variables.
[0022] Furthermore, the target car wash item is continuously tracked in the video stream based on the reference feature vector. When it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined, which also includes:
[0023] Based on the transition of the interactive state variable, the item position data of the 30 frames before the transition in the trajectory manifold are extracted to generate stable boundary values;
[0024] Based on stable boundary values, the actual deviation of the target car wash item position in multiple consecutive frames after the interaction state variable jump is calculated, and the trajectory convergence is generated.
[0025] The trajectory convergence is compared with the stable boundary value to generate the silent anchor position of the target car wash item.
[0026] Furthermore, using the silent anchor point as the center, a precise monitoring zone is defined for the target car wash items, generating a domain image base, including:
[0027] Obtain the three-dimensional dimensions of the target car wash item. Using the static anchor as the core, combine the three-dimensional dimensions to calculate the spatial dimension range covering the shape of the target car wash item and generate anchor-related dimension parameters.
[0028] Based on the anchor location correlation dimension parameters, the diffusion range of water flow and foam splash within the target area is analyzed, and a field adaptation buffer factor is generated.
[0029] Centered on the silent anchor, the boundary coordinate range of the precise monitoring area is calculated by combining the anchor-related dimension parameters and the field adaptation buffer factor, and the domain boundary constraint set is generated.
[0030] Based on the domain boundary constraint set, the corresponding region of the image is extracted from the video stream image at the current moment to generate the domain image basis.
[0031] Furthermore, the video stream is continuously compared with the domain image substrate. When the visual features of the target car wash item within the precise monitoring area are identified as being covered by foam, the morphological features of the covering foam are identified and calculated to generate occupancy confidence, including:
[0032] Based on the region covered by foam and the domain image substrate of 5 consecutive frames in the video stream, the deformation recovery rate of the foam under external force is calculated, and the rheological response factor is generated.
[0033] Based on the contour feature matrix, the contour boundary of the contact area between the target car wash item and the ground is extracted, and the actual contact range is calculated to generate the interface conformal quantity.
[0034] The hardness of the target car wash item is obtained, and the foam energy conduction loss when the target car wash item comes into contact with the foam is calculated by combining the interface conformal amount and the rheological response factor, and the mediated dissipation rate is generated.
[0035] Based on the grayscale value changes in the bubble region of the video stream, the grayscale attenuation rate from the edge to the center is calculated according to spatial coordinates to generate a bubble grayscale matrix.
[0036] The mediated dissipation rate is matched with the foam grayscale matrix to generate the entity contribution value;
[0037] Based on the video stream, the interference of environmental factors in the target area is analyzed to generate environmental modal correction. Then, the entity contribution value and the environmental modal correction are fused to generate the occupancy confidence.
[0038] Furthermore, the occupancy confidence scores from multiple frames are integrated to obtain a historical sequence of occupancy confidence scores, including:
[0039] The occupancy confidence of 20 consecutive frames is collected according to the video stream frame rate, and the rheological response factor and interface conformal quantity corresponding to each frame are recorded to form a time-physical correlation dataset.
[0040] By analyzing the rheological response factors of the first 10 frames in the time-to-physical correlation dataset, the entity presence threshold is obtained;
[0041] Extract the interface conformal quantity of the target car wash items when they are normally placed from the time-series-physical association dataset, and generate the normal threshold of the contact area.
[0042] Data screening of the time-series-physical association dataset is performed based on the entity presence threshold and the normal contact area threshold to generate a historical sequence of occupancy confidence.
[0043] Furthermore, the adaptive decision threshold is calculated from the historical sequence of occupancy confidence, including:
[0044] Perform distribution analysis on the historical sequence of occupancy confidence to generate sequence distribution characteristic values;
[0045] Extract the foam region of the precise monitoring area in the video stream to obtain the current foam coverage intensity of the target area, and calculate the correlation between the current foam coverage intensity and the sequence distribution feature value to generate a scene adaptation correction factor;
[0046] The adaptive decision threshold is obtained by fusing the sequence distribution feature value with the scene adaptation correction factor.
[0047] Furthermore, the occupancy confidence level is compared with the adaptive decision threshold to obtain early warning instructions, including:
[0048] If the occupancy confidence level is greater than or equal to the adaptive decision threshold, then the target car wash item is determined to exist, and no alarm is triggered.
[0049] If the occupancy confidence level is less than the adaptive decision threshold, the item is determined to be suspected of being lost, triggering collaborative analysis.
[0050] Furthermore, the intelligent positioning and anti-loss method for self-service car wash items, applied to the aforementioned intelligent positioning and anti-loss system for self-service car wash items, includes:
[0051] Step S1: Obtain image data of the target car wash items in the target area, perform feature analysis on the preprocessed image data, and obtain the baseline feature vector of the target car wash items. The target area is: self-service car wash area.
[0052] Step S2: Acquire the video stream of the target area in real time, and continuously track the target car wash item in the video stream based on the reference feature vector. When it is recognized that the target car wash item has been released by the user and its position is on the ground in consecutive frames, determine the silent anchor position of the target car wash item.
[0053] Step S3: Using the silent anchor as the center, define a precise monitoring area for the target car wash items and generate a domain image base;
[0054] Step S4: Continuously compare the video stream with the domain image substrate. When the visual features of the target car wash item in the precise monitoring area are identified as being covered by foam, the morphological features of the covering foam are identified and calculated to generate the occupancy confidence score. The occupancy confidence scores of multiple frames are then integrated to obtain the occupancy confidence score history sequence.
[0055] Step S5: Calculate the historical sequence of occupancy confidence to obtain the adaptive decision threshold, compare the occupancy confidence with the adaptive decision threshold, and obtain the warning instruction.
[0056] In summary, the present invention has the following main beneficial effects:
[0057] Through multi-unit collaborative operation, the benchmark analysis unit extracts features by combining visible light and near-infrared images. The contour feature matrix generated by the near-infrared image can penetrate foam, avoiding the loss of contour information due to foam occlusion in traditional algorithms. The object tracking unit accurately determines the user's grip or release state on the car wash brush through inter-frame feature correlation, trajectory manifold generation, and interaction state analysis. It determines the silent anchor position by combining stable boundary values and trajectory convergence, ensuring accurate positioning of the static location of the object. The monitoring unit generates a precise monitoring area and domain image base centered on the silent anchor position, combining the three-dimensional dimensions of the car wash brush with water flow and foam diffusion range, focusing on key monitoring areas. To reduce irrelevant interference, the item recognition unit calculates the occupancy confidence level using multiple parameters, including rheological response factor, interface conformal quantity, mediated dissipation rate, foam grayscale matrix, and environmental modal correction, when the item is covered by foam. After multi-frame integration, a stable historical sequence of occupancy confidence level is obtained. The positioning and early warning unit generates an adaptive decision threshold based on the historical sequence of occupancy confidence level and performs dual verification by combining collaborative analysis. An alarm is triggered only when the item is actually lost, which solves the problem of misjudgment caused by the complete coverage of the car wash brush by foam, reduces the false alarm rate, ensures that the user's normal car wash process is not disturbed, and improves the intelligence and reliability of item management in self-service car wash scenarios. Attached Figure Description
[0058] Figure 1 This is a schematic diagram of the self-service car wash item intelligent positioning and anti-loss system of the present invention;
[0059] Figure 2 This is a flowchart of the self-service car wash item intelligent positioning and anti-loss method of the present invention. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] refer to Figure 1 and Figure 2 A self-service car wash item intelligent positioning and anti-loss system includes:
[0062] The benchmark analysis unit is used to acquire image data of the target car wash items in the target area, perform feature analysis on the preprocessed image data, and obtain the benchmark feature vector of the target car wash items. The target car wash items are: car wash brushes, and the target area is: self-service car wash area.
[0063] The item tracking unit is used to acquire video streams of the target area in real time and continuously track the target car wash items in the video stream based on the reference feature vector; when it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined.
[0064] The monitoring unit is used to define a precise monitoring area for the target car wash items, centered on the silent anchor position, and generate a domain image base.
[0065] The item recognition unit continuously compares the video stream with the domain image substrate. When the visual features of the target car wash item in the precise monitoring area are covered by foam, the morphological features of the covering foam are identified and calculated to generate the occupancy confidence score. The occupancy confidence scores of multiple frames are integrated to obtain the occupancy confidence score history sequence.
[0066] The positioning and early warning unit is used to calculate the historical sequence of occupancy confidence to obtain an adaptive decision threshold. The occupancy confidence is compared with the adaptive decision threshold to obtain an early warning instruction.
[0067] The benchmark analysis unit extracts features by combining near-infrared and visible light images. Near-infrared light can penetrate foam to stably obtain the outline of the car wash brush. The item tracking unit accurately determines the static anchor position of the item through interactive state variables and trajectory convergence, avoiding position judgment errors. The monitoring unit delineates a precise monitoring area by combining the three-dimensional size of the car wash brush with water flow and foam diffusion range. The item recognition unit calculates the occupancy confidence through multiple parameters such as foam deformation recovery speed, contact range, and grayscale attenuation, eliminating interference from simple foam accumulation. The positioning and early warning unit performs secondary verification by combining user hand-related analysis. Thus, it can stably identify the presence of the car wash brush under foam coverage, prevent item loss, reduce false alarm rate, and ensure the normal car wash process for users.
[0068] In one embodiment, feature analysis is performed on the preprocessed image data to obtain the baseline feature vector of the target car wash item, including:
[0069] Image data is divided into visible light images and near-infrared images;
[0070] The near-infrared image data is processed in layers to generate a contour feature matrix. Specifically, the preprocessed near-infrared image is used as an independent processing layer. For this layer, the gray-level difference (i.e., gradient magnitude) between each pixel and its neighboring pixels is calculated separately. Then, the gray-level mean and standard deviation of all pixels in this layer are calculated. The gray-level mean + 1.5 times the standard deviation is used as the high threshold, and the gray-level mean minus 0.5 times the standard deviation is used as the low threshold. For each layer, pixels with gradient magnitude > high threshold are retained, and pixels with gradient magnitude < low threshold are removed. Pixels between the high and low thresholds and connected to pixels with high thresholds are also retained to obtain the edge pixels of each layer.
[0071] For each layer's edge pixels, a 3×3 pixel square structuring element is used as a template. The center of the template traverses each edge pixel, marking pixels not marked as edges within the template's coverage area as edges, thus completing edge breaks and obtaining a complete initial edge map for each layer. The edges must be continuous without breaks. For each coordinate point in the initial edge map, if the point is an edge, it is marked as 1; otherwise, it is marked as 0. Then, all the 0s and 1s of the coordinate points are arranged in the order of (x, y) to form a two-dimensional contour feature matrix.
[0072] Feature extraction is performed on the visible light image of the image data to generate surface texture data clusters. Specifically, this includes: for the preprocessed visible light image, a 5×5 pixel sliding window is used to successively cover the entire image, with each window corresponding to a unique coordinate; the grayscale mean of the image is subtracted by 0.8 times the standard deviation, and the resulting value is used as the grayscale threshold; pixels with grayscale values less than the grayscale threshold are identified as brush area pixels, and the proportion of brush area pixels in each window to the total number of pixels in the window (25) is calculated, which is the density value corresponding to the window; then, the density value of each window is replaced with the average density values of itself and the surrounding 8 windows, and a two-dimensional table is built according to the image pixel coordinates, where the number of rows and columns of the table is consistent with the number of rows and columns of the image pixels, each row of the table corresponds to the image y-axis coordinate, and each column corresponds to the image x-axis coordinate; the density value of each sliding window is filled into the cell in the table corresponding to the pixel coordinate of the window center, forming a brush density distribution table;
[0073] In the contour feature matrix, for a pixel marked as 1, adjacent (up, down, left, right, and diagonally adjacent) 1-value pixels are grouped into the same edge segment, and isolated 1-value pixels are removed. An isolated pixel is one that is itself a 1-value pixel and has no adjacent 1-value pixels, resulting in multiple continuous edge segments. The total number of pixels in each continuous edge segment is calculated and the aspect ratio of the area enclosed by the edge segment is calculated. Only edge segments with a total number of pixels ≥ 50 and an actual length ≥ 5cm are retained, as well as edge segments with an aspect ratio ≥ 8:1, corresponding to a long strip shape. Other edge segments are removed.
[0074] For the retained edge segments, the area enclosed by them is determined as the handle area. If multiple edge segments meet the criteria and can be used as handle areas, the area enclosed by the edge segments connected to the area with a density value <10% in the bristle density distribution table is taken as the handle area. For the handle area, a 3×3 pixel neighborhood is selected centered on the gray value of each pixel. The gray values of the 8 pixels in the neighborhood are compared with the gray value of the center pixel. If the gray value of the neighboring pixels is ≥ the gray value of the center pixel, it is recorded as 1; if the gray value of the neighboring pixels is < the gray value of the center pixel, it is recorded as 0. An 8-bit binary number is formed to represent the grayscale mode. The frequency of occurrence of all grayscale modes in the entire handle area is counted, and this frequency is used as the texture feature value.
[0075] The density values at each coordinate in the bristle density distribution table are associated with the texture feature values at the corresponding coordinates in the handle area, and arranged in coordinate order to form a surface texture data cluster.
[0076] Abnormal data points in the contour feature matrix and surface texture data cluster are identified and removed to generate a feature set. Specifically, for the contour feature matrix, a 3×3 sliding window is used to successively cover the contour feature matrix. Each window contains 9 pixels. The number of pixels with a value of 1 in each window is divided by the total number of pixels in the window to obtain the ratio. If the ratio is <10%, the central pixel of the window is judged to be abnormal and its value is changed directly from 1 or 0 to 0. After traversing all windows, the corrected contour feature matrix is obtained.
[0077] For a surface texture data cluster, calculate the mean and standard deviation of all coordinate density values, then calculate the difference between each density value and the mean, divide the difference by the standard deviation to obtain outliers; if the outlier is >2 or the outlier is <-2 (deviation exceeds 2 times the standard deviation), remove the density value of that coordinate; count the occurrence frequency of all texture feature values, if the frequency of a texture feature value is <5 times, remove the coordinate data corresponding to that feature value;
[0078] For all data in the corrected contour feature matrix and the data in the surface texture data cluster that were not removed, the two are associated according to the same (x, y) coordinates and organized into a dataset sorted by coordinates, which is the feature set.
[0079] The contour features and texture features in the feature set are fused to generate a baseline feature vector for tracking target car wash items. Specifically, this involves: taking contour features and texture features at the same coordinate in the feature set, where contour features include 0 and 1, and texture features include density value and texture feature value frequency. The weight of contour features is set to 0.4, density value to 0.3, and texture feature value frequency to 0.3. The contour features, density values, and texture feature value frequencies are multiplied by their corresponding weights and then summed. The calculation results are normalized to 0-1 to obtain the fused value for each coordinate. The normalized fused values of all coordinates are arranged in order from left to right on the x-axis and from top to bottom on the y-axis to form a one-dimensional array, which is the baseline feature vector.
[0080] Among them, the contour features are derived from near-infrared images. Near-infrared light can penetrate the interference of foam and stably reflect the overall shape of the target car wash items. It is the core basis for target recognition and is therefore given a high weight of 0.4.
[0081] The density value and texture feature frequency both come from visible light images. The density value reflects the distribution characteristics of the bristles, and the texture frequency reflects the texture of the handle. Although they can help with identification, visible light is easily affected by foam reflection and its stability is weaker than that of near-infrared contour features. Moreover, the two correspond to different local features of the target car wash items and their effects are complementary. Therefore, they are given equal low weights of 0.3.
[0082] By using gradient magnitude thresholding and 3×3 template edge completion, the stable extraction of the complete outline of the car wash brush is ensured, avoiding interference from foam on the overall shape recognition. At the same time, outlier removal and weight optimization are performed on the surface texture data clusters extracted from the visible light image. This retains local features to assist in recognition while avoiding the influence of foam reflection on visible light through low weight. The baseline feature vector formed after the fusion of the two-dimensional features balances anti-interference ability and recognition accuracy, reducing the risk of feature loss in foam-covered scenarios.
[0083] By filtering out false edges of foam by edge segment screening, and combining the brush density distribution table to lock in the key areas of the item, and further filtering out foam interference data by removing abnormal data, the core recognition role of entity shape is strengthened during feature fusion, which effectively avoids the problem of traditional algorithms misjudging foam accumulation as lost items and ensures the accuracy of early warning.
[0084] In one embodiment, the target car wash item is continuously tracked in the video stream based on a reference feature vector; when it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined, including:
[0085] Inter-frame feature association is performed on consecutive video frames of the video stream to generate a cross-frame topological sequence containing inter-frame positional association information. Specifically, this includes: extracting the item feature vector of the target car wash item from each frame of the video stream using the same extraction method as the baseline analysis unit; for each frame t, the frame t coordinates are the mean coordinates of all edge pixels (pixels marked as 1) in the contour feature matrix of the item feature vector of that frame, and using the frame t coordinates as the coordinates of the target car wash item; calculating the cosine similarity between the item feature vectors of frame t and frame t+1, and for each frame t, retaining feature points in frame t+1 with a similarity > 0.7 as candidate matches to form a candidate matching pair list of frame t coordinates - frame t+1 coordinates - similarity; using the similarity of the matching pairs as weights, calculating the difference in the x-direction coordinates of all matching pairs. Multiplying the x-coordinate difference by the corresponding weights yields the weighted x-coordinate difference and weighted y-coordinate difference for a single matching pair. The sum of the weighted x-coordinate differences, the sum of the weighted y-coordinate differences, and the sum of the similarities for all matching pairs are calculated. Dividing the sum of the weighted x-coordinate differences by the sum of the similarities yields the x-coordinate displacement component. Dividing the sum of the weighted y-coordinate differences by the sum of the similarities yields the y-coordinate displacement component, thus providing the displacement components of the target car wash item in the x and y directions. Combining the x and y-coordinate displacement components yields the displacement vector. Concatenating the index (which is the timestamp), the coordinates of the target car wash item, and the displacement vector in chronological order creates a cross-frame topological sequence containing time, position, and motion trends.
[0086] Based on the cross-frame topological sequence, the baseline feature vector is dynamically matched with the item features in each frame to generate a trajectory manifold representing the continuous positional changes of the target car wash item. Specifically, this includes: extracting the coordinates and displacement vectors of the target car wash item in each frame from the cross-frame topological sequence; calculating the cosine similarity between the item feature vector and the baseline feature vector in that frame, retaining the frame coordinates with a similarity greater than 0.6; calculating the coordinate distance between adjacent frames, and if the coordinate distance between adjacent frames is less than 40 pixels, it is considered a continuous position; if the coordinate distance between adjacent frames is greater than 40 pixels, the new coordinates are obtained by adding the displacement vector of the current frame to the coordinates of the previous frame; and arranging the coordinates and new coordinates of all continuous positions sequentially according to the time order of the video frames to form a complete sequence, thus obtaining the trajectory manifold representing the continuous positional changes of the item.
[0087] The system analyzes the positional change trend in the trajectory manifold and the pixel features around the target car wash item to identify the contact features between the user's hand and the target car wash item, generating an interaction state quantity. Specifically, this includes: selecting three consecutive frames of the trajectory manifold, calculating the absolute value of the difference between the x-coordinate and the absolute value of the difference between the y-coordinate of frame t and frame t+1 to obtain the number of pixels moved between frames; if the number of pixels moved between frames in two consecutive sets is greater than 20, it is determined to be an active motion state; otherwise, it is determined to be an inactive motion state; extending the edge of the target car wash item by 10 edge pixels to form a 20×20 monitoring area, extracting skin-colored pixels with H values of 5-20, S values of 28-255, and V values of 50-255 in the HSV color space, and dividing the skin-colored pixels by the total pixels in the monitoring area to obtain the proportion of skin-colored pixels; when it is determined to be an active motion state or the proportion of skin-colored pixels is greater than 8%, the interaction state quantity is recorded as 1, indicating that the target car wash item is held by the user; otherwise, it is recorded as 0, indicating that the target car wash item is not held by the user.
[0088] In one embodiment, the target car wash item is continuously tracked in the video stream based on a reference feature vector; when it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined, which further includes:
[0089] Based on the transition of the interaction state variable, the item position data of the 30 frames before the transition in the trajectory manifold are extracted to generate stable boundary values. Specifically, when the interaction state variable changes from 1 to 0, the x-coordinate and y-coordinate of the target car wash item in the 30 frames before the transition time are extracted from the trajectory manifold. The mean and standard deviation of these 30 x-coordinates are calculated respectively. The two values obtained by adding and subtracting twice the standard deviation from the mean are the stable boundary in the x-direction. The mean and standard deviation of the 30 y-coordinates are calculated. The two values obtained by adding and subtracting twice the standard deviation from the mean are the stable boundary in the y-direction. The two boundary values in the x-direction and the two boundary values in the y-direction are combined together to obtain the stable boundary value.
[0090] Based on stable boundary values, the actual deviation of the target car wash item position in multiple consecutive frames after the interaction state variable jumps is calculated to generate trajectory convergence. Specifically, this includes: after the interaction state variable jumps, selecting the target car wash item coordinates for 10 consecutive frames; for each frame's x-coordinate, adding the two values of the stable boundary in the x-direction and dividing by 2 to obtain the x-boundary mean; calculating the absolute value of the x-boundary mean minus the x-coordinate of each frame to obtain the single-frame x-absolute deviation; then subtracting the minimum value from the maximum value of the x-boundary to obtain the x-boundary range; dividing the single-frame x-absolute deviation by the x-boundary range to obtain the single-frame x-deviation rate; calculating the corresponding single-frame y-deviation rate for each frame's y-coordinate using the same calculation method for the x-coordinate; adding the x and y deviation rates of each frame and dividing by 2 to obtain the single-frame actual deviation value; adding the 10 single-frame actual deviation values and dividing by 10 to obtain the trajectory convergence.
[0091] The trajectory convergence is compared with the stable boundary value to generate the silent anchor position of the target car wash item. Specifically, this includes: calculating the average x and y deviation rates of the trajectory in the 30 frames before the interaction state change, adding the two averages and dividing by 2 to obtain the baseline deviation; multiplying the baseline deviation by 0.3 to obtain the dynamic threshold; if the trajectory convergence is less than or equal to the dynamic threshold, the average x and y coordinates of the 1st to 10th frames after the change are used as the silent anchor position, representing the static coordinate position of the target car wash item; if the trajectory convergence is greater than the dynamic threshold, the 11th to 20th frames are used for calculation, and so on. Each time, starting from the next frame after the last 10 frames used in the previous calculation, 10 frames are taken consecutively for recalculation until the convergence reaches the standard, and the average coordinates are the silent anchor position. If the trajectory convergence is greater than the dynamic threshold, the maximum number of consecutive calculations is 3. If more than 3 times are performed, an alarm is triggered, and manual intervention is required.
[0092] By generating cross-frame topology sequences through inter-frame feature association, matching pairs are selected using cosine similarity, and displacement vectors are calculated to ensure the reliability of item feature matching in a foam environment. Then, a trajectory manifold is generated by combining the baseline feature vector and coordinate distance to complete the position sequence that may be broken due to foam interference. The user's holding state is then determined by the interaction state quantity. After the interaction state quantity jumps, the data of the first 30 frames is extracted to generate stable boundary values. The trajectory convergence is calculated in units of 10 frames, and the static anchor position is determined by comparing with the dynamic threshold. A maximum of 3 retries are made to ensure the validity of the result. In this way, the interference of foam on position judgment can be avoided, the static position of the item can be accurately locked, and false alarms can be reduced.
[0093] In one embodiment, a precise monitoring area is defined for the target car wash items, centered on the silent anchor point, and a domain image base is generated, including:
[0094] Obtain the three-dimensional dimensions of the target car wash item, including: the handle length of the car wash brush, the diameter of the brush bristle area, the handle width, and the bristle height. Using a silent anchor as the core, and combining the three-dimensional dimensions, calculate the spatial dimension range covering the shape of the target car wash item, generating anchor-related dimension parameters. Specifically, map the handle length of the target car wash item to the x-axis to represent the handle's extension direction; map the diameter of the brush bristle area to the y-axis to represent the lateral spread direction of the bristles; map the bristle height to the z-axis to represent the height of the bristles perpendicular to the ground; and use half of each of the handle length, brush bristle diameter, and bristle height as the baseline extension values for the x-axis, y-axis, and z-axis, respectively.
[0095] Using the silent anchor as the origin, expand along the positive and negative x, y, and z axes (a total of 8 neighborhoods), based on the reference expansion amount of the corresponding axis, so that the expansion range extends exactly from the center of the object to the edge; for example, when the handle length is 20cm, take 10cm as the reference expansion amount of the x-axis, and the x-axis range is from x0-10cm to x0+10cm. Similarly, the range of the y and z axes can be obtained, and thus the initial three-dimensional space range can be obtained.
[0096] Calculate the difference between half the handle width and half the diameter of the bristle area. If the difference is positive, it means the handle cross-section is wider than the bristle area, and the initial range of the y-axis needs to be expanded in the direction of the difference to ensure complete coverage by the handle. If the difference is negative, no expansion is needed. Calculate half the handle width as a check value, and then calculate the total length of the current range of the x-axis and y-axis respectively. Total length = maximum range value - minimum range value. If the total length of a certain axis is less than the check value, expand to both ends of that axis until the total length is not less than the check value to avoid missing coverage of the edge of the item. The final spatial range is the anchor position associated dimension parameter.
[0097] Based on the anchor position association dimension parameters, the diffusion range of water flow and foam splash within the target area is analyzed, and a field adaptation buffer factor is generated. Specifically, this includes: setting the pixel recognition range of water flow and foam in the HSV color space: H value 0-30, corresponding to the yellow hues of water and foam; S value 40-255 to ensure the exclusion of low-saturation gray noise; V value 50-255 to avoid interference from dark shadows; extracting pixels that fit the range frame by frame from nearly 100 frames of video stream in the target area and identifying them as water flow or foam pixels; marking the water flow and foam pixels in each frame, grouping adjacent (up, down, left, right, and diagonally adjacent) pixels of the same type into a connected component, calculating the straight-line distance from the edge pixel of each connected component to the silent anchor position, and taking the largest straight-line distance in each frame as the diffusion radius of water flow and foam in that frame;
[0098] Add the diffusion radii of 100 frames together, divide by 100 to get the mean radius, and then multiply the mean radius by 0.8. The result is the base buffer value. Multiplying the mean radius by 0.8 is to avoid excessive buffering leading to redundancy in the monitoring area while covering the normal diffusion range. Calculate the total length of the x-axis range and the total length of the y-axis range in the anchor position association dimension parameters, add the two total lengths together and divide by 2 to get the mean. Use 15% of this mean as the supplementary buffer value. Set the weight of the base buffer value to 0.6 and the weight of the supplementary buffer value to 0.4. Multiply the base buffer value by 0.6 and add the supplementary buffer value by 0.4 to get the field adaptation buffer factor.
[0099] Centered on the silent anchor, and combining the anchor-related dimension parameters and the field adaptation buffer factor, the boundary coordinate range of the precise monitoring area is calculated, and a domain boundary constraint set is generated. Specifically, the field adaptation buffer factor is superimposed on the x, y, and z axis values of the anchor-related dimension parameters, centered on the silent anchor. For example, if the original range of the x axis is [x1, x2], the superimposed range is [x1 - field adaptation buffer factor, x2 + field adaptation buffer factor]. The y and z axes are also superimposed in the same way as the x axis to obtain the expanded range.
[0100] Obtain the physical boundary of the target region, and compare the expanded range with the physical boundary of the target region. For example, if the expanded x-axis range is [x0-15cm, x0+15cm], and the x-axis of the target region boundary is [0cm, 200cm], then x0-15cm < 0cm, so the left boundary of the x-axis is corrected to 0cm, and the final x-axis range is [0cm, x0+15cm]. If the expanded range does not exceed the physical boundary, the original expanded range is retained. Organize the corrected x, y, z axis three-dimensional coordinate ranges into an ordered set of coordinate intervals, with each interval corresponding to the boundary constraint of an axis. This set is the domain boundary constraint set.
[0101] Based on the domain boundary constraint set, the corresponding region of the image is extracted from the video stream image at the current moment to generate the domain image basis. Specifically, this includes converting the x, y, and z axis coordinate ranges of the domain boundary constraint set into pixel coordinate ranges according to the conversion ratio between video stream pixels and physical size, such as 1 pixel = 0.5 cm; for example, if the physical range of the x-axis is [0 cm, 30 cm], the pixel range of the x-axis is 0 ÷ 0.5 = 0 pixels to 30 ÷ 0.5 = 60 pixels. The y and z axes are converted in the same way to obtain the three-dimensional pixel coordinate range.
[0102] In the current video stream frame, an image region within the three-dimensional pixel coordinate range is extracted to form an initial cropped image. For the edge pixels of the initial cropped image, that is, the edge pixels at the boundary of the cropped range, the difference in gray value between each edge pixel and its neighboring pixels is calculated. If the difference is greater than 10, the gray value of the edge pixel is adjusted to the gray value of the neighboring pixel + 10, thereby completing the gray-level discontinuity. If the difference is less than or equal to 10, the original gray value is retained, and finally an image with complete edges is obtained, which is the domain image base.
[0103] By analyzing the three-dimensional dimensions of the car wash brush and its silent anchor position, the associated dimensional parameters of the anchor position are calculated. The range of the completed area is verified by the handle width to ensure complete coverage of the item without any omissions. Then, the range of water flow and foam diffusion in nearly 100 frames of video stream are analyzed. A field adaptation buffer factor is generated by weighting the basic buffer value and the supplementary buffer value. The two are then superimposed to obtain the extended range. Combined with the physical boundary correction of the target area, a domain boundary constraint set is generated. The result is then converted into a pixel range for image cropping and grayscale fault filling to form a domain image base with complete edges. This allows for focusing on the core area of the item, eliminating irrelevant interference, and reducing the occurrence of misjudgments.
[0104] In one embodiment, the video stream is continuously compared with the domain image substrate. When the visual features of the target car wash item within the precise monitoring area are identified as being covered by foam, the morphological features of the covering foam are identified and calculated to generate a occupancy confidence level, including:
[0105] Based on the foam-covered area and the domain image substrate in 5 consecutive frames of the video stream, the deformation recovery rate of the foam under external force is calculated, and a rheological response factor is generated. Specifically, this includes: using the coordinates of pixels marked as 1 in the 2D contour feature matrix as the edge pixel coordinates of the target car wash item; extracting the edge pixel coordinates of the target car wash item from the domain image substrate; comparing the edge coordinates of the foam-covered area in the 5 consecutive frames of the video stream with the edge coordinates of the target car wash item in the domain image substrate, calculating the Euclidean distance point by point, and taking the average of the distances of all points in each frame as the single-frame offset distance; adding the single-frame offset distances of the 5 frames and dividing by 5 to obtain the total deformation value; and then... With a frame interval of 0.2 seconds, the total duration of 5 frames is 1 second. Dividing the total deformation value by 1 yields the deformation recovery rate. Multiplying the deformation recovery rate by 0.7 yields the rheological response factor. In self-service car washes, the foam is mostly neutral cleaning foam. In this application, the rheological response factor is used to assist in calculating the occupancy confidence. The core is to distinguish between two situations: foam covering the target car wash item and foam only remaining when there is no target car wash item. If the weight is too high, it may exaggerate the foam recovery ability, leading to misjudging scenarios where there are no items as having items. If the weight is too low, it will underestimate the foam recovery ability and miss scenarios where there are items. Therefore, the weight of the deformation recovery rate is set to 0.7.
[0106] Based on the contour feature matrix, the contour boundary of the contact area between the target car wash item and the ground is extracted, and the actual contact range is calculated to generate the interface conformal quantity. Specifically, this includes: extracting ground pixels with H of 0-360, S of 0-50, and V of 30-200 in the HSV color gamut from the image containing the contour feature matrix to determine the coordinate range of the ground area; traversing all edge pixels marked as 1 in the contour feature matrix and checking whether the coordinates of each edge pixel are within the coordinate range of the ground area, and those within the range are the contact edge points; grouping adjacent (up, down, left, right, and diagonally adjacent) contact edge points into the same group, with each group forming a continuous line segment, and all line segments together constituting the contact boundary between the target car wash item and the ground;
[0107] The total number of pixels within the area enclosed by the contact boundary is counted. The conversion ratio between pixels and actual area is set to 1 pixel = 0.25 cm². The actual area obtained by multiplying the total number of pixels by this ratio is the conformal amount of the interface.
[0108] The hardness of the target car wash item is obtained. Combining interface conformal coefficient and rheological response factor, the foam energy conduction loss when the target car wash item comes into contact with foam is calculated, generating the mediated dissipation rate. Specifically, the hardness of the target car wash item is divided into the base hardness of the car wash brush handle and the base hardness of the brush bristles. The weight of the handle is set to 0.65, and the weight of the bristles to 0.35. The two base hardnesses are multiplied by their respective weights to obtain the overall hardness of the target car wash item. The interface conformal coefficient is multiplied by the rheological response factor, then multiplied by 0.5, and then divided by the overall hardness to obtain the mediated dissipation rate. Since the overall hardness needs to balance the influence of the handle (high hardness) and the bristles (low hardness) on the foam energy conduction loss, if the handle weight is too high, the energy conduction effect of high hardness will be exaggerated; if the handle weight is too low, the energy conduction effect of hardness will be underestimated. Therefore, the weight of the handle is set to 0.65, and the weight of the bristles to 0.35.
[0109] Based on the grayscale value changes in the bubble region of the video stream, the grayscale attenuation rate from the edge to the center is calculated according to spatial coordinates to generate a bubble grayscale matrix. Specifically, this includes: for the edge coordinates of the bubble-covered area in the video stream, the coordinates formed by the average of all edge coordinates are used as the bubble center; with the bubble center as the origin, sampling lines are divided towards the edge at 5-pixel intervals until the edge; the coordinates and corresponding grayscale values of each sampling point are recorded point by point; for each sampling line, the grayscale value of the bubble center sampling point is subtracted from the grayscale value of the edge sampling point and then divided by the total pixel length of the sampling line to obtain the attenuation rate; finally, the coordinates, grayscale values, and attenuation rates of all sampling points are sorted from left to right on the x-axis and from top to bottom on the y-axis to form a two-dimensional bubble grayscale matrix.
[0110] The mediated dissipation rate is matched with the foam grayscale matrix to generate the entity contribution value. Specifically, this involves multiplying the decay rate of each sampling point in the foam grayscale matrix by the grayscale weight (0.3), and then multiplying it by the mediated dissipation rate to obtain the contribution value of a single sampling point; multiplying the mean of the contribution values of all sampling points by the entity correction coefficient (1.2) to obtain the entity contribution value. In this application, the foam is neutral clean foam, and the grayscale decays gradually from the edge to the center. If the weight is too high, it will increase the impact of grayscale decay on entity judgment, leading to misclassification. Thick foam is considered to indicate the presence of an item; if the weight is too low, the auxiliary role of grayscale information will be weakened, so the grayscale weight is set to 0.3; in this application, the entity contribution value needs to distinguish between foam covering the target car wash item and car wash items without a target only having foam: when there is a target car wash item, the mediated dissipation rate is higher and the grayscale decay is more regular, so the difference needs to be amplified by the entity correction coefficient; when there is no target car wash item, the energy loss is low and the grayscale decay is chaotic, so the entity correction coefficient will decrease. In order to distinguish between the two types of scenarios, the entity correction coefficient is set to 1.2.
[0111] Based on the video stream, the interference of environmental factors in the target area is analyzed to generate an environmental modality correction value. Then, the entity contribution value and the environmental modality correction value are fused to generate the occupancy confidence value. Specifically, this includes: extracting water flow pixels that conform to the HSV color gamut (H is 0-30, S is 40-255) from the video stream, selecting two consecutive frames with an interval of 0.2 seconds between the two frames; calculating the displacement distance of the same water flow pixel, dividing the displacement distance by the interval time to obtain the water flow velocity; if the water flow velocity is >5cm / s, the water flow correction coefficient is 0.15; if the water flow velocity is ≤5cm / s, the water flow correction coefficient is 0.08.
[0112] Obtain the standard grayscale value; calculate the grayscale values of all pixels in the current video stream frame, and calculate the difference between the mean and the standard grayscale value; if the difference is >30, the lighting correction coefficient is 0.12; if the difference is ≤30, the lighting correction coefficient is 0.05; add the water flow correction coefficient and the lighting correction coefficient to obtain the environmental modality correction amount; subtract the environmental modality correction amount from 1, multiply it by the entity contribution value, and then normalize the result to 0-1, which is the occupancy confidence level.
[0113] By comparing five consecutive video streams with the domain image substrate, the deformation recovery rate is calculated and a rheological response factor is generated. This distinguishes between foam-covered objects and simple foam accumulation. The interface conformal quantity is then extracted from the ground contact boundary and the mediated dissipation rate is generated by combining the comprehensive hardness of the car wash brush. This reflects the energy interaction characteristics between the object and the foam. At the same time, the grayscale attenuation law is captured by the foam grayscale matrix and matched with the mediated dissipation rate to generate the entity contribution value, which strengthens the difference between the entity and the foam. Finally, the environmental modal correction quantity is generated by combining the water flow velocity and the light grayscale difference. The occupancy confidence is obtained by fusing the data, effectively avoiding foam occlusion interference, reducing false alarms, and ensuring a smooth car wash process for users.
[0114] In one embodiment, the placeholder confidence scores of multiple frames are integrated to obtain a placeholder confidence score history sequence, including:
[0115] The occupancy confidence of 20 consecutive frames is collected according to the video stream frame rate. The rheological response factor and interface conformal quantity corresponding to each frame are recorded to form a time-physical association dataset. Specifically, the process includes: collecting 20 consecutive frames of images at a frame rate of 5 frames / second with a 0.2-second interval between each frame; and arranging the 20 frames of data sequentially according to the structure of frame index (timestamp) - occupancy confidence - rheological response factor - interface conformal quantity to form a time-physical association dataset.
[0116] The rheological response factors of the first 10 frames in the temporal-physical association dataset are analyzed to obtain the entity presence threshold. Specifically, this involves: extracting the rheological response factors of the first 10 frames from the temporal-physical association dataset, calculating the mean of the 10 factors, and using the range formed by 80% and 120% of the mean as the interval range, which is the entity presence threshold. The first 10 frames represent the stage where the target car wash item is not completely covered by foam and is in a stable state. At this time, the rheological response factors are less affected by foam, but are affected by slight water flow and minor changes in lighting, resulting in normal fluctuations within ±20%. The 80%-120% interval can completely cover this stable fluctuation range, avoiding misjudging normal fluctuations as the absence of an entity.
[0117] To extract interface conformal quantities of the target car wash items under normal placement from the temporal-physical association dataset, and generate a normal contact area threshold, the following steps are taken: extract interface conformal quantities from the first 10 frames of the temporal-physical association dataset; calculate the mean of the 10 conformal quantities, and use the interval formed by 70% and 130% of the mean as the normal contact area threshold; the first 10 frames represent the stage where the target car wash items are stably placed, and the interface conformal quantities are affected by slight deformation of the brush bristles and minor protrusions on the ground, resulting in normal fluctuations within ±30%. Therefore, the interval formed by 70% and 130% of the mean can completely cover this fluctuation, avoiding misjudging normal deformation as non-physical.
[0118] The temporal-physical association dataset is screened based on the entity presence threshold and the normal contact area threshold to generate a historical sequence of occupancy confidence. Specifically, this involves: traversing 20 frames of data in the temporal-physical association dataset, retaining frames in each frame where the rheological response factor is within the entity presence threshold and the interface conformal quantity is within the normal contact area threshold, and removing frames that exceed the corresponding thresholds; sorting the retained frames by frame index (timestamp), and arranging the corresponding occupancy confidence in sequence to form a historical sequence of occupancy confidence.
[0119] In one embodiment, the adaptive decision threshold is calculated by analyzing the historical sequence of occupancy confidence, including:
[0120] The historical sequence of occupancy confidence is analyzed for distribution to generate sequence distribution feature values. Specifically, this includes: calculating the mean of all retained frames in the historical sequence of occupancy confidence and using the mean as the central tendency feature value; calculating the standard deviation of all retained frames in the historical sequence of occupancy confidence and using the standard deviation as the dispersion feature value; and combining the central tendency feature value and the dispersion feature value to form the sequence distribution feature value.
[0121] The foam region in the precise monitoring area of the video stream is extracted to obtain the current foam coverage intensity of the target area. The current foam coverage intensity is correlated with the sequence distribution feature value to generate a scene adaptation correction factor. Specifically, this includes: extracting foam pixels that conform to HSV (H is 0-30, S is 40-255, and V is 50-255) from the precise monitoring area of the video stream, calculating the proportion of these foam pixels to the total number of pixels in the precise monitoring area, and using this proportion as the current foam coverage intensity; multiplying the current foam coverage intensity by the foam coefficient (0.4) and dividing by the dispersion feature value to obtain the scene adaptation correction factor. Among these, the foam in the car wash is neutral cleaning foam. The coverage intensity has a linear correlation with the occupancy confidence, but it is not the dominant factor. If the foam coefficient is too high, it will exaggerate the interference of the foam, causing the presence of an entity to be misjudged as the absence of an entity. If the foam coefficient is too low, it will not be able to effectively correct the judgment bias caused by the foam. Therefore, the foam coefficient is set to 0.4.
[0122] The adaptive decision threshold is obtained by fusing the sequence distribution feature value with the scene adaptation correction factor. Specifically, this involves: using the central tendency feature value of the historical occupancy confidence sequence as the base threshold; subtracting 1.5 times the dispersion feature value from the base threshold to obtain the initial threshold; and then subtracting (the scene adaptation correction factor multiplied by 0.1) from the initial threshold to obtain the adaptive decision threshold. The historical occupancy confidence sequence in this scheme is stable data after screening, and its fluctuation range due to interference from water flow, lighting, etc., is mostly within ±1.5 standard deviations. Therefore, 1.5 times the dispersion feature value can cover more than 90% of normal data fluctuations. Multiplying the scene adaptation correction factor by 0.1 is to control the correction magnitude, ensuring that the threshold adapts to the foam scene without deviating from the stable range of the historical sequence.
[0123] By analyzing the historical sequence of the occupancy confidence after screening, a sequence distribution feature value is constructed using the mean as the central tendency feature value and the standard deviation as the dispersion feature value. Then, foam pixels that conform to HSV in the precise monitoring area are extracted, and the proportion is calculated as the current foam coverage intensity. The scene adaptation correction factor is generated by combining the foam coefficient and the dispersion feature value, which can balance the impact of foam interference. The initial threshold is obtained by subtracting 1.5 times the dispersion feature value from the central tendency feature value as the base threshold. Then, the scene adaptation correction factor × 0.1 is subtracted to obtain the adaptive decision threshold. The adaptive decision threshold covers more than 90% of normal data fluctuations such as water flow and light, and adapts to the foam scene to avoid misjudging the loss of items due to foam obstruction, thus ensuring a smooth car wash process for users.
[0124] In one embodiment, the occupancy confidence level is compared with an adaptive decision threshold to obtain an early warning instruction, including:
[0125] If the occupancy confidence level is greater than or equal to the adaptive decision threshold, then the target car wash item is determined to exist, and no alarm is triggered.
[0126] If the occupancy confidence level is less than the adaptive decision threshold, the item is judged to be suspected of being lost, triggering collaborative analysis.
[0127] Collaborative analysis specifically includes:
[0128] The video stream is analyzed to obtain a collaborative identifier, specifically including: extracting skin color pixels from the video stream in HSV format (H = 5-20, S = 28-255, V = 50-255), and defining the region formed by connecting these skin color pixels as the hand region; sampling this hand region at 3×3 pixel intervals to extract 50 feature points, including 20 edge feature points and 30 texture feature points; comparing these 50 feature points with feature points with the same coordinates in the baseline feature vector and counting the number of overlaps; dividing the number of overlaps by 50 to obtain the overlap rate; if the overlap rate is >50%, the collaborative identifier is recorded as 1; if the overlap rate is ≤50% or there is no hand region, the collaborative identifier is recorded as 0.
[0129] When the collaboration flag is 1, it means that the user is holding the target car wash item, and it is determined that the target car wash item has been picked up normally, so the tracking continues;
[0130] When the collaboration flag is 0, it indicates that the user is not identified as holding the target car wash item, and an alarm is triggered.
[0131] The system uses a combination of occupancy confidence and an adaptive decision threshold for determination: when occupancy confidence is greater than or equal to the adaptive decision threshold, the presence of the car wash brush is directly confirmed; when occupancy confidence is less than the adaptive decision threshold, no alarm is triggered directly, but instead, collaborative analysis is initiated for further verification. In the collaborative analysis, skin color pixels from the HSV are first extracted to delineate the hand area, and then 20 edge feature points and 30 texture feature points are extracted at 3×3 pixel intervals. A collaborative identifier is generated by comparing the overlap rate of the feature points. If the collaborative identifier is 1, it indicates that the user is holding the car wash brush normally, and tracking continues; if the collaborative identifier is 0, an alarm is triggered. This effectively eliminates false positives caused by foam obscuring the brush, ensuring that the alarm is only triggered when the item is actually lost, thus protecting the user's car wash experience.
[0132] In one embodiment, the self-service car wash item intelligent positioning and anti-loss method is applied to the aforementioned self-service car wash item intelligent positioning and anti-loss system, including:
[0133] Step S1: Obtain image data of the target car wash item in the target area, perform feature analysis on the preprocessed image data, and obtain the baseline feature vector of the target car wash item. The target car wash item is: car wash brush, and the target area is: self-service car wash area.
[0134] Step S2: Acquire the video stream of the target area in real time, and continuously track the target car wash item in the video stream based on the reference feature vector; when it is recognized that the target car wash item has been released by the user and its position is on the ground in consecutive frames, determine the silent anchor position of the target car wash item.
[0135] Step S3: Using the silent anchor as the center, define a precise monitoring area for the target car wash items and generate a domain image base;
[0136] Step S4: Continuously compare the video stream with the domain image substrate. When the visual features of the target car wash item in the precise monitoring area are identified as being covered by foam, the morphological features of the covering foam are identified and calculated to generate the occupancy confidence score. The occupancy confidence scores of multiple frames are then integrated to obtain the occupancy confidence score history sequence.
[0137] Step S5: Calculate the historical sequence of occupancy confidence to obtain the adaptive decision threshold, compare the occupancy confidence with the adaptive decision threshold, and obtain the warning instruction.
[0138] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A self-service car wash item intelligent positioning and anti-loss system, characterized in that, include: The benchmark analysis unit is used to acquire image data of the target car wash items in the target area, perform feature analysis on the preprocessed image data, and obtain the benchmark feature vector of the target car wash items. The target area is: the self-service car wash area. The item tracking unit is used to acquire video streams of the target area in real time, and continuously track the target car wash items in the video stream based on the reference feature vector. When it is detected that the target car wash item has been released by the user and its position is on the ground in consecutive frames, the silent anchor position of the target car wash item is determined. The monitoring unit is used to define a precise monitoring area for the target car wash items, centered on the silent anchor position, and generate a domain image base. The item recognition unit continuously compares the video stream with the domain image substrate. When the visual features of the target car wash item in the precise monitoring area are covered by foam, the morphological features of the covering foam are identified and calculated to generate the occupancy confidence score. The occupancy confidence scores of multiple frames are integrated to obtain the occupancy confidence score history sequence. The positioning and early warning unit is used to calculate the historical sequence of occupancy confidence to obtain an adaptive decision threshold. The occupancy confidence is compared with the adaptive decision threshold to obtain an early warning instruction.
2. The self-service car wash item intelligent positioning and anti-loss system according to claim 1, characterized in that, Feature analysis is performed on the preprocessed image data to obtain the baseline feature vector of the target car wash items, including: Image data is divided into visible light images and near-infrared images; The near-infrared image data is processed in layers to generate a contour feature matrix; Feature extraction is performed on the visible light image of the image data to generate surface texture data clusters; Abnormal data points in the contour feature matrix and surface texture data cluster are identified and removed to generate a feature set; The contour features and texture features in the feature set are fused to generate a baseline feature vector for tracking target car wash items.
3. The self-service car wash item intelligent positioning and anti-loss system according to claim 2, characterized in that, In the video stream, the target car wash item is continuously tracked based on the reference feature vector. When it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined, including: Inter-frame feature association is performed on consecutive video frames of the video stream to generate a cross-frame topology sequence containing inter-frame position association information; Based on cross-frame topology sequences, the baseline feature vector is dynamically matched with the item features in each frame to generate a trajectory manifold representing the continuous positional changes of the target car wash items. The positional change trend in the trajectory manifold and the pixel features around the target car wash item are analyzed to identify the contact features between the user's hand and the target car wash item, and to generate interaction state variables.
4. The self-service car wash item intelligent positioning and anti-loss system according to claim 3, characterized in that, In the video stream, the target car wash item is continuously tracked based on the reference feature vector. When it is detected that the target car wash item has been released by the user and its position is on the ground for multiple consecutive frames, the silent anchor position of the target car wash item is determined, which also includes: Based on the transition of the interactive state variable, the item position data of the 30 frames before the transition in the trajectory manifold are extracted to generate stable boundary values; Based on stable boundary values, the actual deviation of the target car wash item position in multiple consecutive frames after the interaction state variable jump is calculated, and the trajectory convergence is generated. The trajectory convergence is compared with the stable boundary value to generate the silent anchor position of the target car wash item.
5. The self-service car wash item intelligent positioning and anti-loss system according to claim 4, characterized in that, Centered on the silent anchor point, a precise monitoring zone is defined for the target car wash items, generating a domain image base, including: Obtain the three-dimensional dimensions of the target car wash item. Using the static anchor as the core, combine the three-dimensional dimensions to calculate the spatial dimension range covering the shape of the target car wash item and generate anchor-related dimension parameters. Based on the anchor location correlation dimension parameters, the diffusion range of water flow and foam splash within the target area is analyzed, and a field adaptation buffer factor is generated. Centered on the silent anchor, the boundary coordinate range of the precise monitoring area is calculated by combining the anchor-related dimension parameters and the field adaptation buffer factor, and the domain boundary constraint set is generated. Based on the domain boundary constraint set, the corresponding region of the image is extracted from the video stream image at the current moment to generate the domain image basis.
6. The self-service car wash item intelligent positioning and anti-loss system according to claim 5, characterized in that, The video stream is continuously compared with the domain image substrate. When the visual features of the target car wash item within the precise monitoring area are identified as being covered by foam, the morphological features of the covering foam are identified and calculated to generate a occupancy confidence score, including: Based on the region covered by foam and the domain image substrate of 5 consecutive frames in the video stream, the deformation recovery rate of the foam under external force is calculated, and the rheological response factor is generated. Based on the contour feature matrix, the contour boundary of the contact area between the target car wash item and the ground is extracted, and the actual contact range is calculated to generate the interface conformal quantity. The hardness of the target car wash item is obtained, and the foam energy conduction loss when the target car wash item comes into contact with the foam is calculated by combining the interface conformal amount and the rheological response factor, and the mediated dissipation rate is generated. Based on the grayscale value changes in the bubble region of the video stream, the grayscale attenuation rate from the edge to the center is calculated according to spatial coordinates to generate a bubble grayscale matrix. The mediated dissipation rate is matched with the foam grayscale matrix to generate the entity contribution value; Based on the video stream, the interference of environmental factors in the target area is analyzed to generate environmental modal correction. Then, the entity contribution value and the environmental modal correction are fused to generate the occupancy confidence.
7. The self-service car wash item intelligent positioning and anti-loss system according to claim 6, characterized in that, By integrating the occupancy confidence scores from multiple frames, a historical sequence of occupancy confidence scores is obtained, including: The occupancy confidence of 20 consecutive frames is collected according to the video stream frame rate, and the rheological response factor and interface conformal quantity corresponding to each frame are recorded to form a time-physical correlation dataset. By analyzing the rheological response factors of the first 10 frames in the time-to-physical correlation dataset, the entity presence threshold is obtained; Extract the interface conformal quantity of the target car wash items when they are normally placed from the time-series-physical association dataset, and generate the normal threshold of the contact area. Data screening of the time-series-physical association dataset is performed based on the entity presence threshold and the normal contact area threshold to generate a historical sequence of occupancy confidence.
8. The self-service car wash item intelligent positioning and anti-loss system according to claim 7, characterized in that, The adaptive decision threshold is calculated by analyzing the historical sequence of occupancy confidence, including: Perform distribution analysis on the historical sequence of occupancy confidence to generate sequence distribution characteristic values; Extract the foam region of the precise monitoring area in the video stream to obtain the current foam coverage intensity of the target area, and calculate the correlation between the current foam coverage intensity and the sequence distribution feature value to generate a scene adaptation correction factor; The adaptive decision threshold is obtained by fusing the sequence distribution feature value with the scene adaptation correction factor.
9. The self-service car wash item intelligent positioning and anti-loss system according to claim 8, characterized in that, The occupancy confidence level is compared with the adaptive decision threshold to obtain early warning instructions, including: If the occupancy confidence level is greater than or equal to the adaptive decision threshold, then the target car wash item is determined to exist, and no alarm is triggered. If the occupancy confidence level is less than the adaptive decision threshold, the item is determined to be suspected of being lost, triggering collaborative analysis.
10. A method for intelligent positioning and anti-loss of items in self-service car washes, applied to the intelligent positioning and anti-loss system for self-service car washes as described in any one of claims 1-9, characterized in that, include: Step S1: Obtain image data of the target car wash items in the target area, perform feature analysis on the preprocessed image data, and obtain the baseline feature vector of the target car wash items. The target area is: self-service car wash area. Step S2: Acquire the video stream of the target area in real time, and continuously track the target car wash item in the video stream based on the reference feature vector. When it is recognized that the target car wash item has been released by the user and its position is on the ground in consecutive frames, determine the silent anchor position of the target car wash item. Step S3: Using the silent anchor as the center, define a precise monitoring area for the target car wash items and generate a domain image base; Step S4: Continuously compare the video stream with the domain image substrate. When the visual features of the target car wash item in the precise monitoring area are identified as being covered by foam, the morphological features of the covering foam are identified and calculated to generate the occupancy confidence score. The occupancy confidence scores of multiple frames are then integrated to obtain the occupancy confidence score history sequence. Step S5: Calculate the historical sequence of occupancy confidence to obtain the adaptive decision threshold, compare the occupancy confidence with the adaptive decision threshold, and obtain the warning instruction.
Citation Information
Patent Citations
System and method for analyzing mineral flotation froth size based on embedded platform
CN103442047A
Portable car washing method and portable car washing system
CN118700992A