View-based spatial shape recognition method, device, equipment, and storage medium
By obtaining multiple perspective views of the target product, extracting the structural contour map and using Hu moments and Fourier descriptors for feature extraction, combined with classification models and multi-image fusion rules, the problem of unreliable retrieval results in traditional image search engines is solved, and higher image retrieval accuracy is achieved.
Patent Information
- Application Number
- CN202210866401.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2042-07-22
AI Technical Summary
Traditional image search engines perform searches based on metadata information surrounding the image, resulting in unreliable search results that cannot meet the specific needs of users. In addition, image retrieval based on visual features ignores the relationship between the various views in the appearance patent design, reducing the accuracy of image searches.
By obtaining views of the target product from multiple perspectives, extracting the structural contour map, using Hu moment and Fourier descriptor for feature extraction, combining classification model and multi-image fusion rules, the spatial shape of the target product can be identified.
It improves the accuracy of image retrieval, provides a more reliable retrieval basis, and can more accurately identify the spatial shape of the target product.
Smart Images

Figure CN115272689B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of appearance recognition, and in particular to a view-based spatial shape recognition method, apparatus, device and storage medium. Background Art
[0002] Traditional image search engines typically index multimedia visual information based on metadata surrounding the image, such as titles and tags. However, because this textual information may be inconsistent with the visual information, it is inefficient and cannot meet user-specific requirements. Consequently, search results can be unreliable, especially for product design patents.
[0003] To avoid this problem, traditional retrieval methods based on single image input are no longer able to meet user needs. To address the issue of obtaining input images, combining feature-based image retrieval with annotation-based image retrieval has become a significant trend. Feature-based image retrieval primarily extracts visual features from images. Traditional visual features include shape, texture, color, and edge features, or a fusion of multiple features.
[0004] Design patents contain six views and an axonometric drawing of a product. Traditional visual feature extraction of each view of a design patent ignores some information and fails to fully utilize the relationships between views, significantly reducing the accuracy of image search.
[0005] In view of this, the applicant filed this application after studying the existing technology. Summary of the Invention
[0006] The present invention provides a view-based spatial shape recognition method, device, equipment and storage medium to improve the traditional image feature extraction that ignores the hidden information in the appearance patent design.
[0007] First,
[0008] An embodiment of the present invention provides a view-based spatial shape recognition method, which includes steps S1 to S6.
[0009] S1. Obtain N perspectives of the target product.
[0010] S2. Extracting N-view structural contours of the X structures of the target product based on the N-view views.
[0011] S3. Perform feature extraction based on the structure outline to obtain the Hu moments and Fourier descriptors of the N perspectives of the X structures.
[0012] S4. Obtain feature vectors of N perspectives of the X structures based on the Hu moment and Fourier descriptor.
[0013] S5. Classify the feature vectors using a classification model to obtain the plane shapes of the X structures from N perspectives.
[0014] S6. Based on the plane shape, identify the target product using multi-image fusion rules to obtain the spatial shape of each structure.
[0015] Second aspect,
[0016] An embodiment of the present invention provides a view-based spatial shape recognition device, comprising:
[0017] The view acquisition module is used to obtain views of the target product from N perspectives.
[0018] The structure contour acquisition module is used to extract the structure contour images of the target product from the N perspectives according to the views from the N perspectives.
[0019] The feature extraction module is used to extract features based on the structure outline and obtain the Hu moments and Fourier descriptors of N viewing angles of X structures.
[0020] The feature fusion module is used to obtain feature vectors of N perspectives of X structures based on Hu moments and Fourier descriptors.
[0021] The shape recognition module is used to classify the feature vectors through a classification model to obtain the plane shapes of X structures from N perspectives.
[0022] The shape fusion module is used to identify the planar shape through multi-image fusion rules to obtain the spatial shape of each structure of the target product.
[0023] Thirdly,
[0024] An embodiment of the present invention provides a view-based spatial shape recognition device, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the view-based spatial shape recognition method as described in any paragraph of the first aspect.
[0025] Fourthly,
[0026] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the view-based spatial shape recognition method as described in any paragraph of the first aspect.
[0027] By adopting the above technical solution, the present invention can achieve the following technical effects:
[0028] The spatial shape recognition method of the embodiment of the present invention can obtain the spatial shape of each structure of the target product based on N perspective views, thereby providing a new retrieval basis for image retrieval to obtain more accurate retrieval results, which has great practical significance.
[0029] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0031] Figure 1 It is a flowchart of the spatial shape recognition method provided by the first embodiment of the present invention.
[0032] Figure 2 It is a logic block diagram of the spatial shape recognition method provided by the first embodiment of the present invention.
[0033] Figure 3 It is the logic block diagram of smooth processing.
[0034] Figure 4 It is a logical block diagram obtained from the structural outline diagram.
[0035] Figure 5 This is a picture of the product before smoothing.
[0036] Figure 6 This is the product picture after smoothing.
[0037] Figure 7 It is the front view of the product and the extracted product outline.
[0038] Figure 8 It is a schematic diagram of the handle structure in the product outline diagram circled by the minimum circumscribed rectangle.
[0039] Figure 9 It is a line graph of the shower shape recognition rate under different matching coefficients.
[0040] Figure 10 It is a line graph of the accuracy of shower shape judgment under different matching coefficients.
[0041] Figure 11 It is a line graph of the recognition accuracy of different Fourier descriptor lengths.
[0042] Figure 12 It is a schematic diagram of solving the classification model.
[0043] Figure 13 It is a regular graph of multi-graph fusion rules.
[0044] Figure 14 2 is a schematic structural diagram of a spatial shape recognition device provided by a second embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0046] In order to better understand the technical solution of the present invention, the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0047] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0048] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0049] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0050] The "first" and "second" mentioned in the embodiments are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or precedence of "first" and "second" can be interchanged where appropriate. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.
[0051] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:
[0052] Example 1:
[0053] See also Figures 1 to 13 The first embodiment of the present invention provides a view-based spatial shape recognition method, which can be performed by a spatial shape recognition device. In particular, the method is performed by one or more processors in the spatial shape recognition device to implement steps S1 to S6.
[0054] S1. Obtain N perspectives of the target product. Preferably, N=3.
[0055] The embodiments of the present invention are used to extract spatial shape information that is often overlooked in design patent images, providing a retrieval basis for image retrieval and having great practical significance. The spatial shape recognition device can be an electronic device with computing capabilities, such as a portable notebook computer, desktop computer, server, smartphone, or tablet computer.
[0056] It is understandable that the images captured from different angles of a three-dimensional object may be the same or different. By shooting from three different perspectives, three views can be roughly obtained, thereby determining the basic three-dimensional shape of the three-dimensional object. In this embodiment, preferably, three views of three target products are selected as the initial data. In other embodiments, more than three views can be used to more accurately obtain the spatial shape information of the target product, and the present invention does not specifically limit this.
[0057] S2. Extracting N-view structural contours of the X structures of the target product based on the N-view views.
[0058] Specifically, to more accurately identify the spatial shape information of the target product, the various parts of the target product can be subdivided based on the target product's structural characteristics. For example, if the target product is a showerhead, X = 2, and the X structures include the showerhead structure and the handle structure. If the target product is a monitor, the X structures include the screen body structure and the base structure.
[0059] like Figures 2 to 8 As shown, based on the above embodiment, in an optional embodiment of the present invention, step S2 specifically includes steps S21 to S23:
[0060] S21. Using an image segmentation algorithm, extract images of the target product from N perspectives, obtain product images of the target product from N perspectives, and perform edge smoothing on the product images.
[0061] Step S21 extracts the Figure 7 Patent images have the following characteristics: 1. They have a single scene and clear classification, but lack a uniform color background; 2. Some images have embedded text information; and 3. They are expressed from multiple perspectives. Image segmentation can quickly and efficiently identify the main body of the appearance patent expression; breaking down the main structure and preprocessing the segmentation results facilitates subsequent feature extraction.
[0062] After the target subject is extracted from the view through the image segmentation algorithm, the edges may be not smooth. Figure 5 As shown in the figure, it is obvious that the effect of the unprocessed image edge is not good, there are redundant backgrounds that are judged as entities and the edge optimization no longer presents a step-like shape. Figure 6 As shown, the parts mistakenly regarded as foreground are removed and the effect of softening the outline is achieved.
[0063] Preferably, the image segmentation algorithm is the GrabCut algorithm. Target region segmentation is achieved using the GrabCut algorithm. The GrabCut algorithm is an interactive target segmentation method whose main idea is to map the image into an ST network graph and divide the foreground and background into regions of interest (ROI) using a manually labeled rectangular box. The area outside the rectangle is the background, and the area inside the rectangle contains both background and foreground.
[0064] The texture (color) information and boundary (contrast) information in the image are utilized, the Gaussian mixture model (GMM) is used to model the foreground and background, and the pixels in the ROI are classified using the energy function until convergence.
[0065] This algorithm is better at recognizing contours than several other segmentation algorithms. For example, the Otsu threshold segmentation algorithm is sensitive to noise, has unclear grayscale differences, and has unclear segmentation when there are overlapping grayscale values of different targets, and is not very robust.
[0066] Assume I is an image, S is a function that measures the consistency of pixel attributes, true or false, then image segmentation is to divide the image I into n regions R i (i=1,2,3,…,n), satisfying the following conditions: Condition (1): Condition (2): Condition (3): S(R i)=true,i=1,2,3,…,n;Condition (4): S(R i ∪R j ) = false. Conditions (1) and (2) mean that the entire image is divided into n regions and that the regions do not intersect with each other. Condition (3) states that if they are within the same region, all pixels within that region satisfy the general properties of that region. Condition (4) states that it is impossible for adjacent regions to simultaneously satisfy the general properties of the region and thus be combined into one region.
[0067] For grayscale images, it can be used as a grayscale value matrix to segment the region R i Each pixel uses a "transparency" value parameter γ=(γ1,γ2…γ n ), where γ j ∈(0,1), and finally output γ n = 1, that is, it is judged to be the foreground part, and the segmented foreground-background image is obtained.
[0068] Contour is the key to shape recognition. The smoothness of the contour curve will affect the accuracy of the subsequent shape feature description. The contours obtained by the traditional edge segmentation-based contour extraction algorithm are prone to missed segmentation and over-segmentation, and the curves are not smooth enough. After the target area is segmented, some images have stepped edges or redundant backgrounds and are judged as targets. Therefore, the image edges are smoothed first to remove the edge defects after segmentation and obtain a better target segmentation image. The specific processing steps are as follows: Figure 3 First, the subject image after background removal is converted from the RGB domain to the HSV domain, then a threshold is set to form a mask, and then it is bitwise XORed with the original image and morphological calculations are performed. Finally, it is determined whether the edge of the subject image is smooth. If it is smooth, the smoothing process is completed, otherwise it is smoothed again.
[0069] S22. Extract the overall outline of the target product based on the product image after edge smoothing, and obtain product outline images of the target product from N perspectives.
[0070] Specifically, according to the edge of the product image, a contour image of the product can be obtained.
[0071] S23. Extract structural contour maps of the X structures of the target product according to the product contour map, and obtain structural contour maps of the X structures of the target product from N perspectives.
[0072] Specifically, in order to better express the spatial shape information of the target product, three-view images of each structure are obtained according to the structural characteristics of the target product, so as to identify each structure separately.
[0073] Based on the above embodiment, in an optional embodiment of the present invention, step S23 includes step S231 and step S234.
[0074] S231. Based on the structural characteristics of the target product, find the boundary points of the structure in the product outline diagram to obtain the range to which the X structures belong.
[0075] S232: Identify the maximum area contour of the range, and obtain the maximum area contours of X structures.
[0076] S233: Calculate the minimum bounding rectangle of the maximum area contour, and obtain the minimum bounding rectangles of the X structures.
[0077] S234 , extracting a structural outline from the product outline according to the minimum circumscribed rectangle, and obtaining structural outlines of the target product at N viewing angles for the X structures.
[0078] Specifically, the extracted patent image contours are divided into regions to determine the contours of each structure. The main idea is to find the possible largest contour belonging to the structure based on the contour characteristics and structural characteristics of the image, and draw the minimum circumscribed rectangle to divide the region of interest of each structure. In order to obtain the contours of each structure, the preprocessing of the contours is crucial. Usually, the screening of the contours is the screening of the contour hierarchy and length. By removing the sub-contours and contours that are too short or too long, the complete external contour is obtained. Then, according to the characteristics of the patented item, the boundary points of each structure are found, the range to which the structure belongs is determined, and the minimum circumscribed rectangle of the contour with the largest area in the range is calculated. The hierarchical screening and contour extraction flow chart is shown below. Figure 4 shown.
[0079] Take the shower as an example: the improved edge is divided into regions according to the shower structure. For the water spray cover structure, the uppermost point of the edge is the horizontal coordinate of the upper left corner of the rectangle, and the vertical coordinate is the leftmost point of the outline edge. For the lower right corner, the horizontal coordinate of the lower right corner is to first find a relatively continuous straight line in the edge, that is, the edge of the handle. The position where the slope changes significantly along the straight line is the right end point of the water spray cover, and the vertical coordinate is the rightmost point of the outline edge. For the handle, its rectangle is as follows: Figure 8 After obtaining the minimum bounding rectangle, the contour image within the bounding rectangle is the contour image of the structure.
[0080] S3. Perform feature extraction based on the structure outline to obtain the Hu moments and Fourier descriptors of the N perspectives of the X structures.
[0081] Specifically, after extracting the target edges using an improved edge smoothing algorithm, the team combined the characteristics of the patent image to be identified. Considering that the image capture process, such as placement, angle, and camera height, can cause the resulting image to exhibit translation, rotation, and scaling, the extracted features should exhibit translation, rotation, and scaling characteristics. Therefore, the Hu moment and Fourier descriptor were used to perform contour recognition on the preprocessed image.
[0082] Preferably, before extracting the Fourier descriptor, the following steps are required: obtaining the length of the Fourier descriptor.
[0083] Specifically, when describing shapes using Fourier coefficients, they are rotationally and translationally invariant, and are independent of the chosen contour starting point. The low-frequency components of the Fourier descriptor better reflect the overall shape of the showerhead, while the high-frequency components better capture its detailed features. For recognition, the feature vectors must be of appropriate length. Based on these properties, the Fourier coefficients can be normalized.
[0084] Below, we take a shower as an example to illustrate. The results after the Fourier descriptor is normalized are shown in Table 1.
[0085] Table 1 Normalized Fourier descriptor
[0086]
[0087] Among them, there are 200 groups of sample data of shower head appearance patents. The experiment randomly selects 150 groups of data as training data and 50 groups of Fourier descriptor data as test data. In order to reflect the classification of shower head patent images from different perspectives by Fourier descriptors, Fourier descriptors of different lengths are used as input to SVM for recognition. The accuracy is as follows: Figure 11 shown.
[0088] Depend on Figure 11 It can be seen that the recognition accuracy reaches the highest and most stable value when the lengths of the Fourier descriptors for the front view, side view, and bottom view are 12, 12, and 16, respectively. Therefore, the lengths of the Fourier descriptors for the front view and side view are selected as 12, and the length of the Fourier descriptor for the bottom view is selected as 16.
[0089] Step S3 specifically includes step S31 and step S32.
[0090] S31 . Extracting Hu moments according to the structure outline to obtain Hu moments of N viewing angles of X structures.
[0091] Specifically, the image is recognized by the feature quantity composed of Hu moments, which is very fast, but the disadvantage is that when applied to the contour of complex textures, the corresponding position cannot be correctly framed, and the matching rate is relatively low. Hu moments are generally used to identify large objects in images. When the shape of the object is described better, the matching rate is higher. The gray value of a digital image I of size MxN at point (x, y) is f(x, y), and its (p+q) order moment m pq and center distance μ pq for:
[0092]
[0093]
[0094] Where m 00 is the 0th order moment, m 01 、m 10 is the first-order moment.
[0095]
[0096]
[0097] After normalization, the center distance is:
[0098] η pq =u pq / u r 00
[0099] in,
[0100]
[0101] The second-order and third-order normalized central moments are used to construct seven invariant moments. The invariant moment is a highly concentrated image feature that is invariant to translation, grayscale, scale, and rotation under continuous images. hu (I) is the Hu moment eigenvalue of image I, which is defined as follows:
[0102]
[0103] S32. Extract the Fourier descriptor of the length according to the structure outline, and obtain the Fourier descriptors of the N viewing angles of the X structures.
[0104] Specifically, Fourier descriptors can effectively describe contour features, and only a small number of descriptors are needed to roughly represent the entire contour. Secondly, a simple normalization operation on the Fourier descriptor renders the descriptor invariant to translation, rotation, and scale. This means it is unaffected by the contour's position, angle, or scaling within the image, making it a highly robust image feature.
[0105] As a feature parameter that describes the contour features of an image, the basic idea of the Fourier descriptor is to use the Fourier transform of the object boundary information as the shape feature, transform the contour features from the spatial domain to the frequency domain, and extract the frequency domain information as the feature vector of the image. That is, a vector is used to represent a contour, and the contour is digitized, so that different contours can be better distinguished, thereby achieving the purpose of object recognition. For a contour starting from any point (x0, y0), moving counterclockwise, it will encounter (x1, y1), (x2, y2), ..., (x K-1 ,y K-1), describing the contour in the spatial domain, with a boundary coordinate set D(k) = [x(k), y(k), k = 0, 1, 2…K-1]. Furthermore, each coordinate point is treated as a complex number as shown in the following equation.
[0106] In the complex number domain, the two-dimensional problem is simplified to a one-dimensional problem, but the boundary essence remains unchanged.
[0107] D(k)=x(k)+jy(k)
[0108] The discrete Fourier transform of the coordinate sequence of the contour curve is as follows:
[0109]
[0110] Where u = 0, 1, 2…K-1. The complex coefficient a(u) is the Fourier descriptor of the boundary.
[0111] The coefficient K of the Fourier descriptor fd (i) Perform normalization, that is, divide each amplitude by the 2-norm of a(1), ||a(1)||. This is the Fourier descriptor:
[0112]
[0113] S4. Obtain feature vectors of N perspectives of the X structures based on the Hu moment and Fourier descriptor.
[0114] In some embodiments, step S4 includes steps S41 to S43.
[0115] S41. Obtain a contour template library of the target product.
[0116] S42. Calculate matching coefficients based on the contour template library and the Hu moment to obtain matching coefficients for the N viewing angles of the X structures of the target product.
[0117] Specifically, we first classify each type of shower head according to different viewing angles and randomly collect 5 images to build a contour template library. Then, we extract the Hu moments of the three-view images of the shower head and perform similarity judgment with the corresponding viewing angle templates in the template library. The results are shown in Table 2. i ,Y i ,Z i Represent the shapes of the front view, side view, and top view respectively. When the image matching coefficient is smaller, the contours are more similar. Conversely, the larger the matching coefficient, the lower the matching degree.
[0118] Table 2 Image matching coefficients
[0119]
[0120] The image matching coefficient is set to a threshold of 0.01 to 0.2 for judgment. The image recognition of the three perspectives of the shower is analyzed, and the number of correct classifications and incorrect classifications of the three views are counted. When the matching coefficient between the image to be detected and the template library is greater than the threshold, it is considered a failure. If it is less than the threshold, it is considered a success. The shape recognition rate is as follows: Figure 9 shown.
[0121] The recognition accuracy rate is calculated for each threshold. If it belongs to the first category in the template library, it is marked as 0, otherwise it belongs to the second category and is marked as 1, with the first category being negative.
[0122] Record the TP, TN, FN, and FP values of the three-view image. The experimental results evaluation standard table is shown in Table 3. The shower shape discrimination accuracy under different thresholds is as follows: Figure 10 shown.
[0123] Table 3 Experimental results evaluation criteria
[0124]
[0125]
[0126] In the case of ideal contour extraction, as the threshold increases, the number of successful recognitions also increases, but at the same time more images are mismatched, and the recognition accuracy decreases. When the threshold is set too small, the recognition success rate is low, the accuracy is high, and the number of correctly recognized shape types is small. Figure 9 、 Figure 10 ,Comprehensively considering the successful recognition rate and accuracy, a threshold of 0.08 is selected as the research data, and the precision P, recall R and F score are used to evaluate the model performance. The experimental results are shown in Table 4.
[0127]
[0128]
[0129]
[0130] Table 4 Recognition performance evaluation statistics
[0131]
[0132] Table 4 shows that using the showerhead's outer contour as a feature based on the Hu invariant moment method can accurately capture showerhead features and achieve accurate matching. The side view achieves the best recognition results, while the bottom view performs poorly. The side view has clearer segmentation boundaries and uses the overall shape for matching. However, due to positioning errors and preprocessing of the showerhead's bottom in the bottom view, the discrimination is weaker than in the other two views.
[0133] S43. Normalize the matching coefficients and Fourier descriptors, and then fuse them serially to obtain feature vectors of N perspectives of X structures of the target product.
[0134] As you can understand, serial feature fusion is simple and computationally inefficient, making it a good choice for multi-class feature fusion where the sum of the dimensions is small. Serial fusion combines the advantages of Fourier descriptors and Hu moments to further improve recognition accuracy, which has great practical significance.
[0135] Specifically, assuming an image has n-dimensional eigenvalues α and m-dimensional eigenvalues β, the dimension of the new serial feature ω is (n+m). The Hu moment matching coefficients and Fourier descriptors of the patent image are normalized, removing outliers and evenly distributing the data. After mapping the values to the same interval, the serial features are fused and used as input to the SVM classification model.
[0136] Determine the recognition accuracy of the patent object shape. Serial feature fusion (i.e., feature vector fusion model) is shown in the following formula:
[0137]
[0138] In the formula, K(I) is the feature vector, ε1 is the matching coefficient weight, K hu (I hu ) indicates the first hu Hu moment matching coefficient, ε2 is the weight of Fourier descriptor, K fd (I fd ) is the first fd Fourier descriptors are used to obtain the final fusion features by adjusting the weight values of parameters ε1 and ε2.
[0139] S5. Classify the feature vectors using a classification model to obtain the plane shapes of the X structures from N perspectives.
[0140] In this embodiment, the classification model is an SVM multi-class classifier using a one-to-many scheme. In other embodiments, the classification model may be an existing multi-classification model. The present invention does not specifically limit the classification model.
[0141] SVM is a binary classification model whose basic model is a linear classifier with the largest margin defined in feature space. The basic idea of SVM is to find a separating hyperplane that correctly divides the training dataset and has the largest geometric margin. For linearly separable datasets, there are infinitely many such hyperplanes (i.e., perceptrons), but the separating hyperplane with the largest geometric margin is unique. For a known dataset, the support vector machine uses nonlinear mapping to map the data to a high-dimensional feature space and uses linear problem-solving methods to find the optimal regression function in this high-dimensional feature space. The linear regression function is shown below:
[0142] f(x)=(w,x)+b
[0143] In the formula, f(x) is a nonlinear mapping function, w is a weight, x is the input feature value of the i-th sample, i.e. K(i) in formula (8), and b is a threshold. Figure 12 As shown in Figure 2, the problem of solving the maximum splitting hyperplane of the SVM model can be expressed as the following constrained optimization problem.
[0144] Specifically, a grid search method was used to determine the optimal hyperparameters, the error penalty coefficient c, and the kernel function parameter r, within the model. K-fold cross-validation was then used to perform different partitions of the training set to provide a balanced estimate of the current hyperparameters. Support Vector Machines were trained on patent image samples of showerhead designs to develop a classification model for showerhead shapes from different viewpoints. This classification model was used to identify showerhead shapes, accurately recognizing the target shape from front, side, and top views.
[0145] S6. Based on the plane shape, identify the target product using multi-image fusion rules to obtain the spatial shape of each structure.
[0146] It can be understood that the multi-image fusion rule is based on the characteristics of multiple images in the appearance patent, and the target shape judgment and recognition are performed on multiple perspective images of each type of image to obtain experimental data, and the obtained experimental data are fused according to the multi-image fusion rule. Figure 13 As shown, the multi-image fusion rule includes spatial shapes corresponding to combinations of contour shapes of three different perspectives. In other embodiments, the multi-image fusion rule can be a combination of more than three views, which is not limited in the present invention.
[0147] Specifically, the outline of the target product in each perspective view is flat, and its product shape and classification variable indicators are shown in Table 5.
[0148] Table 5 Product shape and classification variables
[0149]
[0150]
[0151] Let i be an image, and judge the shape of the main view, side view and top view respectively, using X i , Y i , Z i Represents the main view, side view, and top view category labels, namely X i Y i Z i The multi-image fusion rules are as follows Figure 13 As shown, its spatial shape is S i =X i ∩Y i ∩Z i .
[0152] The spatial shape recognition method of the embodiment of the present invention can obtain the spatial shape of each structure of the target product based on N perspective views, thereby providing a new retrieval basis for image retrieval to obtain more accurate retrieval results, which has great practical significance.
[0153] Example 2
[0154] See also Figure 14 , an embodiment of the present invention provides a view-based spatial shape recognition device, which includes:
[0155] The view acquisition module 1 is used to acquire views of the target product from N perspectives.
[0156] The structure contour acquisition module 2 is used to extract the structure contour images of the target product from the N perspectives based on the views from the N perspectives.
[0157] The feature extraction module 3 is used to extract features based on the structure outline and obtain Hu moments and Fourier descriptors of N viewing angles of X structures.
[0158] The feature fusion module 4 is used to obtain feature vectors of N perspectives of X structures based on Hu moments and Fourier descriptors.
[0159] The shape recognition module 5 is used to classify the feature vectors through a classification model to obtain the plane shapes of the X structures at N viewing angles.
[0160] The shape fusion module 6 is used to identify the planar shape using multi-image fusion rules to obtain the spatial shape of each structure of the target product.
[0161] Based on the above embodiment, in an optional embodiment of the present invention, the structure profile acquisition module 2 includes:
[0162] The product image acquisition unit is used to extract images of the target product from N perspectives through an image segmentation algorithm, obtain product images of the target product from N perspectives, and perform edge smoothing on the product images.
[0163] The product outline image acquisition unit is used to extract the overall outline of the target product based on the product image after edge smoothing, and obtain product outline images of the target product from N perspectives.
[0164] The structure profile acquisition unit is used to extract the structure profiles of X structures of the target product according to the product profile, and acquire the structure profiles of the X structures of the target product from N perspectives.
[0165] Based on the above embodiment, in an optional embodiment of the present invention, the structure profile image acquisition unit includes:
[0166] The range acquisition subunit is used to find the boundary points of the structure in the product outline diagram based on the structural characteristics of the target product and obtain the range to which X structures belong.
[0167] The area recognition subunit is used to identify the maximum area contour of the range and obtain the maximum area contours of X structures.
[0168] The rectangle calculation subunit is used to calculate the minimum bounding rectangle of the maximum area contour and obtain the minimum bounding rectangle of X structures.
[0169] The structure contour acquisition subunit is used to extract the structure contour from the product contour according to the minimum circumscribed rectangle, and obtain the structure contour of the target product from N perspectives of X structures.
[0170] Based on the above embodiment, in an optional embodiment of the present invention, the feature fusion module 4 includes:
[0171] The template library unit is used to obtain the outline template library of the target product.
[0172] The coefficient acquisition unit is used to calculate the matching coefficient based on the contour template library and the Hu moment, and obtain the matching coefficient of the N viewing angles of the X structures of the target product.
[0173] The feature vector acquisition unit is used to normalize the matching coefficients and Fourier descriptors, and then serially fuse them to obtain the feature vectors of N perspectives of X structures of the target product.
[0174] Example 3:
[0175] An embodiment of the present invention provides a view-based spatial shape recognition device, comprising a processor, a memory, and a computer program stored in the memory. The computer program can be executed by the processor to implement the view-based spatial shape recognition method described in any of the sections of the first embodiment.
[0176] Example 4:
[0177] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the view-based spatial shape recognition method as described in any paragraph of Example 1.
[0178] In the several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.
[0179] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0180] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk. It should be noted that, in this article, the terms "include", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0181] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A view-based spatial shape recognition method, characterized in that: Include: Get N perspectives of the target product; Extracting N-view structural outlines of X structures of a target product according to the N-view views; The method includes: extracting images of the target product from the N perspectives respectively by using an image segmentation algorithm, obtaining product images of the target product from the N perspectives, and performing edge smoothing on the product images; Perform feature extraction based on the structure outline to obtain Hu moments and Fourier descriptors of N viewing angles of X structures; Obtaining feature vectors of N perspectives of X structures according to the Hu moment and the Fourier descriptor; Classifying the feature vectors using a classification model to obtain the plane shapes of the X structures at N viewing angles; According to the plane shape, identification is performed using multi-image fusion rules to obtain the spatial shape of each structure of the target product; According to the Hu moment and the Fourier descriptor, feature vectors of N perspectives of X structures are obtained, including: Obtain a contour template library of the target product; Calculating matching coefficients based on the contour template library and the Hu moment to obtain matching coefficients for N viewing angles of X structures of a target product; Normalizing the matching coefficients and the Fourier descriptors, and then fusing them serially to obtain feature vectors of N perspectives of X structures of the target product; The image segmentation algorithm is GrabCut algorithm; According to the product outline, extract the structural outline of X structures of the target product, and obtain the structural outline of N perspectives of the X structures of the target product, including: Based on the structural characteristics of the target product, find the boundary points of the structure in the product outline map to obtain the range to which the X structures belong; Identify the maximum area contour of the range, and obtain the maximum area contours of X structures; Calculate the minimum bounding rectangle of the maximum area contour to obtain the minimum bounding rectangles of X structures; According to the minimum circumscribed rectangle, a structural outline diagram is extracted from the product outline diagram to obtain structural outline diagrams of X structures of a target product at N viewing angles.
2. The view-based spatial shape recognition method according to claim 1, characterized in that: Extracting N-view structural outlines of X structures of a target product according to the N-view views, further comprising: Extract the overall outline of the target product from the product image after edge smoothing, and obtain the product outline images of the target product from N perspectives; According to the product outline, structural outlines of the X structures of the target product are extracted, and structural outlines of the X structures of the target product at N viewing angles are obtained.
3. The view-based spatial shape recognition method according to claim 1, characterized in that: The fusion model of the feature vector is: ; Where, is the characteristic vector, is the matching coefficient weight, Indicates the Hu moment matching coefficient, is the weight of the Fourier descriptor, For the Fourier descriptors; The classification model is an SVM multi-class classifier using a one-to-many scheme.
4. The view-based spatial shape recognition method according to claim 1, characterized in that: When the target product is a shower head, X=2, and the X structures include a shower head structure and a handle structure; N=3; the multi-image fusion rule includes spatial shapes corresponding to combinations of contour shapes from three different perspectives.
5. A spatial shape recognition device based on a view, characterized in that: Used to perform a view-based spatial shape recognition method according to any one of claims 1 to 4; The spatial shape recognition device comprises: A view acquisition module is used to obtain views of the target product from N perspectives; A structure contour acquisition module, configured to extract structure contour images of X structures of a target product from N perspectives based on the views from the N perspectives; A feature extraction module is used to extract features based on the structure outline to obtain Hu moments and Fourier descriptors of N viewing angles of X structures; A feature fusion module, configured to obtain feature vectors of N perspectives of X structures based on the Hu moment and the Fourier descriptor; a shape recognition module, configured to classify the feature vectors using a classification model to obtain the planar shapes of X structures at N viewing angles; The shape fusion module is used to identify the planar shape using multi-image fusion rules to obtain the spatial shape of each structure of the target product.
6. The view-based spatial shape recognition device according to claim 5, characterized in that: The structure profile acquisition module includes: a product image acquisition unit, configured to extract images of the target product from the N perspectives using an image segmentation algorithm, obtain product images of the target product from the N perspectives, and perform edge smoothing on the product images; A product contour image acquisition unit is used to extract the overall contour of the target product based on the product image after edge smoothing, and obtain product contour images of the target product from N perspectives; The structure outline acquisition unit is used to extract the structure outlines of the X structures of the target product according to the product outline, and acquire the structure outlines of the X structures of the target product from N perspectives.
7. A view-based spatial shape recognition device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory; the computer program can be executed by the processor to implement a view-based spatial shape recognition method as described in any one of claims 1 to 4.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the view-based spatial shape recognition method according to any one of claims 1 to 4.