A tongue feature recognition method and device, electronic equipment and storage medium

By using a cross-network model of Seg-CNN and Point-CNN to perform multi-scale operations on tongue images, the problem of unclear crack features in tongue images was solved, thus achieving accuracy and objectivity in tongue image analysis.

CN117292233BActive Publication Date: 2026-05-29CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD
Filing Date
2022-06-17
Publication Date
2026-05-29

Smart Images

  • Figure CN117292233B_ABST
    Figure CN117292233B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a tongue feature recognition method and device, electronic equipment and a storage medium. The method comprises: acquiring a tongue image; performing multi-scale operation on the tongue image to determine at least two scale transformation images; inputting the at least two scale transformation images into a recognition model to output target feature information corresponding to the tongue image through the recognition model; and the recognition model is a cross-network model composed of a segmented convolutional neural network (Seg-CNN) structure and a point convolutional neural network (Point-CNN) structure. In this way, when analyzing the tongue image, the multi-scale operation is performed on the tongue image, and the network model composed of the Seg-CNN structure and the Point-CNN structure is used, so that the global feature information with points, lines and surfaces can be recognized, thereby the crack feature in the tongue image can be clearly recognized, and the recognition accuracy of the crack feature is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biometric identification technology, and in particular to a method, device, electronic device and storage medium for tongue image feature recognition. Background Technology

[0002] Tongue analysis in Traditional Chinese Medicine (TCM) is a simple and effective method to aid in diagnosis and differentiation of TCM syndromes by observing changes in the color and shape of the tongue. The tongue is considered the "sprout of the heart" and the external manifestation of the spleen; its coating originates from the stomach qi. The internal organs are connected to the tongue through meridians, making it an important diagnostic tool in TCM. The Hand Shaoyin meridian connects to the root of the tongue, the Foot Shaoyin meridian runs along the root of the tongue, the Foot Jueyin meridian connects to the root of the tongue, and the Foot Taiyin meridian connects to the root of the tongue and disperses under the tongue. Therefore, diseases of the internal organs can be reflected in the tongue's texture and coating. Tongue analysis primarily examines the shape, color, moisture, and dryness of the tongue's texture and coating to determine the nature and severity of the disease, the state of qi and blood, the abundance or deficiency of body fluids, and the state of the internal organs.

[0003] In health service practice, if a compliant tongue photograph can be collected, it can be analyzed using corresponding artificial intelligence recognition algorithms or by professional TCM analysts to obtain real-time analysis conclusions on the user's health status.

[0004] However, in related technologies, most of the acquired tongue images are often blurry and unclear, especially regarding minute features such as cracks. It is also difficult to extract these minute features in subsequent image processing. As a result, when artificial intelligence recognition algorithms or professional TCM analysts analyze tongue images, they cannot comprehensively analyze the tongue image, leading to inaccurate tongue image analysis results. Summary of the Invention

[0005] This application provides a tongue image feature recognition method, device, electronic device, and storage medium, which can clearly identify crack features in the tongue image and improve the recognition accuracy of crack features.

[0006] The technical solution of this application is implemented as follows:

[0007] In a first aspect, embodiments of this application provide a tongue image feature recognition method, the method comprising:

[0008] Obtain tongue image;

[0009] Perform multi-scale operations on the tongue image to determine at least two scale-transformed images;

[0010] The at least two scale-transformed images are input into the recognition model, and the recognition model outputs the target feature information corresponding to the tongue image; wherein, the recognition model is a cross-network model composed of a segmented convolutional neural network (Seg-CNN) structure and a point convolutional neural network (Point-CNN) structure.

[0011] Secondly, embodiments of this application provide a tongue image feature recognition device, including an acquisition unit, a scaling operation unit, and a recognition unit; wherein,

[0012] The acquisition unit is configured to acquire a tongue image;

[0013] The scale operation unit is configured to perform multi-scale operations on the tongue image to determine at least two scale-transformed images.

[0014] The recognition unit is configured to input the at least two scale-transformed images into the recognition model and output the target feature information corresponding to the tongue image through the recognition model; wherein, the recognition model is a cross-network model composed of Seg-CNN structure and Point-CNN structure.

[0015] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein,

[0016] The memory is used to store computer programs that can run on the processor;

[0017] The processor is configured to execute the method as described in the first aspect when running the computer program.

[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by at least one processor, implements the method described in the first aspect.

[0019] This application provides a tongue image feature recognition method, apparatus, electronic device, and storage medium. The method involves acquiring a tongue image; performing multi-scale operations on the tongue image to determine at least two scale-transformed images; inputting the at least two scale-transformed images into a recognition model; and outputting target feature information corresponding to the tongue image through the recognition model. The recognition model is a cross-network model composed of Seg-CNN and Point-CNN structures. Thus, when analyzing the tongue image, by performing multi-scale operations on the tongue image and utilizing a network model composed of Seg-CNN and Point-CNN structures, global feature information with points, lines, and surfaces can be identified, thereby clearly identifying crack features in the tongue image and improving the accuracy of crack feature recognition. Attached Figure Description

[0020] Figure 1 A flowchart illustrating a tongue image feature recognition method provided in an embodiment of this application;

[0021] Figure 2A flowchart illustrating another tongue image feature recognition method provided in this application embodiment;

[0022] Figure 3 A schematic diagram of the network composition of a Seg-CNN structure provided in an embodiment of this application;

[0023] Figure 4 This application provides a schematic diagram of the network composition of a Point-CNN structure.

[0024] Figure 5 This is a schematic diagram of the composition structure of an identification model provided in an embodiment of this application;

[0025] Figure 6 This is a schematic diagram of the composition structure of a tongue image feature recognition device provided in an embodiment of this application;

[0026] Figure 7 A schematic diagram of the composition structure of an electronic device provided in an embodiment of this application;

[0027] Figure 8 This is a schematic diagram of the composition structure of another electronic device provided in an embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the relevant application and not for limiting the application. Furthermore, it should be noted that, for ease of description, only the parts related to the relevant application are shown in the accompanying drawings.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0030] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0031] It should be noted that the terms "first, second, and third" used in the embodiments of this application are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0032] In recent years, with social development, the public has increasingly valued health, and the trend of preventive health management has become evident. At the same time, the state attaches great importance to traditional Chinese medicine (TCM), requiring the leveraging of TCM's original advantages and promoting its inheritance, innovation, and development. The "Internet + TCM Health Services" initiative is being implemented, encouraging the development of internet-based TCM health services based on medical institutions and TCM information technology, and promoting integrated online and offline services as well as remote health services.

[0033] Tongue diagnosis is one of the most direct and fundamental diagnostic methods in Traditional Chinese Medicine (TCM) clinical practice. It has been highly regarded by numerous physicians since ancient times and widely applied in clinical practice. The tongue's appearance contains rich physiological and pathological information. By observing the tongue's coating and related attributes, including color and shape, the location of diseases can be determined, leading to syndrome differentiation and treatment. This has significant reference value for TCM medication and disease diagnosis. However, for a long time, tongue diagnosis results have relied entirely on the doctor's subjective judgment. The accuracy of diagnostic information is influenced by the doctor's experience and environmental factors, resulting in a lack of objective diagnostic methods and standards. Furthermore, most tongue diagnosis experience is difficult to pass on and preserve, hindering the development of tongue diagnosis to some extent. Therefore, based on TCM theory, combining TCM diagnosis and treatment with image analysis technology to quantitatively analyze tongue appearance and achieve objectivity, standardization, and quantification of tongue diagnosis has become an essential path for the development of TCM tongue diagnosis.

[0034] As people's daily habits and physical conditions change, various shapes, depths, and numbers of cracks and grooves may appear on the tongue. Cracks or grooves without a coating are often pathological; those with a coating are more commonly congenital. In Traditional Chinese Medicine (TCM), a fissured tongue is a manifestation of systemic malnutrition, caused by deficiency of essence and blood, malnourishment of the tongue, atrophy of the tongue papillae, or tissue rupture. Therefore, tongue analysis primarily examines the shape, color, and moisture of the tongue and its coating to determine the nature and severity of the disease, the state of qi and blood, the abundance or deficiency of body fluids, and the condition of the internal organs.

[0035] In health service practice, if a compliant tongue photograph can be collected, it can be analyzed using corresponding artificial intelligence recognition algorithms or by professional TCM analysts to obtain real-time analysis conclusions on the user's health status.

[0036] However, in related technologies, most of the acquired tongue images are often blurry and unclear, especially regarding minute features such as cracks. It is also difficult to extract these minute features in subsequent image processing. As a result, when artificial intelligence recognition algorithms or professional TCM analysts analyze tongue images, they cannot comprehensively analyze the tongue image, leading to inaccurate tongue image analysis results.

[0037] Based on this, the embodiments of this application aim to provide a tongue image feature recognition method, which involves acquiring a tongue image; performing multi-scale operations on the tongue image to determine at least two scale-transformed images; inputting the at least two scale-transformed images into a recognition model; and outputting target feature information corresponding to the tongue image through the recognition model. The recognition model is a cross-network model composed of Seg-CNN and Point-CNN structures. Thus, when analyzing the tongue image, by performing multi-scale operations on the tongue image and using a network model composed of Seg-CNN and Point-CNN structures to extract intersection and line segment features, global feature information with points, lines, and surfaces can be identified, thereby clearly identifying crack features in the tongue image and improving the accuracy of crack feature recognition.

[0038] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0039] In one embodiment of this application, see [link to embodiment]. Figure 1 The diagram illustrates a flowchart of a tongue image feature recognition method provided in an embodiment of this application. Figure 1 As shown, the method may include:

[0040] S101. Obtain the tongue image.

[0041] It should be noted that this application provides a tongue image feature recognition method, which can be applied to a tongue image feature recognition device or an electronic device integrated with such a device. Here, the electronic device can be such as a computer, smartphone, tablet computer, laptop computer, PDA, navigation device, wearable device, server, etc., and this application does not specifically limit it in this regard.

[0042] It should also be noted that in this step, the user can use some conventional image acquisition devices (such as cameras, camcorders, or tongue imagers) to acquire images of the tongue. Here, the image acquisition device can be the electronic device used to perform the tongue feature recognition method as described in the embodiments of this application, or it can be a device only used for image acquisition, and then, after acquiring the tongue image, send the tongue image to the electronic device used to perform the tongue feature recognition method.

[0043] Additionally, it's important to note that the tongue image obtained here is a standardized photograph of the tongue. Specifically, the image must fully represent the complete features of the tongue and not interfere with the analysis by artificial intelligence algorithms or traditional Chinese medicine analysts.

[0044] S102. Perform multi-scale operations on the tongue image to determine at least two scale-transformed images.

[0045] It should be noted that, in the embodiments of this application, different scale transformation operations are performed on the tongue image to obtain at least two scale-transformed images. Specifically, this refers to performing scale enlargement / reduction transformation operations on the acquired tongue image to obtain an image with a different size from the tongue image.

[0046] In some embodiments, performing multi-scale operations on a tongue image to determine at least two scale-transformed images may include:

[0047] Different scaling operations are performed on the tongue image to obtain at least two scale-transformed images; wherein, the scaling operation may include: scaling up operation and / or scaling down operation.

[0048] It should be noted that, in the embodiments of this application, the scaling operation can be to enlarge the size of the original tongue image, such as an upsampling operation; the scaling reduction operation can be to reduce the size of the original tongue image, such as a downsampling operation.

[0049] In one specific embodiment, taking scale magnification as an example, different scale transformation operations are performed on the tongue image to obtain at least two scale-transformed images, which may include:

[0050] By using bilinear interpolation to perform different scale magnification operations on the tongue image, at least two scale-transformed images are obtained.

[0051] In this embodiment of the application, the scaling operation can be implemented using bilinear interpolation. Bilinear interpolation, also known as bilinear interpolation, refers to the linear interpolation extension of an interpolation function with two variables. Its core idea is to perform linear interpolation in two directions (first using linear interpolation in one direction, and then using linear interpolation in the other direction to perform bilinear interpolation).

[0052] In another specific embodiment, taking scale reduction as an example, different scale transformation operations are performed on the tongue image to obtain at least two scale-transformed images, which may include:

[0053] By using downsampling to perform different scale reduction operations on the tongue image, at least two scale-transformed images are obtained.

[0054] In this embodiment, the scaling operation can be implemented using downsampling. Downsampling refers to downsampling an image of size M×N by a factor of s, resulting in an image of size (M / s)×(N / s), where s is a common divisor of M and N.

[0055] It should also be noted that, for scaling up operations, in addition to using bilinear interpolation, other methods can be used for scaling up operations, and the embodiments of this application do not specifically limit the methods; for scaling down operations, in addition to using downsampling, other methods can be used for scaling down operations, and the embodiments of this application do not specifically limit the methods.

[0056] Additionally, it should be noted that after performing multi-scale operations on the tongue image, the determined at least two scale-transformed images may or may not include the original tongue image. This is determined based on practical application requirements, and the embodiments in this application do not impose specific limitations.

[0057] S103. Input at least two scale-transformed images into the recognition model, and output the target feature information corresponding to the tongue image through the recognition model; wherein, the recognition model is a cross-network model (or, may also be called a cross extractor) composed of a segmented convolutional neural network (Seg-CNN) structure and a point convolutional neural network (Point-CNN) structure.

[0058] It should be noted that, in the embodiments of this application, the recognition model is mainly used to perform cross-processing on at least two scale-transformed images in order to extract the target feature information corresponding to the tongue image. Here, the recognition model can be composed of a segment convolutional neural network (Seg-CNN) structure and a point convolutional neural network (Point-CNN) structure.

[0059] It should also be noted that, in the embodiments of this application, the at least two scale-transformed images input into the recognition model may include the original tongue image. Thus, for multi-scale tongue images, the Seg-CNN structure uses cross-connection encoders and decoders and feature set information, ensuring that line segments in the tongue cracks exist in only one feature map as much as possible. Furthermore, the multi-scale approach allows for classification using other feature maps even when some areas are occluded, increasing the network's generalization ability. Additionally, the Point-CNN structure extracts horizontal and vertical feature maps and performs mathematical operations on these maps to obtain intersection features (or "cross-point features") in the tongue cracks. Then, the line segment features and intersection features extracted from the Seg-CNN and Point-CNN structures are cross-inputted, ultimately enabling the identification of crack features (i.e., target feature information) in the tongue image.

[0060] This embodiment provides a tongue image feature recognition method. The method involves acquiring a tongue image; performing multi-scale operations on the tongue image to determine at least two scale-transformed images; inputting these at least two scale-transformed images into a recognition model; and outputting the target feature information corresponding to the tongue image through the recognition model. The recognition model is a cross-network model composed of a segmented convolutional neural network (Seg-CNN) structure and a point convolutional neural network (Point-CNN) structure. Thus, when analyzing the tongue image, by performing multi-scale operations on the tongue image and using the network model composed of Seg-CNN and Point-CNN structures to extract intersection and line segment features, global feature information with points, lines, and surfaces can be identified. This allows for clear identification of crack features in the tongue image, improving the accuracy of crack feature recognition.

[0061] In another embodiment of this application, see Figure 2 This illustrates a flowchart of another tongue feature recognition method provided in an embodiment of this application. Figure 2 As shown, the method may include:

[0062] S201. Obtain a tongue image.

[0063] S202. Perform multi-scale operations on the tongue image to determine at least two scale transformation images.

[0064] S203. Use Seg-CNN and Point-CNN structures to perform feature cross-extraction on at least two scale-transformed images to determine the tongue image feature information corresponding to at least two scale-transformed images.

[0065] S204. Perform weighted calculation on the tongue feature information corresponding to at least two scale-transformed images to obtain the target feature information.

[0066] It should be noted that, in this embodiment of the application, a method for identifying minute features (such as crack features) in an image is proposed. Specifically, this method extracts multi-scale, intersection point, and line segment features from a tongue image, enabling the identification of crack length and thus clearly recognizing crack features in the tongue image.

[0067] It should also be noted that, in the embodiments of this application, in the recognition model, the Seg-CNN structure extracts features from at least two scale-transformed images to obtain at least one set of line segment feature information; the Point-CNN structure extracts features from at least two scale-transformed images to obtain at least one set of intersection feature information. Therefore, in some embodiments, using the Seg-CNN structure and the Point-CNN structure to perform feature cross-extraction processing on at least two scale-transformed images to determine the tongue image feature information corresponding to at least two scale-transformed images may include: using the Seg-CNN structure and the Point-CNN structure to extract features from at least two scale-transformed images to obtain at least one set of line segment feature information and at least one set of intersection feature information; and performing cross-processing on the at least one set of line segment feature information and at least one set of intersection feature information to obtain the tongue image feature information corresponding to at least two scale-transformed images.

[0068] In one specific embodiment, feature extraction is performed on at least two scale-transformed images using Seg-CNN and Point-CNN structures to obtain at least one set of line segment feature information and at least one set of intersection point feature information, which may include:

[0069] The Seg-CNN structure is used to extract features from the image at the first scale transformation, resulting in a set of line segment feature information corresponding to the image at the first scale transformation; and

[0070] The Point-CNN structure is used to extract features from the second-scale transformed image to obtain a set of intersection feature information corresponding to the second-scale transformed image.

[0071] It should be noted that, in the embodiments of this application, the first scale transformation image is different from the second scale transformation image, and the first scale transformation image and the second scale transformation image are any two of at least two scale transformation images.

[0072] It should also be noted that the Seg-CNN structure can be used to extract features from one of at least two scale-transformed images to obtain a set of line segment feature information corresponding to this scale-transformed image; the Point-CNN structure can be used to extract features from one of at least two scale-transformed images to obtain a set of intersection feature information corresponding to this scale-transformed image.

[0073] In a more specific embodiment, feature extraction of the first-scale transformed image is performed using a Seg-CNN structure to obtain a set of line segment feature information corresponding to the first-scale transformed image, which may include:

[0074] The first scale transformation image is processed by an encoder and decoder to extract features and obtain a feature set; the feature set includes multiple sets of line segment features, which are line segment features of the tongue image features extracted from the first scale transformation image at different positions.

[0075] Perform a dot product operation on the feature set to obtain grouping information of multiple line segment features;

[0076] By performing a union operation on the grouped information of multiple sets of line segment features, a set of line segment feature information corresponding to the first scale transformed image is obtained.

[0077] It should be noted that, see Figure 3 This illustrates a schematic diagram of the network composition of a Seg-CNN structure provided in an embodiment of this application. Figure 3 As shown, the Seg-CNN structure may include an encoder-decoder 301. Using the first-scale transformed image as input to the Seg-CNN structure, the encoder-decoder 301 extracts features from it to obtain a set of feature groups. Then, through dot product and union operations, the output of the Seg-CNN structure is obtained, which is a set of line segment feature information corresponding to the first-scale transformed image.

[0078] In some embodiments, the codec may include an even number of convolutional layers; wherein the even number of convolutional layers have a symmetrical structure and the input connections between the even number of convolutional layers are made in a cross-connection manner.

[0079] It should be noted that the minimum number of convolutional layers is 4, but the maximum number of layers is not limited; for example, the number of convolutional layers can be 4, 6, 8, 12, etc. In this embodiment, a 6-layer convolutional layer is used as an example, wherein the number of channels of the convolutional kernels of the 6 convolutional layers are 64, 32, 8, 8, 32, and 64, respectively; and these 6 convolutional layers have a symmetrical structure.

[0080] It should also be noted that, such as Figure 3 As shown, the codec 301 may include six convolutional layers, where the first three convolutional layers are encoders and the last three convolutional layers are decoders; the six convolutional layers of the codec are connected to each other via cross-connection. It should be noted that, in addition to using... Figure 3 In addition to the cross-connection method shown, other cross-connection methods can also be used for input connection, and this application embodiment does not specifically limit this.

[0081] Furthermore, a feature set is obtained through an encoder-decoder. Based on this feature set, line segment features of the cracks at different locations in the six tongue images can be extracted. Specifically, in some embodiments, a dot product operation is performed on the feature set to obtain grouping information of multiple sets of line segment features, which may include:

[0082] The sliding window method is used to find local peaks in the line segment features in the feature set and to determine the local peak region corresponding to each line segment feature in the feature set.

[0083] By splicing together the features of adjacent line segments in the local peak region, grouping information of multiple sets of line segment features is obtained.

[0084] It should be noted that, taking the first scale transformed image as the input of the Seg-CNN structure as an example, each line segment feature in the feature set represents a certain part of the first scale transformed image, that is, it will produce a peak response to a certain part of the first scale transformed image. The function of the feature set is to stitch together the adjacent features of these local peak regions.

[0085] It should also be noted that in this embodiment, the sliding window method used here has a window size of 1 / 32. The pixel values ​​within the window are then accumulated and compared with a threshold. The initial threshold value is set to 0.2, and the rate of change is a cosine function. During iterative training, the threshold varies between 0.2 and 1.0. Furthermore, the initial threshold value is typically set to 0.2 based on the Pareto principle commonly used in data analysis; the range of 0.2 to 1.0 helps to retain peak information in the line segment features of the feature set as much as possible while eliminating useless information.

[0086] Furthermore, such as Figure 3 As shown, three distinct regions can be obtained, and each region is then classified to determine the line segment feature information. The loss function of the Seg-CNN structure is as follows:

[0087]

[0088] k(i)=∑[||xu x || 2 -||yu y || 2 ] / [||xu x || 2 +||yu y || 2 (2)

[0089]

[0090] Where x and y represent the pixel coordinates in the first-scale transformed image corresponding to the line segment feature map obtained through the Seg-CNN structure; u x and u y i represents the average of x and y; k (x,y) represents the response value of the final line segment feature map; λ represents the penalty factor.

[0091] In this way, by using a loss function k(i) when performing a union operation on the grouped information of multiple line segment features, the feature map combination can become more clustered within the feature map combination, and the feature map combination can become more dispersed between the feature map combinations. This allows the line segment features of the tongue crack to exist in only one feature map as much as possible. In addition, by having multiple feature maps of different scales, it is possible to classify the network by using other feature maps even when some areas are occluded, thus increasing the generalization ability of the network.

[0092] In other words, combining Figure 3 For the Seg-CNN structure, firstly, the first scale-transformed image is processed by a symmetric and cross-connected encoder-decoder 301 to extract features, resulting in a feature set. The feature set includes multiple sets of line segment features, which are line segment features of the tongue image extracted from the first scale-transformed image at different positions. Then, a dot product operation is performed on the feature set to obtain the grouping information of the multiple sets of line segment features. Finally, a union operation is performed on the grouping information of the multiple sets of line segment features to obtain a set of line segment feature information corresponding to the first scale-transformed image.

[0093] In another, more specific embodiment, feature extraction of the second-scale transformed image is performed using a Point-CNN structure to obtain a set of intersection feature information corresponding to the second-scale transformed image, which may include:

[0094] The initial feature matrix is ​​obtained by extracting features from the second-scale transformed image using a convolution module.

[0095] The initial feature matrix is ​​extracted horizontally using the horizontal extraction module to obtain the first feature matrix.

[0096] The initial feature matrix is ​​subjected to vertical feature extraction using the vertical extraction module to obtain the second feature matrix.

[0097] Matrix operations are performed on the first and second feature matrices to obtain a set of intersection feature information corresponding to the second scale transformed image.

[0098] It should be noted that, in the embodiments of this application, the convolution module may include at least three convolutional layers; the horizontal extraction module may include H×1×W / 2 extractors, and the vertical extraction module may include 1×W×H / 2 extractors; wherein, H represents the height of the second-scale transformed image, and W represents the width of the second-scale transformed image.

[0099] It should also be noted that, see Figure 4 This illustrates a schematic diagram of the network composition of a Point-CNN structure provided in an embodiment of this application. Figure 4As shown, the Point-CNN structure may include a convolutional module 401, a horizontal extraction module 402, and a vertical extraction module 403. Figure 4 In the process, the convolution module 401 includes three convolutional layers, and the three convolutional layers have an asymmetric structure. After the initial feature matrix is ​​extracted by the convolution module 401, on the one hand, the horizontal feature extraction module 402 can be used to extract features in the horizontal direction to obtain the first feature matrix; on the other hand, the vertical feature extraction module 403 can be used to extract features in the vertical direction to obtain the second feature matrix. Finally, by performing matrix operations on the first feature matrix and the second feature matrix, a set of intersection feature information corresponding to the second scale transformed image can be obtained.

[0100] Furthermore, in some embodiments, performing matrix operations on the first feature matrix and the second feature matrix to obtain a set of intersection point feature information corresponding to the second scale-transformed image may include:

[0101] Perform matrix addition on the first characteristic matrix and the second characteristic matrix to obtain the addition result;

[0102] Perform matrix multiplication on the first and second characteristic matrices to obtain the multiplication result;

[0103] Perform matrix division on the results of addition and multiplication, and use the resulting division as a set of intersection feature information corresponding to the second-scale transformed image.

[0104] It should be noted that, in the embodiments of this application, matrix multiplication here can refer to matrix dot product. First, matrix addition and matrix dot product operations are performed on the first feature matrix and the second feature matrix. The matrix addition and matrix dot product operations can be performed simultaneously or in steps, without specific limitations. Then, matrix division is performed on the addition result obtained from the matrix addition operation and the dot product result obtained from the matrix dot product operation to obtain a set of intersection feature information corresponding to the second scale transformed image.

[0105] In other words, combining Figure 4For the Point-CNN structure, firstly, feature extraction is performed on the second-scale transformed image using convolutional module 401 to obtain an initial feature matrix; then, horizontal extraction module 402 extracts features from the initial feature matrix in the horizontal direction to obtain a first feature matrix; and vertical extraction module 403 extracts features from the initial feature matrix in the vertical direction to obtain a second feature matrix. Finally, matrix operations are performed on the first and second feature matrices to obtain a set of intersection feature information corresponding to the second-scale transformed image. In this way, by using the Point-CNN network to extract features from both the horizontal and vertical directions and performing mathematical operations on the extracted feature matrices, the intersection features in the tongue image can be obtained.

[0106] Thus, by cross-processing at least one set of line segment feature information obtained by feature extraction from at least two scale-transformed images using the Seg-CNN structure and at least one set of intersection point feature information obtained by feature extraction from at least two scale-transformed images using the Point-CNN structure, tongue image feature information corresponding to at least two scale-transformed images can be obtained. By weighting the tongue image feature information corresponding to at least two scale-transformed images, the target feature information, namely the micro-features (such as crack feature information) in the tongue image, can be obtained.

[0107] Understandably, in the embodiments of this application, the number of Seg-CNN structures and Point-CNN structures for the recognition model can be multiple. For example, the number of Seg-CNN structures is six, and the number of Point-CNN structures is three.

[0108] In some embodiments, six Seg-CNN structures and three Point-CNN structures form a 3×3 matrix network, and at least two scale-transformed images include a first scale-transformed image, a second scale-transformed image, and a third scale-transformed image. Accordingly, the at least two scale-transformed images are input into the recognition model, and the recognition model outputs target feature information corresponding to the tongue image, which may include:

[0109] A matrix network is used to extract and cross-process features from the first scale transformed image, the second scale transformed image, and the third scale transformed image to obtain the first tongue image feature information, the second tongue image feature information, and the third tongue image feature information.

[0110] The target feature information is obtained by weighting the first, second, and third tongue image feature information.

[0111] For example, see Figure 5 This illustrates a schematic diagram of the composition structure of an identification model provided in an embodiment of this application. For example... Figure 5As shown, this recognition model uses six Seg-CNN structures and three Point-CNN structures. The model is divided into three scales. Bilinear interpolation is used to obtain three scale-transformed images of the original tongue image. These three scale-transformed images are then used as input and processed by a cross-extractor composed of Seg-CNN and Point-CNN structures. Through the combined network structures, global feature information with points, lines, and surfaces is constructed, enabling clear identification of cracks in the tongue image.

[0112] It should also be noted that, for Figure 5 The recognition model shown can have other numbers of Seg-CNN and Point-CNN structures, but it must include at least one Seg-CNN or Point-CNN structure. Furthermore, the Seg-CNN and Point-CNN structures can be placed in any other location; this is not specifically limited here. For example, for... Figure 5 Regarding the recognition model shown, the Seg-CNN structure and the Point-CNN structure can also form a 2×2 matrix network or a 4×4 matrix network, corresponding to two scale-transformed images or four scale-transformed images input to the recognition model. This application embodiment does not make specific limitations. The number of Seg-CNN structures and the number of Point-CNN structures can be arbitrary, but there must be at least one Seg-CNN structure or Point-CNN structure. This application embodiment does not make specific limitations. The Seg-CNN structure and the Point-CNN structure can also be in any other position. This application embodiment does not make specific limitations.

[0113] In summary, the embodiments of this application provide a tongue image feature recognition method, specifically a method for recognizing minute features in an image, which may include the following steps:

[0114] Step 1: Perform multi-scale operations on the acquired tongue images to extract crack information from the tongue images through the intersection of Seg-CNN and Point-CNN structures.

[0115] Step 2: The Seg-CNN structure first obtains a set of feature groups through a symmetric encoder and decoder, then obtains grouped information of multiple line segment features through dot multiplication of the feature group set, and then obtains the line segment features in the tongue image through union operation.

[0116] Step 3: The Point-CNN structure mainly extracts features from the horizontal and vertical directions. The feature matrices extracted from the horizontal and vertical directions are then used to obtain the intersection features in the tongue image.

[0117] Step 4: Extract line segment features and intersection features from the tongue image using the Seg-CNN and Point-CNN structures, and then input these features as cross-features to obtain the crack feature information in the tongue image.

[0118] In one specific implementation, the method for identifying minute features in an image may include the following detailed steps:

[0119] (1) Perform multi-scale operations on the acquired tongue image. Here, it is divided into three scales. The three scale transformation images of the acquired tongue image can be obtained by bilinear interpolation. These three scale transformation images are used as the three inputs of the model and then passed through the cross-extractor composed of Seg-CNN structure and Point-CNN structure.

[0120] (2) The Seg-CNN structure first uses an encoder-decoder, which consists of six symmetrical convolutional layers with channels of 64, 32, 8, 8, 32, and 64 respectively. Furthermore, the convolutional layers between the encoder and decoder are connected to the input via cross-connection, such as... Figure 3 As shown. The encoder and decoder can obtain the feature set of the tongue image, that is, extract the line segment features of the crack at different positions in the six tongue image feature maps respectively; then, after the dot product operation, the adjacent line segment features are combined together to obtain three feature maps. Finally, the three feature maps are combined through the union operation to obtain the pure crack features.

[0121] (3) The Point-CNN structure first passes through a 3-layer convolutional module, and then performs feature extraction in the horizontal and vertical directions respectively. There are 1×W×H / 2 extractors in the vertical direction and H×1×W / 2 extractors in the horizontal direction. After feature extraction in the horizontal and vertical directions respectively, feature matrices A and B in two directions can be obtained. The intersection features of cracks in the tongue image can be obtained by (A+B) / (A·B) operation.

[0122] (4) The cross-extractor consists of 6 Seg-CNN structures and 3 Point-CNN structures, wherein the connection between these structures is as follows: Figure 5 As shown, the line segment features extracted from the tongue image are obtained by using the Seg-CNN structure and the intersection point features extracted by the Point-CNN structure, and then cross-connected to obtain accurate crack feature information in the tongue image.

[0123] (5) The Seg-CNN structure first takes multi-scale tongue images as input. Taking one input tongue image as an example, after passing through the convolutional layer, a series of feature maps are generated. The pixels of each feature map represent a certain part of the tongue image, that is, it will produce a peak response to a certain part of the tongue image. Figure 3There are six feature maps, each containing line segment (crack) regions. The feature set is used to concatenate adjacent features from these peak response regions. A sliding window method is used to find local peaks on the feature maps within the feature set, with a window size of 1 / 32. The pixel values ​​within the window are then accumulated and compared with a threshold. The initial threshold is set to 0.2, with a rate of change of a cosine function, varying between 0.2 and 1.0 during iterative training. Furthermore, the initial threshold is typically set to 0.2 based on the Pareto principle commonly used in data analysis. Varying the threshold range from 0.2 to 1.0 helps to retain peak information in the feature maps while eliminating useless information.

[0124] (6) Figure 3 As shown, three different regions can be obtained, and each region is then classified. The loss function of the Seg-CNN structure is as described in equations (1) to (3) above. Among them, k(i) can make the feature map combination more clustered, and m(i) can make the feature map combination more dispersed. This allows the line segment features of the tongue crack to exist in only one feature map as much as possible. In addition, by having multiple feature maps of different scales, it is possible to classify the network by other feature maps even when some regions are occluded, thus increasing the generalization ability of the network.

[0125] (7) Figure 5 The cross-extraction generator, composed of Seg-CNN and Point-CNN structures, has a feature coefficient before the input. Figure 5 For example, there are 9 networks forming the feature coefficient matrix M, as shown below:

[0126]

[0127] Here, matrix M is a positive semi-definite matrix, which is more conducive to the convergence of the recognition model. Assuming the cross-extractors composed of the Seg-CNN and Point-CNN structures are numbered n11, n12, n13, n21, n22, n23, n31, n32, n33, then they can form matrix N as shown below:

[0128]

[0129] In this way, when training the cross-extractor, the feature coefficients in the matrix M at the corresponding positions can be added, and the network parameters of the cross-extractor can be gradually converged during the training process to determine the final cross-extractor (i.e., the recognition model described in the aforementioned embodiment).

[0130] It should be noted that, taking the input of three scale-transformed images into the recognition model as an example, the first tongue image feature information, the second tongue image feature information, and the third tongue image feature information can be obtained. Thus, the determination of the target feature information (i.e., crack features) can specifically include: when the recognition model is performing calculations, matrix N is added to the coefficients in matrix M at the corresponding positions, that is, when each network in the recognition model is performing calculations, the corresponding feature coefficients are added, and then the calculated first tongue image feature information, second tongue image feature information, and third tongue image feature information are added to obtain the target feature information.

[0131] Understandably, in the embodiments of this application, on the one hand, this technical solution proposes a cross-extractor composed of multi-scale operations, Seg-CNN structure, and Point-CNN structure. In input tongue images of different sizes, a matrix network composed of Seg-CNN and Point-CNN structures is used; this matrix is ​​a positive semi-definite matrix. The matrix network is superimposed to form global feature information with points, lines, and surfaces, thus clearly identifying cracks in the tongue image. On the other hand, the Seg-CNN structure uses cross-connection encoders and decoders and feature set information, which ensures that crack segments in the tongue image exist in only one feature map as much as possible. Furthermore, the multi-scale method allows for classification using other feature maps even when some areas are occluded, increasing the network's generalization ability. Moreover, this technical solution also proposes a Point-CNN structure, which extracts horizontal and vertical feature maps and performs mathematical operations on these feature maps to obtain the intersection information of cracks in the tongue image.

[0132] This embodiment provides a tongue image feature recognition method. The specific implementation of the aforementioned embodiments is described in detail below. It can be seen that, according to the technical solution of the aforementioned embodiments, the matrix network composed of Seg-CNN and Point-CNN structures constitutes global feature information with points, lines, and surfaces. This allows for clear identification of cracks in the tongue image. The cross-connection encoder / decoder and feature group set information method can adjust the line segment feature structure within and between feature map combinations, and provide network generalization ability through multi-scale methods. The Point-CNN structure extracts horizontal and vertical feature maps from the tongue image, and mathematical operations are performed on the feature maps to obtain the intersection information of cracks in the tongue image. Therefore, it can clearly identify crack features in the tongue image, improving the accuracy of crack feature recognition.

[0133] In another embodiment of this application, see [link to application]. Figure 6 This illustration shows a schematic diagram of the composition of a tongue feature recognition device 60 provided in an embodiment of this application. Figure 6As shown, the tongue image feature recognition device 60 may include an acquisition unit 601, a scaling operation unit 602, and a recognition unit 603, wherein,

[0134] Acquisition unit 601 is configured to acquire tongue image;

[0135] The scaling operation unit 602 is configured to perform multi-scale operations on the tongue image to determine at least two scale-transformed images.

[0136] The recognition unit 603 is configured to input at least two scale-transformed images into the recognition model and output the target feature information corresponding to the tongue image through the recognition model; wherein, the recognition model is a cross-network model composed of Seg-CNN structure and Point-CNN structure.

[0137] In some embodiments, the scaling operation unit 602 is configured to perform different scaling operations on the tongue image to obtain at least two scale-transformed images; wherein the scaling operations include scaling up and / or scaling down operations.

[0138] In some embodiments, the scaling unit 602 is further configured to perform different scaling operations on the tongue image using bilinear interpolation to obtain at least two scale-transformed images.

[0139] In some embodiments, the recognition unit 603 is configured to extract features from at least two scale-transformed images using Seg-CNN and Point-CNN structures to obtain at least one set of line segment feature information and at least one set of intersection feature information; and to perform cross-processing on the at least one set of line segment feature information and at least one set of intersection feature information to obtain tongue image feature information corresponding to at least two scale-transformed images; and to perform weighted calculation on the tongue image feature information corresponding to at least two scale-transformed images to obtain target feature information.

[0140] In some embodiments, such as Figure 6 As shown, the tongue image feature recognition device 60 may include a feature extraction unit 604, configured to extract features from a first scale transformation image using a Seg-CNN structure to obtain a set of line segment feature information corresponding to the first scale transformation image; and to extract features from a second scale transformation image using a Point-CNN structure to obtain a set of intersection feature information corresponding to the second scale transformation image; wherein the first scale transformation image and the second scale transformation image are different, and the first scale transformation image and the second scale transformation image are any two of at least two scale transformation images.

[0141] In some embodiments, the Seg-CNN structure includes an encoder and decoder; correspondingly, the feature extraction unit 604 is further configured to extract features from the first scale-transformed image through the encoder and decoder to obtain a feature set; wherein the feature set includes multiple sets of line segment features, which are line segment features of the tongue image features extracted from the first scale-transformed image at different positions; and to perform a dot product operation on the feature set to obtain grouping information of the multiple sets of line segment features; and to perform a union operation on the grouping information of the multiple sets of line segment features to obtain a set of line segment feature information corresponding to the first scale-transformed image.

[0142] In some embodiments, such as Figure 6 As shown, the tongue image feature recognition device 60 may include a computing unit 605, configured to use a sliding window method to search for local peaks in the line segment features in the feature set, determine the local peak region corresponding to each line segment feature in the feature set; and splice the line segment features adjacent to the local peak region to obtain grouping information of multiple sets of line segment features.

[0143] In some embodiments, the codec includes an even number of convolutional layers; wherein the even number of convolutional layers have a symmetrical structure and the input connections between the even number of convolutional layers are made in a cross-connection manner.

[0144] In some embodiments, the Point-CNN structure includes a convolution module, a horizontal extraction module, and a vertical extraction module; correspondingly, the feature extraction unit 604 is further configured to perform feature extraction on the second-scale transformed image through the convolution module to obtain an initial feature matrix; and to perform horizontal feature extraction on the initial feature matrix through the horizontal extraction module to obtain a first feature matrix; and to perform vertical feature extraction on the initial feature matrix through the vertical extraction module to obtain a second feature matrix; and to perform matrix operations on the first feature matrix and the second feature matrix to obtain a set of intersection feature information corresponding to the second-scale transformed image.

[0145] In some embodiments, the operation unit 605 is further configured to perform matrix addition on the first feature matrix and the second feature matrix to obtain the addition result; perform matrix multiplication on the first feature matrix and the second feature matrix to obtain the multiplication result; perform matrix division on the addition result and the multiplication result, and use the obtained division result as a set of intersection feature information corresponding to the second scale transformation image.

[0146] In some embodiments, the convolution module includes at least three convolutional layers; the horizontal extraction module includes H×1×W / 2 extractors, and the vertical extraction module includes 1×W×H / 2 extractors; wherein, H represents the height of the second-scale transformed image, and W represents the width of the second-scale transformed image.

[0147] In some embodiments, the recognition model has six Seg-CNN structures and three Point-CNN structures.

[0148] In some embodiments, six Seg-CNN structures and three Point-CNN structures form a 3×3 matrix network; at least two scale-transformed images include a first scale-transformed image, a second scale-transformed image, and a third scale-transformed image; correspondingly, the recognition unit 603 is further configured to use the matrix network to perform feature extraction and cross-processing on the first scale-transformed image, the second scale-transformed image, and the third scale-transformed image to obtain first tongue image feature information, second tongue image feature information, and third tongue image feature information; and to perform weighted calculation on the first tongue image feature information, the second tongue image feature information, and the third tongue image feature information to obtain target feature information.

[0149] Understandably, in this embodiment, a "unit" can be a portion of a circuit, a portion of a processor, a portion of a program or software, etc., and can also be a module or a non-modular component. Furthermore, the components in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.

[0150] If the integrated unit is implemented as a software functional module and not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method described in this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0151] Therefore, this embodiment provides a storage medium storing a computer program that, when executed by at least one processor, implements the steps of any of the tongue feature recognition methods described in the foregoing embodiments.

[0152] Based on the above-described composition and storage medium of a tongue image feature recognition device 60, see [link to relevant documentation]. Figure 7This illustrates a schematic diagram of the structural composition of an electronic device 70 provided in an embodiment of this application. For example... Figure 7 As shown, electronic device 70 may include: a communication interface 701, a memory 702, and a processor 703; the various components are coupled together via a bus system 704. It is understood that the bus system 704 is used to implement communication between these components. In addition to a data bus, the bus system 704 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 7 The various buses are all labeled as bus system 704. Among them, the communication interface 701 is used for receiving and sending signals during the process of sending and receiving information with other external network elements;

[0153] Memory 702 is used to store computer programs that can run on processor 703;

[0154] Processor 703, when running the computer program, performs the following:

[0155] Obtain tongue image;

[0156] Perform multi-scale operations on tongue images to determine at least two scale-transformed images;

[0157] At least two scale-transformed images are input into the recognition model, and the recognition model outputs the target feature information corresponding to the tongue image; wherein, the recognition model is a cross-network model composed of Seg-CNN structure and Point-CNN structure.

[0158] It is understood that the memory 702 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 702 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0159] The processor 703 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 703 or by software instructions. The processor 703 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 702, and the processor 703 reads the information in memory 702 and, in conjunction with its hardware, completes the steps of the above method.

[0160] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0161] For software implementation, the techniques described herein can be achieved through modules (e.g., procedures, functions, etc.) that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented within the processor or externally.

[0162] Alternatively, as another embodiment, the processor 703 is also configured to perform the method described in any of the foregoing embodiments when running the computer program.

[0163] In another embodiment of this application, see [link to application]. Figure 8 This illustrates a schematic diagram of the structural composition of another electronic device 70 provided in an embodiment of this application. For example... Figure 8 As shown, the electronic device 70 includes at least the tongue feature recognition device 60 described in any of the foregoing embodiments.

[0164] In this embodiment of the application, for the electronic device 70, when analyzing the tongue image, by performing multi-scale operations on the tongue image and using a network model composed of Seg-CNN and Point-CNN structures to extract intersection features and line segment features, it is possible to identify global feature information with points, lines and surfaces, thereby clearly identifying crack features in the tongue image and improving the accuracy of crack feature identification.

[0165] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.

[0166] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0167] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0168] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0169] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0170] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0171] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for tongue image feature recognition, characterized in that, The method includes: Obtain tongue image; Perform multi-scale operations on the tongue image to determine at least two scale-transformed images; The at least two scale-transformed images are input into a recognition model, which then outputs the target feature information corresponding to the tongue image; wherein, the recognition model is a segmented convolutional neural network Seg CNN structure and Point Convolutional Neural Network Cross-network models composed of CNN structures; The step of inputting the at least two scale-transformed images into the recognition model and outputting the target feature information corresponding to the tongue image through the recognition model includes: Using the Seg CNN structure and the Point The CNN structure performs feature extraction on the at least two scale-transformed images to obtain at least one set of line segment feature information and at least one set of intersection point feature information; Cross-processing is performed on the at least one set of line segment feature information and the at least one set of intersection point feature information to obtain tongue image feature information corresponding to the at least two scale-transformed images; The target feature information is obtained by weighting the tongue image feature information corresponding to the at least two scale-transformed images.

2. The method according to claim 1, characterized in that, The multi-scale operation on the tongue image to determine at least two scale-transformed images includes: Different scale transformation operations are performed on the tongue image to obtain at least two scale-transformed images; wherein the scale transformation includes: scale magnification operation and / or scale reduction operation.

3. The method according to claim 2, characterized in that, The process of performing different scale transformations on the tongue image to obtain the at least two scale-transformed images includes: The tongue image is magnified at different scales using bilinear interpolation to obtain at least two scale-transformed images.

4. The method according to claim 1, characterized in that, The use of the Seg CNN structure and the Point The CNN structure performs feature extraction on the at least two scale-transformed images to obtain at least one set of line segment feature information and at least one set of intersection point feature information, including: Using the Seg The CNN structure extracts features from the image at the first scale transformation, obtaining a set of line segment feature information corresponding to the image at the first scale transformation; and Using the Point The CNN structure extracts features from the second-scale transformed image to obtain a set of intersection feature information corresponding to the second-scale transformed image; Wherein, the first scale-transformed image is different from the second scale-transformed image, and the first scale-transformed image and the second scale-transformed image are any two of the at least two scale-transformed images.

5. The method according to claim 4, characterized in that, The Seg The CNN architecture includes an encoder and decoder; the Seg... The CNN structure extracts features from the first-scale transformed image to obtain a set of line segment feature information corresponding to the first-scale transformed image, including: The first scale-transformed image is subjected to feature extraction by the codec to obtain a feature set; wherein, the feature set includes multiple sets of line segment features, which are line segment features of the tongue image features extracted from the first scale-transformed image at different positions. Perform a dot product operation on the set of features to obtain the grouping information of the multiple sets of line segment features; Perform a union operation on the grouping information of the multiple sets of line segment features to obtain a set of line segment feature information corresponding to the first scale-transformed image.

6. The method according to claim 5, characterized in that, The step of performing a dot product operation on the feature set to obtain the grouping information of the multiple sets of line segment features includes: The local peak value of each line segment feature in the feature set is determined by using a sliding window method to find the local peak value region corresponding to each line segment feature in the feature set. The line segment features adjacent to the local peak region are spliced ​​together to obtain the grouping information of the multiple sets of line segment features.

7. The method according to claim 5, characterized in that, The codec includes an even number of convolutional layers; wherein the even number of convolutional layers have a symmetrical structure, and the input connections between the even number of convolutional layers are made using a cross-connection method.

8. The method according to claim 4, characterized in that, The Point The CNN structure includes a convolutional module, a horizontal extraction module, and a vertical extraction module; the point is used... The CNN architecture extracts features from images transformed at the second scale. A set of intersection feature information corresponding to the second scale-transformed image is obtained, including: The convolution module is used to extract features from the second scale-transformed image to obtain an initial feature matrix. The horizontal extraction module extracts features from the initial feature matrix in the horizontal direction to obtain a first feature matrix. The vertical extraction module extracts features from the initial feature matrix in the vertical direction to obtain a second feature matrix. Matrix operations are performed on the first feature matrix and the second feature matrix to obtain a set of intersection feature information corresponding to the second scale-transformed image.

9. The method according to claim 8, characterized in that, The step of performing matrix operations on the first feature matrix and the second feature matrix to obtain a set of intersection point feature information corresponding to the second scale-transformed image includes: Perform matrix addition on the first feature matrix and the second feature matrix to obtain the addition result; Perform matrix multiplication on the first feature matrix and the second feature matrix to obtain the multiplication result; Perform matrix division on the addition and multiplication results, and use the resulting division result as a set of intersection feature information corresponding to the second scale-transformed image.

10. The method according to claim 8, characterized in that, The convolutional module includes at least three convolutional layers; The horizontal extraction module includes H×1×W / 2 extractors, and the vertical extraction module includes 1×W×H / 2 extractors; where H represents the height of the second scale-transformed image and W represents the width of the second scale-transformed image.

11. The method according to any one of claims 1 to 10, characterized in that, In the recognition model, the Seg The number of CNN structures is six, and the Point The number of CNN structures is three.

12. The method according to claim 11, characterized in that, The six Seg CNN structure and the three points mentioned The CNN structure consists of a 3×3 matrix network, and the at least two scale-transformed images include a first scale-transformed image, a second scale-transformed image, and a third scale-transformed image; correspondingly, the step of inputting the at least two scale-transformed images into the recognition model and outputting the target feature information corresponding to the tongue image through the recognition model includes: The matrix network is used to extract and cross-process features from the first scale-transformed image, the second scale-transformed image, and the third scale-transformed image to obtain first tongue image feature information, second tongue image feature information, and third tongue image feature information. The target feature information is obtained by weighting the first tongue image feature information, the second tongue image feature information, and the third tongue image feature information.

13. A tongue image feature recognition device, characterized in that, It includes an acquisition unit, a scale operation unit, and a recognition unit; among which, The acquisition unit is configured to acquire a tongue image; The scale operation unit is configured to perform multi-scale operations on the tongue image to determine at least two scale-transformed images. The recognition unit is configured to input the at least two scale-transformed images into a recognition model, and output target feature information corresponding to the tongue image through the recognition model; wherein, the recognition model is based on Seg CNN structure and Point A cross-network model composed of CNN structures; It also includes a feature extraction unit, configured to utilize Seg The CNN structure extracts features from the first-scale transformed image, obtaining a set of line segment feature information corresponding to the first-scale transformed image; and utilizes Point... The CNN structure extracts features from the second scale-transformed image to obtain a set of intersection feature information corresponding to the second scale-transformed image; wherein the first scale-transformed image is different from the second scale-transformed image, and the first scale-transformed image and the second scale-transformed image are any two of at least two scale-transformed images.

14. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory is used to store computer programs that can run on the processor; The processor is configured to perform the method as described in any one of claims 1 to 12 when running the computer program.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by at least one processor, implements the method as described in any one of claims 1 to 12.