Human tissue ultrasound image processing method, apparatus, device, and medium

By using ResNet and UNet network models to identify ultrasound images of human tissue, the problem of ultrasound cosmetic devices being unable to identify tissue distribution has been solved, enabling rapid and accurate damage risk warnings and improving treatment safety and intelligence.

CN119399191BActive Publication Date: 2025-11-21NANJING JIXINGGE MEDICAL TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411978267.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-11-21
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing ultrasound cosmetic equipment cannot identify or recognize the tissue distribution at the treatment depth, leading to accidental damage to prohibited tissues such as bones and blood vessels, causing problems such as skin bruising, prolonged nerve pain, or numbness.

Method used

A hybrid neural network model combining ResNet and UNet is used to extract and enhance features from ultrasound images of human tissues, identify pixel location information of different human tissues, determine whether there are untreatable tissues in the current treatment area, and provide damage risk warnings.

Benefits of technology

It enables rapid and accurate identification of human tissues, avoids accidental damage, improves treatment safety, reduces side effects such as skin bruising and nerve pain, and enhances the intelligence level of ultrasound cosmetic equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399191B_ABST
    Figure CN119399191B_ABST
Patent Text Reader

Abstract

The application discloses a human tissue ultrasound image processing method and device, equipment and medium, which are applied to the technical field of medical ultrasound image processing, and comprises the following steps: acquiring a human tissue ultrasound image under a current treatment depth; performing feature extraction on the human tissue ultrasound image through a ResNet network to obtain human tissue feature images of different semantic levels; performing semantic segmentation on the human tissue feature images of different semantic levels based on a UNet network to obtain pixel position information of different human tissues; and when it is determined that there is prohibited treatment tissue under the current treatment depth based on the pixel position information of different human tissues, performing damage risk prompting based on the prohibited treatment tissue under the current treatment depth, so that the mixed model of the ResNet and the UNet can quickly and accurately identify different human tissues, and then can quickly and accurately judge the interference between the current treatment depth and the prohibited treatment tissue and prompt, thereby improving treatment safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical ultrasound image processing technology, and in particular to a method, apparatus, device and medium for processing ultrasound images of human tissue. Background Technology

[0002] Ultrasonic cosmetic devices utilize ultrasound technology for skin treatment. They emit ultrasound waves of specific frequencies, allowing them to penetrate the skin's surface and reach the dermis and even deeper tissue layers, stimulating high-frequency mechanical vibrations in the body's tissues to achieve various therapeutic effects. However, currently, due to the inability to identify or understand the tissue distribution at the treatment depth, there is a risk of accidentally damaging prohibited tissues such as bones and blood vessels, leading to skin bruising, prolonged nerve pain or numbness, and even more serious medical accidents. Summary of the Invention

[0003] This application provides a method, apparatus, device, and medium for processing ultrasound images of human tissue, to solve the problem that existing ultrasound cosmetic devices cannot identify or recognize the distribution of human tissue at the treatment depth, thus affecting treatment safety. The technical solution provided by this application is as follows:

[0004] On the one hand, this application provides a method for processing ultrasound images of human tissue, applied to ultrasound cosmetic equipment, including:

[0005] Acquire ultrasound images of human tissue in the current treatment area based on the target imaging depth; wherein the target imaging depth is greater than the initial treatment depth of the current treatment area;

[0006] Human tissue recognition is performed on ultrasound images of human tissues using a human tissue recognition model to obtain pixel location information of different human tissues in the ultrasound images. The human tissue recognition model includes a ResNet network and a UNet network connected in sequence. The human tissue recognition model uses the ResNet network to extract features from the ultrasound images of human tissues to obtain first human tissue feature images at different first semantic levels. Then, it uses the UNet network to perform feature enhancement processing on the first human tissue feature images at each first semantic level to obtain second human tissue feature images at different second semantic levels. Finally, semantic segmentation is performed on the second human tissue feature images at each second semantic level to obtain pixel location information of different human tissues in the ultrasound images of human tissues.

[0007] Based on the pixel location information of different human tissues in ultrasound images, when it is determined that there are untreatable tissues in the current treatment area at the initial treatment depth, a damage risk warning is given based on the untreatable tissues in the current treatment area at the initial treatment depth.

[0008] Optionally, the ResNet network includes multiple downsampling layers connected in sequence; the image resolution output by each downsampling layer decreases by a factor of two, the number of feature channels output by the first downsampling layer is the same as the number of feature channels output by the second downsampling layer, and the number of feature channels output by each subsequent downsampling layer increases by a factor of two.

[0009] The UNet network consists of an adaptation layer, multiple upsampling layers, and a fully connected layer connected sequentially. The multiple upsampling layers are connected to the multiple downsampling layers of the ResNet network in a one-to-one skip connection. The image resolution output by each upsampling layer increases exponentially while the number of feature channels decreases exponentially.

[0010] Optionally, feature extraction processing of human tissue ultrasound images is performed using a ResNet network to obtain first human tissue feature images at different first semantic levels, including:

[0011] By using the first downsampling layer in the ResNet network, the ultrasound images of human tissue are sequentially downsampled and dimensionality reduced to obtain a high-resolution first human tissue feature image.

[0012] By using the second and subsequent downsampling layers in the ResNet network, the high-resolution first human tissue feature image is iteratively downsampled to obtain first human tissue feature images at different low-resolution levels.

[0013] Based on the first human tissue feature image at a high resolution level and the first human tissue feature images at different low resolution levels, first human tissue feature images at different first semantic levels are obtained.

[0014] Optionally, after performing feature enhancement processing on the first human tissue feature images at each first semantic level using the UNet network to obtain second human tissue feature images at different second semantic levels, semantic segmentation is performed based on the second human tissue feature images at each second semantic level to obtain pixel location information of different human tissues in the human tissue ultrasound image, including:

[0015] The first human tissue feature image of the first semantic level output by the last downsampling layer in the ResNet network is adapted through the adaptation layer in the UNet network to obtain the first human tissue feature image of the first semantic level with the target feature channel number.

[0016] The first human tissue feature image of the first semantic level with the first number of target feature channels is upsampled through the first upsampling layer in the UNet network, and then fused with the first human tissue feature image of the first semantic level output by the downsampling layer in the corresponding skip connection ResNet network to obtain the second human tissue feature image of the second semantic level.

[0017] For the second and subsequent upsampling layers in the UNet network, the second human tissue feature image of the second semantic level output by the previous upsampling layer is upsampled and then fused with the first human tissue feature image of the first semantic level output by the downsampling layer in the corresponding skip connection ResNet network to obtain the second human tissue feature image of the second semantic level.

[0018] By using the fully connected layers in the UNet network, semantic segmentation is performed on the second human tissue feature image of the second semantic level output by the last upsampling layer to obtain the pixel location information of different human tissues in the human tissue ultrasound image.

[0019] Optionally, based on the pixel location information of different human tissues in the ultrasound image of human tissue, determine whether there are untreatable tissues in the current treatment area at the initial treatment depth, including:

[0020] Based on the target imaging depth, the pixel position information of different human tissues in the ultrasound image of human tissue is converted into the spatial depth information of different human tissues in the current treatment area.

[0021] Based on the spatial depth information of different human tissues in the current treatment area, when it is detected that there are prohibited treatment tissues in different human tissues whose spatial depth information overlaps with the initial treatment depth, it is determined that there are prohibited treatment tissues in the current treatment area at the initial treatment depth.

[0022] Optionally, when determining that there are untreatable tissues in the current treatment area at the initial treatment depth based on the pixel location information of different human tissues in the ultrasound image of human tissue, the method further includes:

[0023] Based on the spatial depth information of the prohibited tissues in the current treatment area at the initial treatment depth, the initial treatment depth of the current treatment area is adjusted to obtain the target treatment depth of the current treatment area.

[0024] Optionally, after performing human tissue recognition on the ultrasound image based on the human tissue recognition model to obtain the pixel location information of different human tissues in the ultrasound image, the method further includes:

[0025] The pixel location information of different human tissues in ultrasound images of human tissues is labeled and displayed.

[0026] On the other hand, this application provides a human tissue ultrasound image processing device, applied to ultrasound cosmetic equipment, comprising:

[0027] An ultrasound image acquisition unit is used to acquire ultrasound images of human tissue in the current treatment area detected according to the target imaging depth; wherein the target imaging depth is greater than the initial treatment depth of the current treatment area;

[0028] The human tissue recognition unit is used to identify human tissues in ultrasound images based on a human tissue recognition model, thereby obtaining pixel location information of different human tissues in the ultrasound images. The human tissue recognition model includes a ResNet network and a UNet network connected in sequence. The human tissue recognition model uses the ResNet network to perform feature extraction processing on the ultrasound images to obtain first human tissue feature images at different first semantic levels. Then, it uses the UNet network to perform feature enhancement processing on the first human tissue feature images at each first semantic level to obtain second human tissue feature images at different second semantic levels. Finally, it performs semantic segmentation based on the second human tissue feature images at each second semantic level to obtain pixel location information of different human tissues in the ultrasound images.

[0029] The damage risk warning unit is used to determine the presence of untreatable tissue in the current treatment area at the initial treatment depth based on the pixel location information of different human tissues in the ultrasound image of human tissue, and to provide a damage risk warning based on the untreatable tissue in the current treatment area at the initial treatment depth.

[0030] On the other hand, this application provides an ultrasonic cosmetic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned human tissue ultrasonic image processing method.

[0031] On the other hand, this application provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the above-described human tissue ultrasound image processing method.

[0032] The beneficial effects of this application are as follows:

[0033] This application, during skin treatment according to the initial treatment line defined for the treatment site, utilizes a human tissue recognition model based on ResNet and UNet networks to identify human tissue in the ultrasound image of the current treatment area at the initial treatment depth. This allows for rapid and accurate identification of different human tissues in the ultrasound image, enabling quick and precise determination of whether there are prohibited treatment tissues in the current treatment area at the initial treatment depth. Consequently, it can provide damage risk warnings based on these prohibited treatment tissues, effectively avoiding problems such as skin bruising, prolonged neuropathic pain or numbness, and more serious medical accidents caused by accidental damage to prohibited treatment tissues such as bones and blood vessels, thereby significantly improving treatment safety.

[0034] Other features and advantages of this application will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0035] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0036] Figure 1 This is a schematic flowchart illustrating the human tissue ultrasound image processing method in the embodiments of this application.

[0037] Figure 2 This is a schematic diagram illustrating the general structure of the human tissue recognition model in an embodiment of this application.

[0038] Figure 3 This is a schematic diagram of the specific structure of the human tissue recognition model in the embodiments of this application;

[0039] Figure 4 This is a schematic diagram outlining the training method for the human tissue recognition model in the embodiments of this application.

[0040] Figure 5 This is a functional structural diagram of the human tissue ultrasound image processing device in the embodiments of this application;

[0041] Figure 6 This is a schematic diagram of the hardware structure of the ultrasonic beauty device in the embodiments of this application. Detailed Implementation

[0042] To make the objectives, technical solutions, and beneficial effects of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] To facilitate a better understanding of this application by those skilled in the art, the technical terms used in this application will be briefly introduced below.

[0044] The human tissue recognition model is a neural network model used to identify different human tissues and their pixel location information in ultrasound images. Its neural network architecture is an improvement on the UNet network, a hybrid architecture based on ResNet (e.g., ResNet34) and UNet. The ResNet network, as the front-end network (encoder), can extract low-resolution features from ultrasound images more deeply without causing gradient vanishing and mesh degradation problems, and can also accelerate network convergence. The UNet network, as the back-end network (decoder), can fully extract high-resolution features (segmentation detail features) from ultrasound images through feature stitching and multi-scale fusion. This allows for the full fusion of different low-resolution features (deep features) and high-resolution features (shallow features) in ultrasound images, thus enabling more accurate identification of different human tissues and their pixel location information in ultrasound images.

[0045] The target imaging depth is the depth of ultrasound detection used when the ultrasound cosmetic device acquires ultrasound images of human tissue; in this application, the target imaging depth is a fixed value and is greater than the maximum treatment depth of the ultrasound cosmetic device, for example, the target imaging depth is 8mm.

[0046] The initial treatment depth is one of the different treatment depths selected from the ultrasonic cosmetic device according to the patient's treatment needs; in this application, the different treatment depths equipped with the ultrasonic cosmetic device include 1.5mm, 3mm, 4.5mm, etc., and the initial treatment depth is one of 1.5mm, 3mm, 4.5mm, etc.

[0047] The target treatment depth is the treatment depth obtained by adjusting the initial treatment depth based on the spatial depth information of the prohibited treatment tissues in the current treatment area identified by the human tissue recognition model at the initial treatment depth.

[0048] After introducing the technical terms used in this application, the application scenarios and design concepts of this application will be briefly introduced next.

[0049] The treatment principle of ultrasonic cosmetic equipment is to focus ultrasound waves at a specific depth within the target treatment area, stimulating high-frequency mechanical vibrations in the tissue to generate heat, approximately 60-70°C. If this heat is focused onto bone or blood vessels, it can cause side effects and should be avoided. Currently, to accommodate different treatment depths (i.e., ultrasound detection depths), ultrasonic cosmetic equipment is equipped with ultrasound probes for various treatment depths, such as 1.5mm, 3mm, and 4.5mm subcutaneously. However, during treatment using a probe at a specific depth, the inability to identify or understand the tissue distribution at that depth may lead to accidental damage to prohibited tissues such as bone or blood vessels, resulting in skin bruising, prolonged nerve pain or numbness, and even more serious medical accidents.

[0050] Therefore, in this application, the ultrasonic cosmetic device is equipped with an ultrasonic probe with a target imaging depth, which is greater than all treatment depths of the ultrasonic cosmetic device. For example, the target imaging depth is 8mm subcutaneously, enabling the ultrasonic probe to detect ultrasonic images of human tissue that meet all treatment depths. Based on this, when the ultrasonic cosmetic device is ready to treat the current treatment area of ​​a target patient, it can use the ultrasonic probe to detect the current treatment area according to the target imaging depth to obtain an ultrasonic image of the human tissue in the current treatment area. Then, through the ResNet network in the human tissue recognition model, feature extraction processing is performed on the human tissue ultrasonic image to obtain first human tissue feature images at different first semantic levels, and then through the U... The Net network performs feature enhancement processing on first human tissue feature images at different first semantic levels to obtain second human tissue feature images at different second semantic levels. Then, semantic segmentation is performed based on the second human tissue feature images at different second semantic levels. This enables the rapid and accurate acquisition of the location information of different human tissues in the ultrasound image. Subsequently, it can quickly and accurately determine whether there are prohibited tissues in the current treatment area at the initial treatment depth. Based on the prohibited tissues in the current treatment area at the initial treatment depth, it can provide damage risk warnings. This can effectively avoid problems such as skin bruising, prolonged neuropathic pain or numbness, and more serious medical accidents caused by accidental damage to prohibited tissues such as bones and blood vessels, thereby improving the safety of ultrasound treatment.

[0051] After introducing the application scenarios and design concepts of this application, the technical solutions provided by this application will be described in detail below.

[0052] This application provides a method for processing ultrasound images of human tissue, applied to ultrasound cosmetic equipment. (See attached document.) Figure 1 As shown, the general flow of the human tissue ultrasound image processing method provided in this application embodiment is as follows:

[0053] Step 101: Acquire an ultrasound image of the human tissue in the current treatment area detected according to the target imaging depth; wherein the target imaging depth is greater than the initial treatment depth of the current treatment area.

[0054] In this embodiment, the ultrasonic cosmetic device uses an ultrasonic probe with a target imaging depth to detect the current treatment area and obtain an ultrasonic image of the human tissue in the current treatment area. In one embodiment, the ultrasonic probe with the target imaging depth can be a new ultrasonic probe equipped with the ultrasonic cosmetic device, and this new ultrasonic probe has a fixed ultrasonic detection depth (i.e., the target imaging depth). In another embodiment, the ultrasonic probe with the target imaging depth can also be one of the original ultrasonic probes of the ultrasonic cosmetic device, for example, the ultrasonic probe with the largest treatment depth. When acquiring an ultrasonic image of human tissue, the servo motor corresponding to the ultrasonic probe with the largest treatment depth is driven to raise the ultrasonic detection depth of the ultrasonic probe with the largest treatment depth to the target imaging depth, thereby enabling the detection of the current treatment area according to the target imaging depth and obtaining an ultrasonic image of the human tissue in the current treatment area. In specific implementation, when acquiring the ultrasonic image of the human tissue in the current treatment area detected according to the target imaging depth, the waveform data obtained by the ultrasonic probe detecting the current treatment area according to the target imaging depth can be acquired first, and then the waveform data can be cleaned, Fourier transformed, and low-pass filtered to obtain the ultrasonic image of the human tissue in the current treatment area.

[0055] Step 102: Perform human tissue recognition on the ultrasound image of human tissue based on the human tissue recognition model to obtain the pixel location information of different human tissues in the ultrasound image of human tissue; wherein, the human tissue recognition model includes a ResNet network and a UNet network connected in sequence; the human tissue recognition model performs feature extraction processing on the ultrasound image of human tissue through the ResNet network to obtain first human tissue feature images at different first semantic levels, and performs feature enhancement processing on the first human tissue feature images at each first semantic level through the UNet network to obtain second human tissue feature images at different second semantic levels, and then performs semantic segmentation based on the second human tissue feature images at each second semantic level to obtain the pixel location information of different human tissues in the ultrasound image of human tissue.

[0056] In the embodiments of this application, see the following: Figure 2As shown, the human tissue recognition model includes a ResNet network and a UNet network connected in sequence. The ResNet network includes multiple downsampling layers connected in sequence. The image resolution output by each downsampling layer decreases by a factor of two. The number of feature channels output by the first downsampling layer is the same as that output by the second downsampling layer, and the number of feature channels output by each subsequent downsampling layer increases by a factor of two. The UNet network includes an adaptation layer, multiple upsampling layers, and a fully connected layer connected in sequence. The multiple upsampling layers are connected to the multiple downsampling layers of the ResNet network in a one-to-one skip connection. The image resolution output by each upsampling layer increases by a factor of two, and the number of feature channels decreases by a factor of two, thus enabling each upsampling layer in the UNet network to gradually restore the resolution level, ensuring that the input and output resolution levels of the human tissue recognition model are the same. Figure 2 This is merely an example of a human tissue recognition model. The number of downsampling layers in the ResNet network and the number of upsampling layers in the UNet network in the human tissue recognition model can be flexibly set according to actual needs, as long as the number of downsampling layers in the ResNet network is the same as the number of upsampling layers in the UNet network and they are connected in a one-to-one skip connection.

[0057] Based on such Figure 2 The human tissue recognition model shown, in the process of human tissue recognition based on ultrasound images, when performing feature extraction processing on the ultrasound images of human tissue through a ResNet network to obtain first human tissue feature images at different first semantic levels, can adopt, but is not limited to, the following methods:

[0058] First, the ultrasound images of human tissue are downsampled and dimensionality reduced sequentially through the first downsampling layer in the ResNet network to obtain a high-resolution first human tissue feature image.

[0059] Then, through the second and subsequent downsampling layers in the ResNet network, the high-resolution first human tissue feature image is iteratively downsampled to obtain first human tissue feature images at different low-resolution levels.

[0060] Finally, based on the high-resolution first human tissue feature image and the first human tissue feature images at different low-resolution levels, first human tissue feature images at different first semantic levels are obtained.

[0061] Based on such Figure 2The human tissue recognition model shown, in the process of recognizing human tissue in ultrasound images based on the human tissue recognition model, after performing feature enhancement processing on the first human tissue feature images of each first semantic level through the UNet network to obtain second human tissue feature images of different second semantic levels, and then performing semantic segmentation based on the second human tissue feature images of each second semantic level to obtain the pixel location information of different human tissues in the ultrasound image, can adopt, but is not limited to, the following methods:

[0062] First, the first human tissue feature image of the first semantic level output by the last downsampling layer in the ResNet network is adapted through the adaptation layer in the UNet network to obtain the first human tissue feature image of the first semantic level with the target feature channel number.

[0063] Next, the first human tissue feature image of the first semantic level with the first number of target feature channels is upsampled through the first upsampling layer in the UNet network, and then fused with the first human tissue feature image of the first semantic level output by the downsampling layer in the corresponding skip connection ResNet network to obtain the second human tissue feature image of the second semantic level.

[0064] Then, for each of the second and subsequent upsampling layers in the UNet network, the second human tissue feature image of the second semantic level output by the previous upsampling layer is upsampled and then fused with the first human tissue feature image of the first semantic level output by the downsampling layer in the corresponding skip connection ResNet network to obtain the second human tissue feature image of the second semantic level.

[0065] Finally, through the fully connected layers in the UNet network, semantic segmentation is performed on the second human tissue feature image of the second semantic level output by the last upsampling layer to obtain the pixel location information of different human tissues in the human tissue ultrasound image. The pixel location information of different human tissues can be represented by a human tissue distribution image; that is, the human tissue recognition model outputs a human tissue distribution image, in which the location regions of different human tissues are represented by different identifiers. For example, this human tissue distribution image includes superficial fascia tissue, deep fascia tissue, and vascular tissue. Each pixel in the location region of superficial fascia tissue is represented by identifier "1", each pixel in the location region of deep fascia tissue is represented by identifier "2", and each pixel in the location region of vascular tissue is represented by identifier "3".

[0066] Step 103: Based on the pixel location information of different human tissues in the ultrasound image of human tissue, if it is determined that there are untreatable tissues in the current treatment area at the initial treatment depth, a damage risk warning is given based on the untreatable tissues in the current treatment area at the initial treatment depth.

[0067] In this embodiment of the application, when determining that there are untreatable tissues in the current treatment area at the initial treatment depth based on the pixel location information of different human tissues in the ultrasound image of human tissue, the following methods may be used, but are not limited to:

[0068] First, based on the target imaging depth, the pixel position information of different human tissues in the ultrasound image of human tissue is converted into the spatial depth information of different human tissues in the current treatment area.

[0069] Then, based on the spatial depth information of different human tissues in the current treatment area, when it is detected that there are prohibited treatment tissues in different human tissues whose spatial depth information overlaps with the initial treatment depth, it is determined that there are prohibited treatment tissues in the current treatment area at the initial treatment depth.

[0070] In this way, by utilizing a human tissue recognition model based on ResNet and UNet networks, human tissue recognition can be performed on the ultrasound image of the current treatment area at the initial treatment depth within the initial treatment line. This allows for the rapid and accurate identification of different human tissues in the ultrasound image, enabling the quick and precise determination of whether there are any untreatable tissues in the current treatment area at the initial treatment depth. Consequently, damage risk warnings can be provided based on these untreatable tissues, for example, through pop-up windows and / or voice prompts. This can assist in the treatment process and improve the intelligence level of the ultrasound cosmetic equipment and the effectiveness of skin treatment.

[0071] Furthermore, in this embodiment, after determining that there are prohibited treatment tissues in the current treatment area at the initial treatment depth based on the pixel position information of different human tissues in the ultrasound image of human tissue, on the one hand, the initial treatment depth of the current treatment area can be adjusted based on the spatial depth information of the prohibited treatment tissues at the initial treatment depth to obtain the target treatment depth of the current treatment area so as to avoid the depth range where the prohibited treatment tissues are located; on the other hand, the pixel position information and / or spatial depth information of different human tissues can be marked and displayed to facilitate the operator to view the distribution of different human tissues in the current treatment area and adjust the initial treatment depth, thereby further improving the intelligence level and auxiliary treatment effect of the ultrasound cosmetic equipment.

[0072] The following section provides a detailed explanation of the human tissue recognition model and the human tissue ultrasound image processing method based on the human tissue recognition model. (See also...) Figure 3 As shown, the human tissue recognition model includes a ResNet network (e.g., a ResNet34 network) and a UNet network connected in sequence; wherein:

[0073] The ResNet network is a process of continuously fitting residuals. By learning the residuals between the input and output, the network converges faster and achieves higher recognition accuracy. As the encoder part of an encoder-decoder structure, the ResNet network consists of multiple sequentially connected downsampling layers. Figure 3 Taking the ResNet network, consisting of 5 downsampling layers, as an example, the first downsampling layer consists of one convolutional layer and one pooling layer. The convolutional layer has 64 convolutional kernels, a size of 7x7, and a stride of 2. The pooling layer has 64 pooling windows, a size of 2x2, and a stride of 2. Each of the remaining downsampling layers consists of one convolutional layer, and the number of residual blocks in the convolutional layers of each of the remaining downsampling layers may be the same or different. Specifically, the second downsampling layer contains three residual blocks, each with 64 convolutional kernels, a size of 3x3, and a stride of 1. The third downsampling layer contains four residual blocks. The first residual block has 128 convolutional kernels, a size of 3x3, and a stride of 2. The second to fourth residual blocks each have 128 convolutional kernels. 8. The size is 3x3 and the stride is 1. The fourth downsampling layer contains a convolutional layer consisting of 6 residual blocks. The first residual block has 256 convolutional kernels, a size of 3x3, and a stride of 2. The second to sixth residual blocks also have 256 convolutional kernels, a size of 3x3, and a stride of 1. The fifth downsampling layer contains a convolutional layer consisting of 3 residual blocks. The first residual block has 512 convolutional kernels, a size of 3x3, and a stride of 2. The second and third residual blocks also have 512 convolutional kernels, a size of 3x3, and a stride of 1.

[0074] Based on such Figure 3 The human tissue recognition model shown below illustrates the specific process by which the ResNet network in the model extracts features from the ultrasound images of human tissue after they are input into the model:

[0075] First, in the initial downsampling layer, the ultrasound image of human tissue (taking only an 800x800x1 grayscale ultrasound image of human tissue as an example) is convolved with a size of 7x7 and a stride of 2 through a convolutional layer, resulting in 64 first human tissue feature images with a dimension of 400x400x1. Then, these 64 first human tissue feature images with a dimension of 400x400x1 are passed through a pooling layer for dimensionality reduction, outputting 64 first human tissue feature images with a dimension of 200x200x1. This method, by connecting a pooling layer, reduces data dimensionality, lowers computational complexity, and also reduces the risk of overfitting.

[0076] Next, in the second downsampling layer, the 64 images of the first human tissue feature with a dimension of 200x200x1 are sequentially passed through the three residual blocks in the convolutional layer for convolution operation with a size of 3x3 and a stride of 1. At the same time, short connection operation is performed between each residual block to perform residual learning between residual blocks, and the output is the 64 images of the first human tissue feature with a dimension of 200x200x1.

[0077] Then, in the third downsampling layer, 64 images of the first human tissue feature with dimensions of 200x200x1 are convolved with a size of 3x3 and a stride of 2 through the first residual block in the convolutional layer. Then, they are convolved with a size of 3x3 and a stride of 1 through the second to fourth residual blocks in the convolutional layer. At the same time, short connection operations are performed between each residual block to perform residual learning between residual blocks, and 128 images of the first human tissue feature with dimensions of 100x100x1 are output.

[0078] Secondly, in the fourth downsampling layer, 128 images of the first human tissue feature with a dimension of 100x100x1 are convolved with a size of 3x3 and a stride of 2 through the first residual block in the convolutional layer. Then, they are convolved with a size of 3x3 and a stride of 1 through the second to sixth residual blocks in the convolutional layer. At the same time, short connection operations are performed between each residual block to perform residual learning between residual blocks, and 128 images of the first human tissue feature with a dimension of 50x50x1 are output.

[0079] Finally, in the fifth downsampling layer, 128 images of the first human tissue feature with a dimension of 50x50x1 are convolved with a size of 3x3 and a stride of 2 through the first residual block in the convolutional layer. Then, they are convolved with a size of 3x3 and a stride of 1 through the second to third residual blocks in the convolutional layer. At the same time, short connection operations are performed between each residual block to perform residual learning between residual blocks, and 512 images of the first human tissue feature with a dimension of 25x25x1 are output.

[0080] In this way, after the deep convolution operation of the ResNet network, sufficiently macroscopic and abstract human tissue features can be extracted. That is, the first human tissue feature images at different first semantic levels can be obtained through the deep convolution operation of the ResNet network.

[0081] The UNet network is a continuous feature concatenation process. It gradually restores image resolution through a series of deconvolutions (also known as transposed convolutions or upsampling) while simultaneously fusing features through a series of skip connections, resulting in higher accuracy in the output human tissue distribution. As the decoder part of an encoder-decoder structure, the UNet network consists of an adaptation layer, multiple upsampling layers, and a fully connected layer connected sequentially. Figure 3 Taking the UNet network, which consists of one adaptation layer, five upsampling layers, and one fully connected layer, as an example, the adaptation layer consists of two convolutional layers. The number of convolutional kernels in the first convolutional layer is one of 512, 768, or 1024. Figure 3Taking 1024 kernels as an example, the size is 3x3 with a stride of 1. The second convolutional layer has 1024 kernels, a size of 3x3, and a stride of 1. The first upsampling layer consists of one deconvolutional layer, one concatenation layer, and two convolutional layers. The concatenation layer is skipped to the fifth downsampling layer in the ResNet network. The deconvolutional layer has 1024 kernels, a size of 2x2, and a stride of 2. The two convolutional layers have 512 kernels, a size of 3x3, and a stride of 1. The second upsampling layer consists of one deconvolutional layer, one concatenation layer, and two convolutional layers. The concatenation layer is skipped to the fourth downsampling layer in the ResNet network. The deconvolutional layer has 512 kernels, a size of 2x2, and a stride of 1. 2. The stride is 2. The two convolutional layers have 256 kernels each, with a size of 3x3 and a stride of 1. The third upsampling layer consists of one deconvolutional layer, one concatenation layer, and two convolutional layers. The concatenation layer is skip-connected to the third downsampling layer in the ResNet network. The deconvolutional layer has 256 kernels each, with a size of 2x2 and a stride of 2. The two convolutional layers have 128 kernels each, with a size of 3x3 and a stride of 1. The fourth upsampling layer consists of one deconvolutional layer, one concatenation layer, and two convolutional layers. The concatenation layer is skip-connected to the second downsampling layer in the ResNet network. The deconvolutional layer has 128 kernels each, with a size of 2x3 and a stride of 1. 2. The stride is 2. The number of convolutional kernels in the two convolutional layers is 64, the size is 3x3, and the stride is 1. The fifth upsampling layer consists of one deconvolutional layer, one concatenation layer, and two convolutional layers. The concatenation layer is skip-connected to the first downsampling layer in the ResNet network. The number of convolutional kernels in the deconvolutional layer is 64, the size is 2x2, and the stride is 2. The number of convolutional kernels in the two convolutional layers is 32, the size is 3x3, and the stride is 1. The fully connected layer consists of one point convolutional layer and one multi-class label mapping layer. The size of the convolutional kernel in the point convolutional layer is 1x1, and the stride is 1.

[0082] Based on such Figure 3 The human tissue recognition model shown in the diagram extracts features from ultrasound images of human tissue using the ResNet network and outputs first human tissue feature images at different first semantic levels. The UNet network in the same model then performs feature enhancement on these first human tissue feature images at different first semantic levels before semantic segmentation. The specific process for semantic segmentation is as follows:

[0083] First, in the adaptation layer, the 512 first human tissue feature images with a dimension of 25x25x1 output from the last downsampling layer in the ResNet network are sequentially passed through two convolutional layers with a size of 3x3 and a stride of 1, outputting 1024 first human tissue feature images with a dimension of 25x25x1.

[0084] Next, in the first upsampling layer, 1024 images of the first human tissue feature with dimensions of 25x25x1 are upsampled by a deconvolution layer with a size of 2x2 and a stride of 2 to obtain 1024 images of the first human tissue feature with dimensions of 50x50x1. These images are then fused with 128 images of the first human tissue feature with dimensions of 50x50x1 output from the fifth downsampling layer in the ResNet network through a concatenation layer. Subsequently, two convolutional layers are used to perform convolution operations with a size of 3x3 and a stride of 1, outputting 512 images of the second human tissue feature with dimensions of 50x50x1.

[0085] Then, in the second upsampling layer, the 512 images of the second human tissue feature with a dimension of 50x50x1 are upsampled by a deconvolution layer with a size of 2 x 2 and a stride of 2 to obtain 512 images of the human tissue feature with a dimension of 100x100x1. These images are then fused with the 256 images of the first human tissue feature with a dimension of 100x100x1 output from the fourth downsampling layer in the ResNet network through a concatenation layer. Subsequently, the images are passed through two convolutional layers with a size of 3x3 and a stride of 1 to output 256 images of the second human tissue feature with a dimension of 100x100x1.

[0086] Secondly, in the third upsampling layer, 256 images of the second human tissue feature with a dimension of 100x100x1 are upsampled by a deconvolution layer with a size of 2 x 2 and a stride of 2 to obtain 256 images of the second human tissue feature with a dimension of 200x200x1. These images are then fused with the 128 images of the first human tissue feature with a dimension of 200x200x1 output from the third downsampling layer in the ResNet network through a concatenation layer. Subsequently, the images are passed through two convolutional layers with a size of 3x3 and a stride of 1 to output 128 images of the second human tissue feature with a dimension of 200x200x1.

[0087] Subsequently, in the fourth upsampling layer, 128 images of the second human tissue feature with a dimension of 200x200x1 are upsampled by a deconvolution layer with a size of 2 x 2 and a stride of 2 to obtain 128 images of the human tissue feature with a dimension of 400x400x1. These images are then fused with the 64 images of the first human tissue feature with a dimension of 400x400x1 output from the second downsampling layer in the ResNet network through a concatenation layer. Afterward, the images are sequentially convolved by the kernels in two convolutional layers with a size of 3x3 and a stride of 1, outputting 64 images of the second human tissue feature with a dimension of 400x400x1.

[0088] Subsequently, in the fifth upsampling layer, 64 images of the second human tissue feature with a dimension of 400x400x1 are upsampled by a deconvolution layer with a size of 2 x 2 and a stride of 2 to obtain 64 images of the second human tissue feature with a dimension of 800x800x1. These images are then fused with the 32 images of the first human tissue feature with a dimension of 800x800x1 output from the first downsampling layer in the ResNet network through a concatenation layer. Afterward, the images are sequentially convolved by the kernels in two convolutional layers with a size of 3x3 and a stride of 1, outputting 32 images of the second human tissue feature with a dimension of 800x800x1.

[0089] Finally, in the fully connected layer, the 32 second human tissue feature images with dimensions of 800x800x1 are subjected to point convolution operations of size 1x1 and stride 1 through the point convolution layer, and then semantic segmentation is performed through the multi-class label mapping layer to output a human tissue distribution image that represents the pixel position information of different human tissues in the human tissue ultrasound image; wherein, the pixel position information of different human tissues in the human tissue distribution image is represented by different identifiers.

[0090] In this embodiment, after identifying the pixel location information of different human tissues in the ultrasound image of human tissue using a human tissue recognition model, the pixel location information of different human tissues in the ultrasound image of human tissue is converted into spatial depth information of different human tissues in the current treatment area based on the target imaging depth. Based on the spatial depth information of different human tissues in the current treatment area, when it is detected that there are prohibited treatment tissues in different human tissues whose spatial depth information overlaps with the initial treatment depth, it is determined that there are prohibited treatment tissues in the current treatment area at the initial treatment depth. Damage risk warnings for prohibited treatment tissues in the current treatment area at the initial treatment depth are provided through pop-up windows and / or voice prompts. At the same time, the pixel location information and / or spatial depth information of different human tissues are marked and displayed to facilitate the operator to view the distribution of human tissues in the current treatment area and adjust the initial treatment depth. In addition, after determining that there are prohibited treatment tissues in the current treatment area at the initial treatment depth, the initial treatment depth of the current treatment area can be adjusted based on the spatial depth information of the prohibited treatment tissues in the current treatment area to obtain the target treatment depth of the current treatment area to avoid the depth range where the prohibited treatment tissues are located. This enables automatic adjustment of the treatment depth, improving the treatment effect and intelligence of the ultrasound cosmetic equipment.

[0091] It is worth mentioning that, in the embodiments of this application, the ultrasonic cosmetic device usually predetermines an initial treatment line for the target treatment area (e.g., the face) of the target patient. The initial treatment line includes different treatment areas of the target treatment area and the initial treatment depth corresponding to each treatment area. The current treatment area mentioned above can be the treatment area currently undergoing skin treatment in the initial treatment line, or it can be the next treatment area in the initial treatment line that will undergo skin treatment. That is, while treating one of the treatment areas in the initial treatment line, interference detection between the initial treatment depth and the prohibited treatment tissue is performed on the next treatment area, thereby improving both the treatment effect and the treatment efficiency.

[0092] The training method for the human tissue recognition model in the embodiments of this application will be described in detail below. (See reference...) Figure 4 As shown, the specific process of the human tissue recognition model training method provided in this application embodiment is as follows:

[0093] Step 401: Obtain a sample set of the target human body part (e.g., face); wherein, the sample set includes multiple sample groups, each sample group including the original human tissue ultrasound image and standard mask image of the target human body part (e.g., face), the standard mask image is a binary mask image or a grayscale mask image; each pixel value in the binary mask image is 0 or 1, 1 represents the target human tissue, and 0 represents other human tissues besides the target human tissue; each pixel value in the grayscale mask image is an integer value between 0 and 255, different integer values ​​represent different human tissues.

[0094] In practical implementation, when obtaining a sample set of the target human body part (e.g., the face), the following methods can be used, but are not limited to:

[0095] First, multiple raw human tissue ultrasound images are acquired; each raw human tissue ultrasound image contains relatively rich information on various human tissues such as bones, blood vessels, and dermis under the skin of the target human body part (e.g., the face).

[0096] Then, using Labelme software, pixel location information of different human tissues is labeled for each original human tissue ultrasound image to generate a standard mask image. Taking the standard mask image as a binary mask image as an example, each pixel in the location area of ​​the target human tissue in the original human tissue ultrasound image is marked as 1, and each pixel in the location area of ​​the other human tissues is marked as 0, thereby generating a binary mask image with the same image size as the original human tissue ultrasound image. The target human tissue label is added to the binary mask image. For example, assuming that the target human tissue is bone, the bone label is added to the binary mask image for multi-label mapping processing (i.e., semantic segmentation).

[0097] Finally, multiple sample sets are formed based on multiple original human tissue ultrasound images and standard mask images of multiple original human tissue ultrasound images to obtain a sample set of the target human body parts (e.g., the face).

[0098] Step 402: After performing image enhancement processing on the sample set of target human body parts (e.g., face), the image enhancement sample set is randomly divided into a training sample set and a test sample set.

[0099] Step 403: Train the human tissue recognition model based on the training sample set.

[0100] In practical implementation, when training the human tissue recognition model based on the training sample set, the following methods can be used, but are not limited to:

[0101] The initial human tissue recognition model is iteratively trained based on the training sample set until the iteration termination condition is met (e.g., the number of iterations is greater than a first threshold and / or the current loss value is less than a second threshold, etc.). Then, the human tissue recognition model is obtained based on the weights of the ResNet and UNet networks in the initial human tissue recognition model updated during the last training operation. Each training operation includes:

[0102] Select the target sample group from the training sample set;

[0103] The original human tissue ultrasound images in the target sample group are input into the initial human tissue recognition model to obtain a human tissue distribution image that represents the pixel location information of different human tissues in the original human tissue ultrasound images.

[0104] The current loss value is calculated based on the human tissue distribution image and the standard mask image in the target sample group;

[0105] Update the weights of the ResNet and UNet networks in the initial human tissue recognition model based on the current loss value.

[0106] Step 404: Evaluate the trained human tissue recognition model based on the test sample set.

[0107] Step 405: The human tissue recognition model that passes the evaluation is identified as the final human tissue recognition model and saved to the computer storage medium.

[0108] Based on the above embodiments, this application provides a human tissue ultrasound image processing device, applied to ultrasound cosmetic equipment, see reference. Figure 5 As shown, the human tissue ultrasound image processing apparatus 500 provided in this application embodiment includes at least:

[0109] The ultrasound image acquisition unit 501 is used to acquire ultrasound images of human tissue in the current treatment area detected according to the target imaging depth; wherein the target imaging depth is greater than the initial treatment depth of the current treatment area;

[0110] The human tissue recognition unit 502 is used to perform human tissue recognition on human tissue ultrasound images based on a human tissue recognition model to obtain pixel location information of different human tissues in the human tissue ultrasound images. The human tissue recognition model includes a ResNet network and a UNet network connected in sequence. The human tissue recognition model performs feature extraction processing on the human tissue ultrasound images through the ResNet network to obtain first human tissue feature images at different first semantic levels. After performing feature enhancement processing on the first human tissue feature images at each first semantic level through the UNet network to obtain second human tissue feature images at different second semantic levels, semantic segmentation is performed based on the second human tissue feature images at each second semantic level to obtain pixel location information of different human tissues in the human tissue ultrasound images.

[0111] The damage risk warning unit 503 is used to provide a damage risk warning based on the pixel location information of different human tissues in the ultrasound image of human tissue, when it is determined that there are untreatable tissues in the current treatment area at the initial treatment depth.

[0112] In one possible implementation, the ResNet network includes multiple downsampling layers connected in sequence; the image resolution output by each downsampling layer decreases by a factor of two, the number of feature channels output by the first downsampling layer is the same as the number of feature channels output by the second downsampling layer, and the number of feature channels output by the second and subsequent downsampling layers increases by a factor of two.

[0113] The UNet network consists of an adaptation layer, multiple upsampling layers, and a fully connected layer connected sequentially. The multiple upsampling layers are connected to the multiple downsampling layers of the ResNet network in a one-to-one skip connection. The image resolution output by each upsampling layer increases exponentially while the number of feature channels decreases exponentially.

[0114] In one possible implementation, the human tissue recognition unit 502 is used to sequentially downsample and reduce the dimensionality of the human tissue ultrasound image through the first downsampling layer in the ResNet network to obtain a high-resolution first human tissue feature image; through the second and subsequent downsampling layers in the ResNet network, the high-resolution first human tissue feature image is iteratively downsampled to obtain first human tissue feature images of different low-resolution levels; based on the high-resolution first human tissue feature image and the first human tissue feature images of different low-resolution levels, first human tissue feature images of different first semantic levels are obtained.

[0115] In one possible implementation, the human tissue recognition unit 502 is used to adapt the first human tissue feature image of the first semantic level output by the last downsampling layer in the ResNet network through the adaptation layer in the UNet network to obtain the first human tissue feature image of the first semantic level with the target number of feature channels; after upsampling the first human tissue feature image of the first semantic level with the target number of feature channels through the first upsampling layer in the UNet network, it performs feature fusion processing with the first human tissue feature image of the first semantic level output by the downsampling layer in the ResNet network with the corresponding skip connection to obtain the second semantic level. The second human tissue feature image is obtained by upsampling the second semantic level human tissue feature image output by the previous upsampling layer through the upsampling layer and then fusing it with the first semantic level human tissue feature image output by the downsampling layer in the corresponding skip connection ResNet network to obtain the second semantic level human tissue feature image. The second semantic level human tissue feature image is then semantically segmented through the fully connected layer in the UNet network to obtain the pixel location information of different human tissues in the human tissue ultrasound image.

[0116] In one possible implementation, the damage risk warning unit 503 is used to convert the pixel position information of different human tissues in the ultrasound image of human tissue into spatial depth information of different human tissues in the current treatment area based on the target imaging depth; and when it detects that there are prohibited treatment tissues in different human tissues whose spatial depth information overlaps with the initial treatment depth based on the spatial depth information of different human tissues in the current treatment area, it determines that there are prohibited treatment tissues in the current treatment area at the initial treatment depth.

[0117] In one possible implementation, the human tissue ultrasound image processing apparatus 500 provided in this application embodiment further includes:

[0118] The treatment line adjustment unit 504 is used to adjust the initial treatment depth of the current treatment area based on the spatial depth information of the prohibited treatment tissue in the current treatment area at the initial treatment depth to obtain the target treatment depth of the current treatment area.

[0119] In one possible implementation, the human tissue ultrasound image processing apparatus 500 provided in this application embodiment further includes:

[0120] The annotation display unit 505 is used to annotate and display the pixel position information of different human tissues in ultrasound images of human tissues.

[0121] It should be noted that the principle of the human tissue ultrasound image processing device 500 provided in the embodiments of this application to solve the technical problem is similar to that of the human tissue ultrasound image processing method provided in the embodiments of this application. Therefore, the implementation of the human tissue ultrasound image processing device 500 provided in the embodiments of this application can refer to the implementation of the human tissue ultrasound image processing method provided in the embodiments of this application, and the repeated parts will not be described again.

[0122] After introducing the human tissue ultrasound image processing method and apparatus provided in the embodiments of this application, the ultrasound cosmetic device provided in the embodiments of this application will be briefly introduced next.

[0123] See Figure 6 As shown, the ultrasonic beauty device 600 provided in this application embodiment includes at least a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program, it implements the above-mentioned human tissue ultrasonic image processing method provided in this application embodiment.

[0124] The ultrasonic beauty device 600 provided in this application embodiment may further include a bus 603 connecting different components (including a processor 601 and a memory 602). The bus 603 represents one or more types of bus structures, including a memory bus, a peripheral bus, a local area bus, etc.

[0125] Memory 602 may include readable media in the form of volatile memory, such as random access memory 6021 and / or cache memory 6022, and may further include read-only memory 6023. Memory 602 may also include program tools 6025 having a set (at least one) of program modules 6024, including but not limited to operating subsystems, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0126] Processor 601 can be a single processing element or a collective term for multiple processing elements. For example, processor 601 can be a microcontroller unit (MCU), a central processing unit (CPU), or one or more integrated circuits configured to implement the above-described human tissue ultrasound image processing method provided in the embodiments of this application. Specifically, processor 601 can be a general-purpose processor, including but not limited to CPUs, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0127] The ultrasonic beauty device 600 can also communicate with one or more devices that allow a user to interact with the ultrasonic beauty device 600 (e.g., mobile phones, computers, etc.), and / or with various external devices 604 such as devices that enable the ultrasonic beauty device 600 to communicate with one or more other ultrasonic beauty devices (e.g., routers, modems, etc.). This communication can be performed through input / output interfaces 605. Furthermore, the ultrasonic beauty device 600 can also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via network adapter 606. Figure 6 As shown, network adapter 606 communicates with other modules of the ultrasonic beauty device 600 via bus 603. It should be understood that, although... Figure 6 As not shown, other hardware and / or software modules can be used in conjunction with the ultrasound beauty device 600, including but not limited to microcode, device drivers, redundant processors, external disk drive arrays, Redundant Arrays of Independent Disks (RAID) subsystems, tape drives, and data backup storage subsystems.

[0128] It should be noted that, Figure 6 The ultrasonic beauty device 600 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0129] Furthermore, this application embodiment also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the human tissue ultrasound image processing method provided in this application embodiment. Specifically, the computer instructions can be built into or installed in a processor, so that the processor can implement the above-mentioned human tissue ultrasound image processing method provided in this application embodiment by executing the built-in or installed computer instructions.

[0130] In addition, the human tissue ultrasound image processing method provided in this application embodiment can also be implemented as a program product, which includes program code. When the program code is executed by a processor, it implements the above-mentioned human tissue ultrasound image processing method provided in this application embodiment.

[0131] The program product provided in this application embodiment can be any combination of one or more readable media, wherein the readable media can be a readable signal medium or a readable storage medium, and the readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. Specifically, more specific examples of readable storage media (a non-exhaustive list) include electrical connections with one or more wires, portable disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] The program product provided in this application embodiment can be a CD-ROM and include program code, and can also run on an ultrasonic beauty device. However, the program product provided in this application embodiment is not limited to this. In this application embodiment, the readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or apparatus.

[0133] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0134] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0135] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0136] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.

Claims

1. A method for processing ultrasound images of human tissue, characterized in that, Applications in ultrasonic cosmetic equipment include: An initial treatment line is determined for the target patient's face, and the current treatment area is determined from the initial treatment line; wherein, the initial treatment line includes different treatment areas of the target patient's face and an initial treatment depth corresponding to each treatment area; the initial treatment depth is one of the treatment depths corresponding to different ultrasonic probes equipped in the ultrasonic cosmetic device; By driving the servo motor corresponding to the ultrasound probe with the largest treatment depth, the ultrasound detection depth of the ultrasound probe with the largest treatment depth is increased to the target imaging depth. By controlling the ultrasound probe with the largest treatment depth to detect the current treatment area according to the target imaging depth, an ultrasound image of human tissue in the current treatment area is obtained; wherein, the target imaging depth is greater than the initial treatment depth of the current treatment area. Human tissue recognition is performed on the ultrasound image of the human tissue based on a human tissue recognition model to obtain a human tissue distribution image of the current treatment area; wherein, the pixel position information of different human tissues in the human tissue distribution image is represented by different identifiers; the human tissue recognition model includes a ResNet network and a UNet network connected in sequence; the ResNet network includes five downsampling layers connected in sequence; the UNet network includes an adaptation layer, five upsampling layers, and a fully connected layer connected in sequence; the five upsampling layers and the five downsampling layers are connected in a one-to-one skip connection; the human tissue recognition model is implemented through the ResNet network. The human tissue ultrasound images are processed to extract features to obtain first human tissue feature images at different first semantic levels. Then, through the UNet network, feature enhancement processing is performed on each of the first human tissue feature images at the first semantic level to obtain second human tissue feature images at different second semantic levels. Finally, multi-label mapping is performed on each of the second human tissue feature images at the second semantic level to obtain the human tissue distribution image. The human tissue recognition model is trained using multiple original human tissue ultrasound images of the human face and multiple binary mask images of different human tissues corresponding to the original human tissue ultrasound images, each with added human tissue labels. Based on the target imaging depth, the pixel position information of different human tissues in the ultrasound image of human tissue is converted into the spatial depth information of the different human tissues in the current treatment area. Based on the spatial depth information of the different human tissues in the current treatment area, when it is detected that there are prohibited treatment tissues in the different human tissues whose spatial depth information overlaps with the initial treatment depth, it is determined that there are prohibited treatment tissues in the current treatment area at the initial treatment depth. Based on the prohibited treatment tissues in the current treatment area at the initial treatment depth, a damage risk warning is provided. Based on the spatial depth information of the prohibited treatment tissues in the current treatment area at the initial treatment depth, the initial treatment depth of the current treatment area is adjusted to obtain the target treatment depth of the current treatment area. The pixel position information and / or spatial depth information of different human tissues in the ultrasound image of human tissue are marked and displayed so that the operator can view the distribution of human tissues in the current treatment area and adjust the initial treatment depth.

2. The method for processing human tissue ultrasound images as described in claim 1, characterized in that, The image resolution output by each downsampling layer decreases by a factor of two; the number of feature channels output by the first downsampling layer is the same as the number of feature channels output by the second downsampling layer; and the number of feature channels output by the second and subsequent downsampling layers increases by a factor of two. The image resolution output by each upsampling layer increases by a factor of two, and the number of feature channels decreases by a factor of two.

3. The method for processing human tissue ultrasound images as described in claim 1, characterized in that, The ResNet network is used to perform feature extraction processing on the ultrasound images of human tissue to obtain first human tissue feature images at different first semantic levels, including: The ultrasound image of human tissue is downsampled and dimensionality reduced sequentially through the first downsampling layer in the ResNet network to obtain a high-resolution first human tissue feature image. The high-resolution first human tissue feature image is iteratively downsampled through the second and subsequent downsampling layers in the ResNet network to obtain first human tissue feature images at different low-resolution levels. Based on the first human tissue feature image at the high resolution level and the first human tissue feature images at different low resolution levels, first human tissue feature images at different first semantic levels are obtained.

4. The method for processing human tissue ultrasound images as described in claim 1, characterized in that, After performing feature enhancement processing on the first human tissue feature images at each of the first semantic levels using the UNet network to obtain second human tissue feature images at different second semantic levels, semantic segmentation is performed based on the second human tissue feature images at each of the second semantic levels to obtain pixel location information of different human tissues in the human tissue ultrasound image, including: The first human tissue feature image of the first semantic level output by the last downsampling layer in the ResNet network is adapted through the adaptation layer in the UNet network to obtain the first human tissue feature image of the first semantic level with the target feature channel number. The first human tissue feature image of the first semantic level with the target feature channel number is upsampled through the first upsampling layer in the UNet network, and then fused with the first human tissue feature image of the first semantic level output by the downsampling layer in the ResNet network with the corresponding skip connection to obtain the second human tissue feature image of the second semantic level. For the second and subsequent upsampling layers in the UNet network, the second human tissue feature image of the second semantic level output by the previous upsampling layer is upsampled and then fused with the first human tissue feature image of the first semantic level output by the downsampling layer in the ResNet network with the corresponding skip connection to obtain the second human tissue feature image of the second semantic level. The pixel location information of different human tissues in the ultrasound image is obtained by semantic segmentation of the second human tissue feature image of the second semantic level output by the last upsampling layer through the fully connected layer in the UNet network.

5. A human tissue ultrasound image processing device, characterized in that, Applications in ultrasonic cosmetic equipment include: A treatment area determination unit is used to determine an initial treatment line for the target patient's face and to determine the current treatment area from the initial treatment line; wherein, the initial treatment line includes different treatment areas of the target patient's face and an initial treatment depth corresponding to each treatment area; the initial treatment depth is one of the treatment depths corresponding to different ultrasonic probes equipped in the ultrasonic cosmetic device; An ultrasound image acquisition unit is used to drive a servo motor corresponding to the ultrasound probe with the largest treatment depth to raise the ultrasound detection depth of the ultrasound probe with the largest treatment depth to a target imaging depth, and to control the ultrasound probe with the largest treatment depth to detect the current treatment area according to the target imaging depth to acquire an ultrasound image of human tissue in the current treatment area; wherein, the target imaging depth is greater than the initial treatment depth of the current treatment area. A human tissue recognition unit is used to perform human tissue recognition on the ultrasound image of the human tissue based on a human tissue recognition model, thereby obtaining a human tissue distribution image of the current treatment area. The pixel location information of different human tissues in the human tissue distribution image is represented by different identifiers. The human tissue recognition model includes a ResNet network and a UNet network connected in sequence. The ResNet network includes five downsampling layers connected in sequence. The UNet network includes an adaptation layer, five upsampling layers, and a fully connected layer connected in sequence. The five upsampling layers and the five downsampling layers are connected in a one-to-one skip connection. The human tissue recognition model uses the ResNet network to identify human tissues in the ultrasound image of the current treatment area. The sNet network extracts features from the ultrasound images of human tissue to obtain first human tissue feature images at different first semantic levels. The UNet network then enhances the features of each first human tissue feature image at the first semantic level to obtain second human tissue feature images at different second semantic levels. Finally, multi-label mapping is performed on each second human tissue feature image at the second semantic level to obtain the human tissue distribution image. The human tissue recognition model is trained using multiple original ultrasound images of human face tissue and binary mask images of different human tissues corresponding to each of the original ultrasound images, each with added human tissue labels. The damage risk warning unit is used to convert the pixel position information of different human tissues in the ultrasound image of human tissue into spatial depth information of the different human tissues in the current treatment area based on the target imaging depth; when it detects that there are prohibited treatment tissues in the different human tissues whose spatial depth information overlaps with the initial treatment depth, it determines that there are prohibited treatment tissues in the current treatment area below the initial treatment depth; and provides a damage risk warning based on the prohibited treatment tissues in the current treatment area below the initial treatment depth. The treatment line adjustment unit is used to adjust the initial treatment depth of the current treatment area based on the spatial depth information of the prohibited tissue in the current treatment area at the initial treatment depth to obtain the target treatment depth of the current treatment area; The annotation display unit is used to annotate and display the pixel position information and / or spatial depth information of different human tissues in the ultrasound image of human tissue, so that the operator can view the distribution of human tissues in the current treatment area and adjust the initial treatment depth.

6. An ultrasonic cosmetic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the human tissue ultrasound image processing method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the human tissue ultrasound image processing method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Ultrasonic system and method based on image guidance

    CN117618811A

  • Ultrasonic image multi-tissue segmentation method and system based on heterogeneous UNet

    CN117876690A

  • Accurate ultrasonic eyebag removing apparatus

    CN202409873U