Joint point detection method and device based on N2S and bidirectional cross attention mechanism

By applying the joint node detection method based on the N2S and bidirectional cross attention mechanism in CT images, the problem of relying on labeled data and ignoring individual characteristics in the prior art is solved, and a higher accuracy and reliability joint node detection is achieved.

CN120182185APending Publication Date: 2025-06-20WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510189002.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

When the prior art is directly applied to CT maps, there are problems of relying on large-scale and high-quality labeling data and ignoring the individual characteristics of patients, resulting in inaccurate detection results of the joint nodes.

Method used

The joint detection method based on the N2S and bidirectional cross attention mechanism is adopted, and the image noise is removed through the N2S denoising network, and the multimodal feature fusion is performed by combining the patient's individual information, and the thermal map is generated by inputting the deconvolution network to finally obtain the joint detection result.

Benefits of technology

It significantly improves the accuracy and reliability of joint position detection, can more accurately reflect the anatomical characteristics and pathological changes of individual patients, reduces the work burden of doctors, and improves the accuracy of diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182185A_ABST
    Figure CN120182185A_ABST
Patent Text Reader

Abstract

The invention discloses a joint point detection method and device based on N2S and a bidirectional cross attention mechanism, a storage medium and electronic equipment. The method comprises the following steps: acquiring a CT image, and acquiring individual information of a patient; performing denoising processing on the CT image through an N2S denoising network to obtain a denoised image; performing multi-scale feature extraction on the de-noised image to obtain high-dimensional image features, and performing semantic feature extraction on the individual information of the patient to obtain text features; the high-dimensional image features and the text features are fused, and multi-modal fusion features are obtained; inputting the multi-modal fusion features into a deconvolution network to obtain a thermodynamic diagram; and obtaining a joint point detection result based on the thermodynamic diagram. The accuracy of the joint point detection result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of medicine and artificial intelligence, and particularly to a joint point detection method, device, storage medium and electronic device based on N2S and bidirectional cross-attention mechanism. Background Art

[0002] Accurately obtaining the positions of joint points from CT images has important practical significance. It can not only improve the diagnostic accuracy of medical images, but also provide scientific support for treatment and preoperative planning. In clinical diagnosis, accurately positioning joint points helps detect lesions such as fractures, arthritis, dislocations, etc., and analyze skeletal developmental abnormalities, providing a basis for early disease detection and intervention. In addition, in orthopedic surgery, accurately obtaining the positions of joint points can provide key references for surgical navigation and the correct placement of internal fixation devices, support the design of personalized surgical plans, and significantly improve the surgical success rate and treatment effect. For postoperative rehabilitation, accurate joint point data can also guide rehabilitation training, optimize movement patterns, and help patients recover normal activity ability as soon as possible. At the same time, this technology is also crucial for biomechanics research, and can provide core data for human motion analysis, mechanical distribution research, and artificial joint design. In addition, accurate joint point extraction can combine with AI models to achieve automated assisted diagnosis, reduce the workload of doctors, improve the efficiency of image analysis, and effectively reduce the risk of misdiagnosis and missed diagnosis, especially in complex cases.

[0003] Traditional CT images do not pre-label the positions of joint points, which means that physicians need to rely on their professional knowledge and experience to observe the anatomical structures in the CT images and infer and judge the positions of joint points. This process usually depends on the physician's understanding of bone morphology and tissue relationships and intuitive judgment of the lesion area, but this method may have certain subjectivity and errors. Especially in complex cases or when the anatomical structure is unclear, it is easy to lead to inaccurate diagnosis or missed diagnosis. For clinical applications that require precise position data (such as surgical planning, prosthesis implantation, motion analysis, etc.), this method may be insufficient. Therefore, developing an automated key point annotation method based on deep learning or image processing technology, which can accurately identify and label the positions of joint points on CT images, can not only reduce the workload of physicians, but also improve the accuracy of diagnosis and treatment in a more efficient and standardized way, providing strong technical support for personalized medicine and clinical assisted diagnosis.

[0004] Currently, human pose estimation and key point detection have been widely used in many fields, but there are two main problems in directly applying existing technologies to CT images: (1) Rely on large-scale and high-quality labeled data. (2) Ignore the influence of individual patient characteristics on the prediction results. Therefore, simply relying on image information may not comprehensively reflect the anatomical characteristics and pathological changes of individual patients, and the obtained joint point detection results are not accurate enough. Summary of the Invention

[0005] An embodiment of the present application provides a joint point detection method, device, storage medium and electronic device based on N2S and bidirectional cross-attention mechanism, which can improve the detection accuracy of joint point positions.

[0006] An embodiment of the present application provides a joint point detection method based on N2S and bidirectional cross-attention mechanism, including: Obtain a CT image and obtain individual patient information; Perform denoising processing on the CT image through an N2S denoising network to obtain a denoised image; Perform multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and perform semantic feature extraction on the individual patient information to obtain text features; Fuse the high-dimensional image features and the text features to obtain multi-modal fusion features; Input the multi-modal fusion features into a transposed convolution network to obtain a heat map; Based on the heat map, obtain the joint point detection result.

[0007] Further, in the above joint point detection method based on N2S and bidirectional cross-attention mechanism, wherein, the performing denoising processing on the CT image through an N2S denoising network to obtain a denoised image includes: Divide the CT image into several local regions, and perform masking processing on random partial pixel values in each local region; Input the several local regions subjected to masking processing into a trained N2S denoising network, and use the pixel values of the unmasked neighborhood to predict the masked pixel values to obtain a denoised image.

[0008] Further, in the above joint point detection method based on N2S and bidirectional cross-attention mechanism, wherein the training process of the N2S denoising network includes: Obtain an image data training set; Divide the images in the image data training set into several local regions, and perform masking processing on random partial pixel values in each local region; Input the local regions into the N2S denoising network, and predict the masked pixel values to obtain a predicted denoised image; Calculate the loss function based on the predicted pixel values and the true pixel values, and iteratively train the N2S denoising network through the loss function.

[0009] Further, in the above joint point detection method based on N2S and the bidirectional cross-attention mechanism, wherein, the multi-scale feature extraction of the denoised image to obtain high-dimensional image features is represented by the following formula:

[0010] Wherein, is the input denoised image, is the CNN convolutional neural network, is the high-dimensional image feature; The semantic feature extraction of the patient individual information to obtain text features is represented by the following formula:

[0011] Wherein, is the patient individual information, is the self-attention mechanism, is the text feature.

[0012] Further, in the above joint point detection method based on N2S and the bidirectional cross-attention mechanism, wherein, the fusion of the high-dimensional image feature and the text feature to obtain a multi-modal fusion feature includes: Taking the high-dimensional image feature as the query and the text feature as the key-value, calculating the first attention weight of the high-dimensional image feature to the text feature, and generating a first interaction feature; Taking the text feature as the query and the high-dimensional image feature as the key-value, calculating the second attention weight of the text feature to the high-dimensional image feature, and generating a second interaction feature; Based on the first attention weight and the second attention weight, fusing the first interaction feature and the second interaction feature to obtain a multi-modal fusion feature.

[0013] Further, in the above joint point detection method based on N2S and the bidirectional cross-attention mechanism, wherein, the input of the multi-modal fusion feature into the deconvolution network to obtain a heat map includes: Mapping the multi-modal fusion feature back to the spatial resolution through the deconvolution network and generating a heat map; wherein, the heat map includes multiple marked key points.

[0014] Further, in the above joint point detection method based on N2S and the bidirectional cross-attention mechanism, wherein, the obtaining of the joint point detection result based on the heat map includes: Use the SoftArgMax algorithm to perform weighted averaging on the heat map and calculate the two-dimensional coordinates of each key point; Take the key point coordinates as the coordinates of the joint points in the CT image.

[0015] An embodiment of the present application also provides a joint point detection device based on N2S and a bidirectional cross-attention mechanism, including: An acquisition module, configured to acquire a CT image and acquire patient individual information; A denoising module, configured to perform denoising processing on the CT image through an N2S denoising network to obtain a denoised image; A feature extraction module, configured to perform multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and perform semantic feature extraction on the patient individual information to obtain text features; A feature fusion module, configured to fuse the high-dimensional image features and the text features to obtain multi-modal fusion features; A heat map generation module, configured to input the multi-modal fusion features into a transposed convolution network to obtain a heat map; A joint point positioning module, configured to obtain a joint point detection result based on the heat map.

[0016] An embodiment of the present application also provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are suitable for being loaded by a processor to execute any one of the above joint point detection methods based on N2S and a bidirectional cross-attention mechanism.

[0017] An embodiment of the present application also provides an electronic device, including a processor and a memory, the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used for the steps in any one of the above joint point detection methods based on N2S and a bidirectional cross-attention mechanism.

[0018] The joint point detection method, device, storage medium and electronic device provided by the present application. The present application removes image noise (such as metal artifacts, scatter noise, etc.) generated by medical devices through the N2S network. Through the bidirectional cross-self-attention mechanism, combined with individual semantic information such as the patient's historical clinical features, the position of the joint points is predicted more accurately. In addition, the present invention also introduces SoftArgMax to overcome the non-differentiability and lack of robustness of the traditional ArgMax, and further improves the accuracy of joint point prediction. Description of the Drawings

[0019] The following will make the technical solutions and other beneficial effects of the present application obvious by describing the specific embodiments of the present application in detail in conjunction with the drawings.

[0020] Figure 1 This is a flowchart of the joint point detection method based on N2S and bidirectional cross-attention mechanism provided by the embodiments of this application.

[0021] Figure 2 This is a schematic structural diagram of the joint point detection device based on N2S and bidirectional cross-attention mechanism provided by the embodiments of this application.

[0022] Figure 3 This is a schematic structural diagram of the electronic device provided by the embodiments of this application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0024] Currently, human pose estimation and key point detection have been widely used in many fields. However, there are two main problems that need to be solved when directly applying existing technologies to CT images: (1) Dependence on large-scale and high-quality labeled data. Deep neural networks are essentially probabilistic statistical models, and their fitting effect on real-world data is affected by the scale and quality of sample data. Therefore, training such models requires a relatively high quality of data samples. Limited by factors such as medical equipment, the obtained data often contains more noise (such as metal artifacts, beam hardening artifacts, etc.), which affects feature calculation and parameter fitting during the entire forward and backward propagation processes, resulting in inaccurate final prediction results and affecting the judgment of physicians. (2) Ignoring the influence of individual patient characteristics on the prediction results. Individual characteristics include age, gender, weight, bone morphology, physiological state, and past medical history, etc. These information can significantly affect the anatomical location and manifestation form of key points. For example, there are significant differences in bone density and joint structure between children and the elderly, and the pelvic anatomical characteristics of women may be significantly different from those of men. And certain medical histories (such as fractures, arthritis, or osteoporosis) will cause key points to appear in abnormal positions or forms in CT images. Therefore, simply relying on image information may not be able to fully reflect the anatomical characteristics and pathological changes of individual patients. Combining individual characteristics to adjust or optimize the model can significantly improve the accuracy and robustness of key point detection, providing a more reliable basis for personalized diagnosis and treatment.

[0025] To solve the above problems, an embodiment of the present application provides a joint point detection method, device, storage medium, and electronic device based on N2S and a bidirectional cross-attention mechanism. An embodiment of the present application provides a joint point detection device based on N2S and a bidirectional cross-attention mechanism, which can be integrated in an electronic device. The electronic device can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0026] Please refer to Figure 1 , Figure 1 which is a flowchart of the joint point detection method based on N2S and a bidirectional cross-attention mechanism provided by an embodiment of the present application. It is applied to an electronic device. The joint point detection method based on N2S and a bidirectional cross-attention mechanism includes the following steps: S1, Obtain a CT image and obtain patient individual information.

[0027] Perform tomography on the human body through X-rays to obtain a CT (Computed Tomography) image of the internal structure of the body.

[0028] S2, Denoise the CT image through an N2S denoising network to obtain a denoised image.

[0029] In one embodiment, step S2 includes the following steps: S21, Divide the CT image into several local regions, and perform masking processing on random partial pixel values in each local region.

[0030] Specifically, divide the input CT image into several local regions , and perform masking processing on random partial pixel values therein to make the masked pixel values invisible.

[0031] S22, Input the several local regions subjected to masking processing into the trained N2S denoising network, and use the pixel values in the unmasked neighborhood to predict the masked pixel values to obtain a denoised image.

[0032] Specifically, use an N2S (Noise2Self) network model of unsupervised learning to denoise the input CT image. This method generates a high-quality denoised image by randomly masking some pixel values in the image and using neighborhood pixels to predict the masked pixels.

[0033] The training process of the N2S denoising network includes the following steps: A1, Obtain an image data training set.

[0034] Collect noisy images and use the collected noisy images as input to be provided to the model for processing.

[0035] A2, Divide the images in the image data training set into several local regions, and perform masking processing on random partial pixel values in each local region.

[0036] In the input noisy image, perform occlusion (through a neural network) processing on some pixel points to predict the values of these pixel points. The purpose of the occlusion operation is to simulate partial loss of the image and force the model to learn to infer the values of the occluded pixels from the surrounding information.

[0037] A3, Input the local regions into the N2S denoising network to predict the masked pixel values and obtain the predicted denoised image.

[0038] The model generates a denoised image with the same shape as the input image by predicting the values of the occluded pixel points.

[0039] A4, Calculate the loss function based on the predicted pixel values and the true pixel values, and iteratively train the N2S denoising network through the loss function.

[0040] Use the loss function (such as mean squared error) to calculate the error between the pixel values predicted by the model and the true values of the original noisy image. This step evaluates the prediction accuracy of the model.

[0041] Iterative optimization, according to the results of the loss function, use an optimization algorithm (such as SGD or Adam) to optimize the parameters of the neural network. The goal of optimization is to minimize the error between the predicted value and the true value, thereby gradually improving the denoising performance.

[0042] In an embodiment of the present invention, the medical CT images are denoised based on the Noise2Self (N2S) algorithm, improving the accuracy and reliability of key point detection. N2S is an unsupervised denoising method that directly learns the noise characteristics from noisy images to achieve efficient denoising without the need for clean images as supervision signals. The processing flow is as follows: First, the original CT images are preprocessed, including normalization and pixel intensity adjustment, to adapt to the input requirements of N2S. Then, using the core idea of N2S, a loss function is constructed by masking some pixel values as "pseudo-labels" to predict the values of the masked pixels with neighboring pixels. In this way, the model can learn the true signal patterns in the image rather than random noise patterns. Further, the entire image is gradually predicted through a sliding window operation, combined with multiple repeated samplings, to form an unbiased estimate and generate the denoised image. Finally, the denoised image is used for the key point detection task, significantly improving the detection accuracy of the model at key parts such as the lesion boundary and tissue structure by reducing noise interference. This method innovatively applies N2S to the unsupervised denoising of CT images, with the advantages of high efficiency, robustness, and unsupervised training, and can be widely applied to the field of medical image analysis and diagnostic assistance.

[0043] S3. Perform multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and perform semantic feature extraction on the patient individual information to obtain text features.

[0044] In one embodiment, step S3 includes the following steps: S31. Use a deep convolutional neural network (CNN) to perform multi-scale feature extraction on the denoised CT image to obtain high-dimensional image features containing spatial and semantic information. It can be represented by the following formula:

[0045] Where, is the input denoised image, is the CNN convolutional neural network, is the high-dimensional image feature.

[0046] S32. For patient individual information (such as age, gender, medical history, etc.), use a feature extraction model based on the self-attention mechanism to process the text information and generate a semantically rich and high-dimensional text feature representation. It can be represented by the following formula:

[0047] Where, is the patient individual information, is the self-attention mechanism, is the text feature.

[0048] Separate the extraction processes of the image stream and the information stream to ensure the independence of features and provide sufficient representation for subsequent fusion.

[0049] In one embodiment, a deep convolutional neural network model can be constructed with a two-stream network architecture to achieve multi-scale feature extraction from CT images and semantic feature extraction from patient individual information.

[0050] S4. Fuse the high-dimensional image features and text features to obtain multi-modal fusion features.

[0051] In one embodiment, step S4 includes the following steps: S41. Use the high-dimensional image features as queries and the text features as key values to calculate the first attention weights of the high-dimensional image features for the text features and generate the first interaction features.

[0052] S42. Use the text features as queries and the high-dimensional image features as key values to calculate the second attention weights of the text features for the high-dimensional image features and generate the second interaction features.

[0053] S43. Based on the first attention weights and the second attention weights, fuse the first interaction features and the second interaction features to obtain multi-modal fusion features.

[0054] Dynamically adjust the interaction weights of the two types of features through a bidirectional attention mechanism to ensure that the key regions and important semantic features receive higher attention weights.

[0055] The following explains step S4 through a specific embodiment: Multiply the collected CT image feature groups (combinations of high-dimensional image features) with W Q_img , W K_img , W V_img respectively to obtain the corresponding Q, K, V groups: {Q img1 , Q img2… Q imgn}, {K img1 , K img2… K imgn}, {V img1 , V img2… V imgn}.

[0056] Multiply the collected CT image feature groups with W Q_txt , W K_txt , W V_txt respectively to obtain the corresponding Q, K, V groups: {Q txt1 , Q txt2… Q txtn}, {K txt1 , K txt2… Ktxtn}, {V txt1 , V txt2… V txtn}.

[0057] Multiply each tensor in the Q group of image features with each tensor in the V group of text features and perform softmax normalization to obtain Attention_Score imgTotxt , multiply each tensor in the Q group of image features with each tensor in the V group of text features and perform softmax normalization to obtain Attention_Score txtToimg .

[0058] Multiply the obtained Attention_Score with V weighted respectively to obtain the image feature group F fused with semantic values imgMergetxt and the semantic feature group F fused with image values txtMergeimg .

[0059] In the embodiment of the present invention, based on the bidirectional cross-attention mechanism, the CT image features and the patient individual information features are fused for precise diagnosis and personalized treatment in medical image analysis. The processing process is as follows: Jointly model the CT image data and the patient individual information (such as age, gender, medical history, etc.), and realize the efficient fusion of multi-modal features through the bidirectional cross-attention mechanism. The processing flow is as follows: First, perform convolutional neural network (CNN) encoding on the CT image data to extract spatial features, and at the same time map the patient individual information to a high-dimensional feature space through a fully connected layer. Then, construct a bidirectional cross-attention module, where the image features and the individual features are used as inputs for each other's query (Query) and key-value (Key-Value), and the significance of feature interaction is dynamically adjusted through attention weights to achieve the full fusion of the two features. Specifically, the image features can enhance the lesion-related information in the individual features, while the individual features guide the image features to extract fine-grained regions matching the individual features. Finally, input the fused multi-modal features into a classification or regression model to complete the disease prediction or key point detection task. This method innovatively introduces the bidirectional cross-attention mechanism to realize the deep interaction between the CT image features and the individual information features, and improves the application value of multi-modal medical data in diagnosis and treatment decision-making.

[0060] S5. Input the multi-modal fusion features into a deconvolution network to obtain a heat map.

[0061] Specifically, map the multi-modal fusion features back to the spatial resolution through a deconvolution network and generate a heat map with the same size as the input image; wherein, the heat map includes multiple marked key points.

[0062] S6. Based on the heat map, obtain the joint point detection result In one embodiment, step S6 includes the following steps: S61, using the SoftArgMax algorithm to perform weighted averaging on the heatmap to calculate the two-dimensional coordinates of each key point; S62, taking the key point coordinates as the coordinates of the joint points in the CT image.

[0063] The joint point coordinates can be used for medical image analysis tasks such as lesion localization and organ structure annotation. The heatmap can be used to assist doctors in accurate diagnosis and support the application of various clinical scenarios, including automatically marking the center point of the lesion, detecting anatomical key points, and other precision medicine tasks.

[0064] According to the method described in the above embodiment, this embodiment will be further described from the perspective of a joint point detection device based on N2S and bidirectional cross-attention mechanism. The joint point detection device based on N2S and bidirectional cross-attention mechanism can be specifically implemented as an independent entity or integrated in an electronic device. The electronic device can be a terminal, a server, or other devices. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0065] Please refer to Figure 2 , Figure 2 Specifically describes the joint point detection device provided in the embodiment of the present application, which is applied to an electronic device. The joint point detection device based on N2S and bidirectional cross-attention mechanism may include: An acquisition module, configured to acquire a CT image and acquire patient individual information; A denoising module, configured to perform denoising processing on the CT image through an N2S denoising network to obtain a denoised image; A feature extraction module, configured to perform multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and perform semantic feature extraction on the patient individual information to obtain text features; A feature fusion module, configured to fuse the high-dimensional image features and the text features to obtain multi-modal fusion features; A heatmap generation module, configured to input the multi-modal fusion features into a transposed convolution network to obtain a heatmap; A joint point localization module, configured to obtain a joint point detection result based on the heatmap.

[0066] In specific implementation, each of the above modules and / or units can be implemented as an independent entity, or can be arbitrarily combined and implemented as the same or several entities. For the specific implementation of each of the above modules and / or units, reference can be made to the foregoing method embodiments. For the specific beneficial effects that can be achieved, please also refer to the beneficial effects in the foregoing method embodiments, which will not be elaborated herein.

[0067] In addition, an embodiment of the present application further provides an electronic device, which can be a device such as a computer or a tablet computer. The electronic device can implement the steps in any of the embodiments of the joint point detection method based on N2S and bidirectional cross-attention mechanism provided by the embodiments of the present application. Therefore, it can achieve the beneficial effects that can be achieved by any of the joint point detection methods based on N2S and bidirectional cross-attention mechanism provided by the embodiments of the present invention. For details, please refer to the foregoing embodiments, which will not be elaborated herein.

[0068] Figure 3 The specific structural block diagram of the electronic device provided by the embodiment of the present invention is shown. The electronic device can be used to implement the joint point detection method based on N2S and bidirectional cross-attention mechanism provided in the above embodiments. The electronic device 500 can be a device such as a terminal or a server. Among them, the terminal can include a tablet computer, a notebook computer, a personal computer (PC), a microprocessing box, or other devices, etc.

[0069] The RF circuit 510 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 510 may include various existing circuit elements for performing these functions. For example, antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 510 can communicate with various networks such as the Internet, intranets, wireless networks or communicate with other devices through wireless networks. The above-mentioned wireless networks may include cellular phone networks, wireless local area networks or metropolitan area networks. The above-mentioned wireless networks can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not been developed yet.

[0070] The memory 520 can be used to store software programs and modules, such as the corresponding program instructions / modules in the above embodiments. The processor 580 executes various functional applications and data processing by running the software programs and modules stored in the memory 520, that is, to implement functions such as taking pictures with the front camera, processing the captured images, and switching the display colors of the display content on the display screen. The memory 520 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 520 may further include a memory remotely disposed relative to the processor 580, and these remote memories can be connected to the electronic device 500 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0071] The input unit 530 can be used to receive input digital or character information, as well as generate a keyboard and a mouse related to user settings and function controls. The display unit 540 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces, and these graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 540 may include a display panel 541. Optionally, the display panel 541 can be configured in the form of an LCD (Liquid Crystal Display) or an OLED (Organic Light-Emitting Diode).

[0072] The audio circuit 560, the speaker 561, and the microphone 562 can provide an audio interface between the user and the electronic device 500. The audio circuit 560 can transmit the electrical signal converted from the received audio data to the speaker 561, and the speaker 561 converts it into a sound signal for output; on the other hand, the microphone 562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 560 and then converted into audio data. After the audio data is output to the processor 580 for processing, it is sent through the RF circuit 510 to, for example, another terminal, or the audio data is output to the memory 520 for further processing. The audio circuit 560 may also include an earphone jack to provide communication between the peripheral earphone and the electronic device 500.

[0073] The electronic device 500 can help the user receive requests, send information, etc. through the transmission module 570 (such as a Wi-Fi module), and it provides the user with wireless broadband Internet access. Although the transmission module 570 is shown in the figure, it can be understood that it does not belong to the essential composition of the electronic device 500 and can be omitted completely within the scope of not changing the essence of the invention according to needs.

[0074] The processor 580 is the control center of the electronic device 500, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 520, and by invoking the data stored in the memory 520, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 580 may include one or more processing cores; in some embodiments, the processor 580 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 580 either.

[0075] The electronic device 500 also includes a power supply 590 (such as a battery) for supplying power to each component. In some embodiments, the power supply can be logically connected to the processor 580 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 590 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.

[0076] Although not shown, the electronic device 500 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors. The one or more programs include instructions for performing the following operations: Obtain a CT image and obtain patient individual information; Perform denoising processing on the CT image through an N2S denoising network to obtain a denoised image; Perform multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and perform semantic feature extraction on the patient individual information to obtain text features; Fuse the high-dimensional image features and the text features to obtain multi-modal fusion features; Input the multi-modal fusion features into a deconvolution network to obtain a heat map; Based on the heat map, obtain a joint point detection result.

[0077] In specific implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned each module, reference can be made to the previous method embodiments, which will not be elaborated here.

[0078] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, an embodiment of the present invention provides a storage medium in which multiple instructions are stored. These instructions can be loaded by a processor to execute the steps of any of the embodiments of the joint point detection method based on N2S and bidirectional cross-attention mechanism provided by the embodiments of the present invention.

[0079] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0080] Since the instructions stored in this storage medium can execute the steps of any of the embodiments of the joint point detection method based on N2S and bidirectional cross-attention mechanism provided by the embodiments of the present invention, the beneficial effects achievable by any of the joint point detection methods based on N2S and bidirectional cross-attention mechanism provided by the embodiments of the present invention can be realized. For details, please refer to the previous embodiments and will not be elaborated here.

[0081] The above has introduced in detail a joint point detection method, device, storage medium, and electronic device based on N2S and bidirectional cross-attention mechanism provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A joint point detection method based on N2S and bidirectional cross attention mechanism, characterized in that: The method comprises: Acquire CT images and obtain individual patient information; Performing denoising on the CT image through an N2S denoising network to obtain a denoised image; Performing multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and performing semantic feature extraction on the individual patient information to obtain text features; Fusing the high-dimensional image features and the text features to obtain multimodal fusion features; Inputting the multimodal fusion features into a deconvolution network to obtain a heat map; Based on the heat map, a joint point detection result is obtained.

2. The joint point detection method based on N2S and bidirectional cross attention mechanism according to claim 1 is characterized in that: The denoising process is performed on the CT image by using the N2S denoising network to obtain a denoised image, including: Dividing the CT image into a plurality of local regions, and performing masking processing on a random portion of pixel values ​​in each local region; The masked local areas are input into the trained N2S denoising network, and the masked pixel values ​​are predicted using the unmasked pixel values ​​in the neighborhood to obtain the denoised image.

3. The joint point detection method based on N2S and bidirectional cross attention mechanism according to claim 2 is characterized in that: The training process of the N2S denoising network includes: Get the image data training set; Dividing the image in the image data training set into a plurality of local areas, and performing masking processing on a random portion of pixel values ​​in each local area; Inputting the local area into the N2S denoising network, predicting the pixel values ​​of the masked image, and obtaining a predicted denoised image; A loss function is calculated based on the predicted pixel values ​​and the actual pixel values, and the N2S denoising network is iteratively trained using the loss function.

4. The joint point detection method based on N2S and bidirectional cross attention mechanism according to claim 1 is characterized in that: The multi-scale feature extraction is performed on the denoised image to obtain high-dimensional image features, which are expressed by the following formula: in, is the input denoised image, is a convolutional neural network, is a high-dimensional image feature; The semantic feature extraction of the patient individual information is performed to obtain text features, which are expressed by the following formula: in, Individual patient information, is the self-attention mechanism, is a text feature.

5. The joint point detection method based on N2S and bidirectional cross attention mechanism according to claim 1 is characterized in that: The step of fusing the high-dimensional image features and the text features to obtain multimodal fusion features includes: Taking the high-dimensional image feature as a query and the text feature as a key value, calculating a first attention weight of the high-dimensional image feature to the text feature, and generating a first interaction feature; Taking the text feature as a query and the high-dimensional image feature as a key value, calculating a second attention weight of the text feature to the high-dimensional image feature, and generating a second interaction feature; Based on the first attention weight and the second attention weight, the first interaction feature and the second interaction feature are fused to obtain a multimodal fusion feature.

6. The joint point detection method based on N2S and bidirectional cross attention mechanism according to claim 1 is characterized in that: The multimodal fusion feature is input into the deconvolution network to obtain a heat map, including: The multimodal fusion features are mapped back to the spatial resolution through a deconvolution network, and a heat map is generated; wherein the heat map includes a plurality of marked key points.

7. The joint point detection method based on N2S and bidirectional cross attention mechanism according to claim 6 is characterized in that: The step of obtaining a joint point detection result based on the heat map includes: The SoftArgMax algorithm is used to perform weighted averaging on the heat map to calculate the two-dimensional coordinates of each key point; The key point coordinates are used as the coordinates of the joint points in the CT image.

8. A joint point detection device based on N2S and bidirectional cross attention mechanism, characterized in that: include: An acquisition module is used to acquire CT images and obtain individual patient information; A denoising module, used for denoising the CT image through an N2S denoising network to obtain a denoised image; A feature extraction module is used to perform multi-scale feature extraction on the denoised image to obtain high-dimensional image features, and to perform semantic feature extraction on the individual patient information to obtain text features; A feature fusion module, used for fusing the high-dimensional image features and the text features to obtain multimodal fusion features; A heat map generation module, used for inputting the multimodal fusion features into a deconvolution network to obtain a heat map; The joint point positioning module is used to obtain the joint point detection result based on the heat map.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the hyperspectral image classification method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: It comprises a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in the hyperspectral image classification method according to any one of claims 1 to 7.