Retina pulse signal decoding method and device, electronic equipment and storage medium

By training the initial pulse encoder and text generator and building a target pulse signal decoding model, the problem of insufficient accuracy in retinal pulse signal decoding is solved, accurate decoding of retinal pulse signals to text is achieved, and the parsing depth of visual information is improved.

CN120632474APending Publication Date: 2025-09-12SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510595404.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies have difficulty in accurately evaluating the visual information contained in retinal pulse signals, resulting in poor decoding accuracy.

Method used

By training the initial pulse encoder and the initial text generator, a target pulse signal decoding model is constructed. The target pulse encoder and the target text generator are combined to achieve accurate decoding of retinal pulse signals into text.

Benefits of technology

The accuracy of retinal pulse signal decoding is improved, enabling a deeper understanding of the visual information in retinal pulse signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632474A_ABST
    Figure CN120632474A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a retinal pulse signal decoding method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: encoding a sample retina pulse signal into an initial retina pulse feature through an initial pulse encoder; encoding the label text description into label text features; training an initial pulse encoder according to the similarity between the initial retina pulse feature and the label text feature to obtain a target pulse encoder; encoding the sample retina pulse signal into a target retina pulse feature through a target pulse encoder; generating sample text description for the text cue words and the target retina pulse features through an initial text generator; and training an initial text generator according to the sample text description and the label text description to obtain a target text generator so as to obtain a target pulse signal decoding model. According to the embodiment of the invention, the visual information in the retina pulse signal can be deeply and accurately understood.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a retinal pulse signal decoding method and device, electronic equipment, and storage medium. Background Art

[0002] Retinal spikes are electrical signals generated by the retina in response to visual stimulation, such as light. A retinal prosthesis is a medical electronic device that can assist the retina in generating retinal spikes to help visually impaired patients restore their visual perception. For example, for patients with retinal damage, a retinal prosthesis stimulates undamaged neurons in the retina through electrodes, activating the neurons to generate electrical signals, thereby helping the patient restore their visual perception. For the retinal spikes generated after stimulation by the retinal prosthesis, it is necessary to evaluate whether they contain rich visual information (such as light intensity, contrast, color, objects, etc.) in order to evaluate the performance of the retinal prosthesis.

[0003] However, it is currently difficult to evaluate the accuracy of the visual information contained in the retinal pulse signal. For example, it is currently common to rely on the patient's subjective recognition of the image after the retinal prosthesis is implanted (such as recognizing the presence of numbers, English letters, airplanes and other objects in the image) to evaluate the retinal pulse signal generated by the retinal prosthesis. However, this evaluation method is easily interfered with by human factors, and can only indirectly obtain the shallow visual information contained in the retinal pulse signal, but may not be able to recognize the light intensity, contrast, movement, etc. in the image, and the recognized visual information is not accurate enough. Current technology makes it difficult to extract the visual information contained in the retinal pulse signal, resulting in poor accuracy in signal decoding.

[0004] Therefore, how to improve the decoding accuracy of retinal pulse signals has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to propose a retinal pulse signal decoding method and device, electronic device, and storage medium, aiming to more deeply understand the visual information in the retinal pulse signal and improve the decoding accuracy.

[0006] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application provides a retinal pulse signal decoding method, the method comprising:

[0007] Acquire sample retinal pulse signals;

[0008] performing initial pulse signal encoding on the sample retinal pulse signal by a pre-built initial pulse encoder to obtain an initial retinal pulse feature;

[0009] Obtaining a label text description and performing text encoding on the label text description to obtain a label text feature; wherein the label text description indicates visual information contained in the sample retinal pulse signal, or indicates visual information contained in other pulse signals;

[0010] performing comparative learning training on the initial pulse encoder based on the similarity between the initial retinal pulse feature and the label text feature to obtain a target pulse encoder;

[0011] Performing target pulse signal encoding on the sample retinal pulse signal by the target pulse encoder to obtain a target retinal pulse feature;

[0012] Generating text from a preset text prompt word and the target retinal pulse feature using a pre-built initial text generator to obtain a sample text description; wherein the text prompt word indicates that the retinal pulse feature is described as text;

[0013] Performing model training on the initial text generator according to the sample text description and the label text description to obtain a target text generator;

[0014] The target text generator is connected after the target pulse encoder to obtain a target pulse signal decoding model, so as to decode the retinal pulse signal into text through the target pulse signal decoding model.

[0015] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a retinal pulse signal decoding device, the device comprising:

[0016] A pulse acquisition module, used for acquiring sample retinal pulse signals;

[0017] an initial pulse encoding module, configured to perform initial pulse signal encoding on the sample retinal pulse signal using a pre-built initial pulse encoder to obtain an initial retinal pulse feature;

[0018] a text encoding module, configured to obtain a label text description and perform text encoding on the label text description to obtain a label text feature; wherein the label text description indicates visual information contained in the sample retinal pulse signal, or indicates visual information contained in other pulse signals;

[0019] a contrastive learning training module, configured to perform contrastive learning training on the initial pulse encoder based on the similarity between the initial retinal pulse feature and the label text feature to obtain a target pulse encoder;

[0020] a target pulse encoding module, configured to perform target pulse signal encoding on the sample retinal pulse signal through the target pulse encoder to obtain a target retinal pulse feature;

[0021] a text generation module, configured to generate text from a preset text prompt word and the target retinal pulse feature using a pre-built initial text generator to obtain a sample text description; wherein the text prompt word indicates that the retinal pulse feature is described as text;

[0022] A text generation training module, configured to perform model training on the initial text generator based on the sample text description and the label text description to obtain a target text generator;

[0023] A model combination module is used to connect the target text generator after the target pulse encoder to obtain a target pulse signal decoding model, so as to decode the retinal pulse signal into text through the target pulse signal decoding model.

[0024] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0025] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0026] The retinal pulse signal decoding method and device, electronic device, and storage medium proposed in this application are designed to decode retinal pulse signals into text by separately training an initial pulse encoder and an initial text generator, and then using the target pulse encoder and target text generator obtained after training to obtain a target pulse signal decoding model. Specifically, in the process of training the initial pulse encoder, the sample retinal pulse signal is input into the initial pulse encoder and encoded into an initial retinal pulse feature. The initial pulse encoder is then trained for comparative learning based on the similarity between the initial retinal pulse feature and the label text feature corresponding to the label text description to obtain a target pulse encoder. This allows the target pulse encoder to learn the potential semantic association between the retinal pulse signal and the visual information contained in the label text description. For example, during contrastive learning training, if the label text description indicates the visual information contained in the sample retinal pulse signal, the parameters of the initial pulse encoder are adjusted to increase the similarity between the initial retinal pulse feature and the label text feature, so that the initial retinal pulse feature is close to the correct visual information; if the label text description indicates the visual information contained in other pulse signals, the parameters of the initial pulse encoder are adjusted to reduce the similarity between the initial retinal pulse feature and the label text feature, so that the initial retinal pulse feature is far away from the erroneous visual information. This allows the pulse feature encoded by the target pulse encoder (i.e., the target retinal pulse feature) to be more accurate. Then, in the process of training the initial text generator, the initial text generator generates a text description of the target retinal pulse feature (i.e., the sample text description) under the guidance of the text prompt word, and then the target text generator is trained based on the sample text description and the label text description. This allows the text description generated by the target text generator to reflect the true visual information of the retinal pulse signal as accurately as possible, so as to more accurately decode the retinal pulse signal into text containing visual information. In summary, the target pulse signal decoding model can accurately decode retinal pulse signals into text descriptions containing visual information, thereby more deeply analyzing the visual information contained in the retinal pulse signals and improving the accuracy of decoding retinal pulse signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flow chart of a retinal pulse signal decoding method provided in an embodiment of the present application;

[0028] Figure 2 yes Figure 1 Flowchart of step 102 in FIG.

[0029] Figure 3 This is provided in one embodiment of the present application Figure 1 Flowchart of step 103 in FIG.

[0030] Figure 4 Another embodiment of the present application provides Figure 1 Flowchart of step 103 in FIG.

[0031] Figure 5 yes Figure 1 Flowchart of step 104 in FIG.

[0032] Figure 6 yes Figure 5 Flowchart of step 502 in FIG.

[0033] Figure 7 yes Figure 5 Flowchart of step 503 in FIG.

[0034] Figure 8 This is a schematic diagram of the process of constructing a target pulse signal decoding model provided by an embodiment of the present application;

[0035] Figure 9 is a schematic structural diagram of a retinal pulse signal decoding device provided in an embodiment of the present application;

[0036] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0038] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0040] First, let’s analyze some of the terms used in this application:

[0041] Artificial Intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It can also refer to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning. This application can acquire and process relevant data based on AI technologies.

[0042] Multimodal Large Models (MM-LLMs): These are AI models capable of processing multiple data types, such as text, images, audio, and video. MM-LLMs improve comprehensive analysis capabilities for complex scenarios by learning relationships between data of different modalities (i.e., different types).

[0043] The retinal pulse signal decoding method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically described through the following embodiments. First, the retinal pulse signal decoding method in the embodiments of the present application is described.

[0044] The retinal pulse signal decoding method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network, content distribution network) and big data and artificial intelligence platforms; the software can be an application that implements the retinal pulse signal decoding method, etc., but is not limited to the above forms.

[0045] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0046] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0047] Figure 1 This is an optional flowchart of the retinal pulse signal decoding method provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps 101 to 108.

[0048] Step 101, obtaining a sample retinal pulse signal;

[0049] Step 102, performing initial pulse signal encoding on the sample retinal pulse signal by a pre-built initial pulse encoder to obtain an initial retinal pulse feature;

[0050] Step 103: Obtain a label text description and perform text encoding on the label text description to obtain a label text feature; wherein the label text description indicates visual information contained in the sample retinal pulse signal, or indicates visual information contained in other pulse signals;

[0051] Step 104: performing comparative learning training on the initial pulse encoder based on the similarity between the initial retinal pulse features and the label text features to obtain a target pulse encoder;

[0052] Step 105, encoding the sample retinal pulse signal into a target pulse signal by a target pulse encoder to obtain a target retinal pulse feature;

[0053] Step 106: generating text from the preset text prompt word and the target retinal pulse feature using a pre-built initial text generator to obtain a sample text description; wherein the text prompt word indicates that the retinal pulse feature is described as text;

[0054] Step 107: Perform model training on the initial text generator based on the sample text description and the label text description to obtain a target text generator;

[0055] Step 108: Connect a target text generator after the target pulse encoder to obtain a target pulse signal decoding model, so as to decode the retinal pulse signal into text through the target pulse signal decoding model.

[0056] The beneficial effects of the embodiments of the present application include but are not limited to: by separately training the initial pulse encoder and the initial text generator, and combining the target pulse encoder and the target text generator obtained after training to obtain a target pulse signal decoding model, the retinal pulse signal is decoded into text, thereby understanding the visual information contained in the retinal pulse signal. Specifically, in the process of training the initial pulse encoder, the sample retinal pulse signal is input into the initial pulse encoder and encoded into the initial retinal pulse feature. Then, based on the similarity between the initial retinal pulse feature and the label text feature corresponding to the label text description, the initial pulse encoder is subjected to comparative learning training to obtain a target pulse encoder. This allows the target pulse encoder to learn the potential semantic association between the retinal pulse signal and the visual information contained in the label text description. For example, during contrastive learning training, if the label text description indicates the visual information contained in the sample retinal pulse signal, the parameters of the initial pulse encoder are adjusted to increase the similarity between the initial retinal pulse feature and the label text feature, so that the initial retinal pulse feature is close to the correct visual information; if the label text description indicates the visual information contained in other pulse signals, the parameters of the initial pulse encoder are adjusted to reduce the similarity between the initial retinal pulse feature and the label text feature, so that the initial retinal pulse feature is far away from the erroneous visual information. This allows the pulse feature encoded by the target pulse encoder (i.e., the target retinal pulse feature) to be more accurate. Then, in the process of training the initial text generator, the initial text generator generates a text description of the target retinal pulse feature (i.e., the sample text description) under the guidance of the text prompt word, and then the target text generator is trained based on the sample text description and the label text description. This allows the text description generated by the target text generator to reflect the true visual information of the retinal pulse signal as accurately as possible, so as to more accurately decode the retinal pulse signal into text containing visual information. In summary, the target pulse signal decoding model can accurately decode retinal pulse signals into text descriptions containing visual information, thereby more deeply analyzing the visual information contained in the retinal pulse signals and improving the accuracy of decoding retinal pulse signals.

[0057] In step 101 of some embodiments, the sample retinal pulse signal refers to a retinal pulse signal used to train the initial pulse encoder. It should be noted that the retinal pulse signal is an electrical signal generated by the retina under visual stimulation (such as light stimulation, image stimulation), and the retinal pulse signal can also be an electrical signal generated by the retina under visual stimulation with the assistance of a retinal prosthesis. For example, a retinal pulse signal can be obtained from a preset retinal pulse database and used as a sample retinal pulse signal. For another example, for a patient with a retinal prosthesis implanted, an image can be used to stimulate the patient's retina so that the retina generates a retinal pulse signal with the assistance of the retinal prosthesis. In addition, the sample retinal pulse signal can also be obtained by other means, not limited to this.

[0058] In step 102 of some embodiments, the initial pulse encoder is a neural network model used to encode the retinal pulse signal. For example, the initial pulse encoder can be an encoder based on a multi-layer perceptron (MLP). The initial pulse encoder can extract efficient feature representations from the retinal pulse signal. It should be noted that the initial retinal pulse feature is the pulse feature encoded by the initial pulse encoder from the sample retinal pulse signal. The initial retinal pulse feature can specifically be a feature vector.

[0059] In step 103 of some embodiments, the label text description may be a text description of the visual information contained in the sample retinal pulse signal, or a text description of the visual information contained in pulse signals other than the sample retinal pulse signal (i.e., other pulse signals). Specifically, the label text description may be a sentence or a paragraph of text. For example, assuming that the sample retinal pulse signal is an electrical signal generated based on the stimulation of image A, and the visual information contained in image A includes: "A white cat sitting on the ground." If the label text description indicates the visual information contained in the sample retinal pulse signal, the label text description may be "A white cat sitting on the ground." If the label text description indicates the visual information contained in other pulse signals, the label text description is a sentence that is unrelated to or quite different from the visual information contained in image A.

[0060] In some embodiments, the label text feature is a text feature of the label text description. For example, the label text feature is defined as shown in the following formula:

[0061] H c =T θ (C), Formula (1);

[0062] Among them, H c Represents the label text feature; C represents the label text description; T θ (C) represents the encoding function used by the preset text encoder to perform text encoding on the label text description C.

[0063] In step 104 of some embodiments, the similarity between the initial retinal pulse feature and the label text feature may be the vector cosine similarity between the initial retinal pulse feature and the label text feature. The target pulse encoder refers to the trained initial pulse encoder.

[0064] In some embodiments, the specific process of contrastive learning training may include: calculating a contrastive loss value for the similarity between the initial retinal pulse features and the label text features using a contrastive learning function of the initial pulse encoder to obtain a contrastive loss value; and adjusting the parameters of the initial pulse encoder to minimize the contrastive loss value to obtain a target pulse encoder. In some embodiments, the contrastive loss value may be defined as follows:

[0065]

[0066] Among them, L contrastive Indicates the contrast loss value; H s represents the initial retinal pulse characteristics; H c Represents label text features; sim i,i (H s ,H c ) represents the i-th initial retinal pulse feature H s and the i-th label text feature H c The cosine similarity between i,j (H s ,H c ) represents the i-th spike eigenvector H s and the jth text feature vector H c where i∈[1,n] and j∈[1,n], N represents the total number of initial retinal pulse features and the total number of label text features; τ represents the temperature parameter of the initial pulse encoder; exp(·) represents the natural exponential function.

[0067] It should be noted that the temperature parameter is used to control the smoothness of the contrast loss function. The i-th label text feature is the positive label text description of the i-th initial retinal pulse feature. When i is not equal to j, the j-th text feature vector is the negative label text description of the i-th initial retinal pulse feature. In some embodiments, the similarity between the retinal pulse signal and the matching label text description (i.e., the positive label text description) can be made as high as possible by minimizing the contrast loss value, while the similarity between the retinal pulse signal and the unmatched label text description (i.e., the negative label text description) is made as low as possible. As for the meaning of the positive label text description and the negative label text description, please refer to the specific explanation of step 701 below, which will not be repeated here.

[0068] In step 105 of some embodiments, the target retinal pulse feature refers to a pulse feature obtained by encoding the sample retinal pulse signal through a target pulse encoder. The target retinal pulse feature may specifically be a feature vector.

[0069] In step 106 of some embodiments, the initial text generator is a neural network model for generating text. Specifically, the initial text generator can be a large language model (LLM), such as a large language model based on the Transfommer architecture, or a multimodal large language model. For example, the initial text generator can be an LLaMA model. It should be noted that the LLaMA model is a neural network model based on multi-layer self-attention and adopts the Transfommer architecture.

[0070] It should be noted that the text prompt is an instruction for describing the retinal pulse feature as text. For example, the text prompt may be "Describe the image corresponding to the retinal pulse feature in one sentence" or "Describe the visual information of the image corresponding to the retinal pulse feature in one sentence", etc.

[0071] It should be noted that the sample text description is text generated by the initial text generator to describe the visual information in the target retinal pulse feature. As in the previous example, assuming that the sample retinal pulse signal is an electrical signal generated based on the stimulation of image A, the sample text description could be "A cat sitting on the ground."

[0072] In some embodiments, the sample text description may be a phrase sequence consisting of at least two phrases. If the initial text generator is a large language model, the initial text generator may predict the probability of each phrase and then construct a text sequence from the phrases with the highest probability to generate the sample text description. For example, the process of predicting a text sequence is shown in the following formula:

[0073]

[0074] Among them, X S Indicates the target retinal pulse signal; X P Indicates text prompt words; Y * represents the text sequence output by the initial text generator, i.e., the sample text description; p(Y * |X S ,X P ) represents the target retinal pulse signal X S and text prompt word X P Generate text sequence Y under the condition of input data * probability; Indicates the target retinal pulse signal X S and text prompt word X P For input data, a text sequence has been generated Under the condition of i where i∈[1,n]; L represents the total number of phrases in the sample text description; and θ represents the trainable parameters of the initial text generator.

[0075] In step 107 of some embodiments, the target text generator refers to the trained initial text generator.

[0076] In some embodiments, step 107 may include: calculating the loss value of the sample text description and the label text description through a preset cross-entropy loss function to obtain a cross-entropy loss value; and adjusting the parameters of the initial text generator to minimize the cross-entropy loss value to obtain a target text generator.

[0077] In step 108 of some embodiments, the target pulse signal decoding model comprises a target pulse encoder and a target text generator. The target pulse signal decoding model is capable of decoding the retinal pulse signal into text. For example, assuming that the retinal pulse signal S0 is input into the target pulse signal decoding model, the target pulse encoder can encode the retinal pulse signal S0 into pulse features, and then the target text generator can convert the pulse features of the retinal pulse signal S0 into text, thereby obtaining text containing the visual information of the retinal pulse signal S0.

[0078] See also Figure 2 In some embodiments, the initial pulse encoder includes at least two hidden layers arranged in sequence, wherein the output data of the previous hidden layer is the input data of the next hidden layer;

[0079] Step 102 may include but is not limited to steps 201 to 206:

[0080] Step 201, determining the hidden layer of the first layer of the initial pulse encoder as the target layer, and determining the sample retinal pulse signal as the target input value;

[0081] Step 202: Perform a linear calculation based on the target input value, the weight matrix of the target layer, and the bias vector to obtain a linear value.

[0082] Step 203: If the target layer is not the last hidden layer, perform batch normalization on the linear value to obtain a normalized value;

[0083] Step 204: performing activation calculation on the normalized value using a preset activation function to obtain an activation value;

[0084] Step 205: determining the next hidden layer of the target layer as the target layer, determining the activation value as the target input value, and returning to the step of performing linear calculation based on the target input value, the weight matrix and the bias vector of the target layer to obtain a linear value;

[0085] In step 206 , if the target layer is the last hidden layer, the linear value is determined as the initial retinal pulse feature.

[0086] The advantage of this embodiment is that the features of the sample retinal pulse signal are extracted through each hidden layer of the initial pulse encoder. Specifically, for hidden layers other than the last layer, a linear calculation is performed on the target input value using a weight matrix and bias vector, followed by normalization and activation calculation. The resulting data is then input into the next hidden layer for more in-depth feature extraction. For the last hidden layer, a linear calculation is performed and the data is output to obtain the initial retinal pulse features. This allows the retinal pulse signal to be accurately encoded as features, thereby improving the accuracy of the encoding of the retinal pulse signal and providing a deeper and more accurate understanding of the visual information contained in the retinal pulse signal.

[0087] In step 201 of some embodiments, the initial pulse encoder may include an input layer, at least two hidden layers, and an output layer. Specifically, each hidden layer includes a linear transformation layer, a batch normalization layer, and a nonlinear activation layer. The hidden layer may also include other layers, such as a dropout layer, which is used to randomly close a portion of the nodes in the hidden layer during training to prevent model overfitting and improve generalization ability. It should be noted that the target layer can be any hidden layer. The target input value is the data used to input into the target layer. The target input value can be a sample retinal pulse signal or the output data of the previous hidden layer.

[0088] In step 202 of some embodiments, the product of the target input value and the weight matrix of the target layer may be added to the bias vector of the target layer to obtain a linear value.

[0089] In step 203 of some embodiments, the linear values ​​may be batch normalized using a batch normalization (BN) technique to increase the model training speed.

[0090] In step 204 of some embodiments, the activation function may be a ReLU function. Other types of activation functions may also be used, without limitation.

[0091] In some embodiments, the activation value can be obtained by the following formula:

[0092] h i =σ(BN(Wi h i-1 +b i )), formula (4);

[0093] Among them, h i Represents the output data of the i-th layer, that is, the activation value; h i-1 Represents the input data of the i-th layer, that is, the target input value; where i∈[1,n], n represents the total number of hidden layers; W i represents the weight matrix of the i-th layer; b i represents the bias vector of the i-th layer; σ(·) represents the activation function; BN(·) represents the batch normalization operation.

[0094] In some embodiments, it should be noted that W i h i-1 +b i Indicates linear value; BN(W i h i-1 +b i ) represents the normalized value.

[0095] In step 205 of some embodiments, the activation value is input into the next hidden layer for calculation. The calculation process can refer to the detailed description of steps 202 to 204 above and will not be repeated here.

[0096] In step 206 of some embodiments, the formula for the linear value output by the last hidden layer is the same as the definition of the linear value in the formula above and is not repeated here. It should be noted that the initial pulse encoder can capture the rich visual information contained in the retinal pulse signal by gradually reducing the dimension of the features of the retinal pulse signal using different hidden layers.

[0097] See also Figure 3 ,In some embodiments, the label text description includes a positive label text description, the positive label text description indicating visual information contained in the sample retinal pulse signal;

[0098] Step 103 may include but is not limited to steps 301 to 307:

[0099] Step 301, obtaining an image corresponding to a sample retinal pulse signal to obtain an original image;

[0100] Step 302: performing binary image segmentation on the original image to obtain a foreground area and a background area of ​​the original image;

[0101] Step 303: extract foreground image features from the foreground area to obtain foreground image features;

[0102] Step 304: extract background image features from the background area to obtain background image features;

[0103] Step 305, performing a first text conversion on the foreground image features to obtain a foreground text description;

[0104] Step 306, performing a second text conversion on the background image features to obtain a background text description;

[0105] Step 307: Perform text fusion on the foreground text description and the background text description to obtain a positive label text description.

[0106] The advantage of this embodiment is that, in the process of generating a positive label text description, the original image corresponding to the sample retinal pulse signal is segmented into a foreground area and a background area, and then the image features of the foreground area and the background area are extracted respectively, and the image features are independently converted into text. In this way, the visual information contained in the foreground of the image corresponding to the sample retinal pulse signal and the visual information contained in the background can be accurately extracted to avoid missing or failing to fully extract the visual information of the foreground or background, thereby reducing the possibility of loss of foreground details or weakening of background information due to global feature mixing during the feature extraction process. A positive label text description is then generated based on the foreground text description and the background text description, so that the positive label text description contains the visual information of the foreground text description and the background text description, thereby improving the completeness of the visual information contained in the positive label text description, and thereby improving the accuracy and reliability of the visual information extracted from the retinal pulse signal by the trained model.

[0107] It should be noted that if the label text description is used to describe the visual information contained in the sample retinal pulse signal, the label text description is a positive label text description.

[0108] In step 301 of some embodiments, the original image is used to stimulate the retina to generate the sample retinal pulse signal. In other words, the visual information contained in the sample retinal pulse signal is equivalent to the visual information contained in the original image. For example, if the original image contains visual information A as shown in the following sentence: "A white airplane flies in the blue sky," then the visual information contained in the sample retinal pulse signal is also visual information A.

[0109] In step 302 of some embodiments, the original image may be segmented into two categories using an image segmentation algorithm, such as an image threshold segmentation algorithm, an edge detection algorithm, a convolutional neural network algorithm, or the like.

[0110] In step 303 of some embodiments, the foreground image features are features in the foreground region. Specifically, the foreground image features may include the color (e.g., red, blue), light intensity, contrast, object type (e.g., airplane, car, bird), and other features of the foreground region.

[0111] In step 304 of some embodiments, the background image features are features in the background area. Specifically, the background image features may include features such as color, light intensity, contrast, and background type (such as sky, ground, or building) of the background area.

[0112] In step 305 of some embodiments, the foreground text description is a text description of the foreground image features. As in the above example, assuming the original image contains visual information A, and the foreground region of the original image includes an airplane, and the background region of the original image includes the sky, then the foreground text description may be Text Description 1: "A white airplane."

[0113] In step 306 of some embodiments, the background text description is a text description of the background image features. In the same example as above, the background text description can be text description 2: "blue sky".

[0114] In step 307 of some embodiments, a natural language processing model can be used to perform text fusion on the foreground text description and the background text description to generate a positive label text description, so that the positive label text description includes the semantic information of the foreground text description and the background text description. As in the above example, a sentence containing visual information A can be generated based on text descriptions 1 and 2 (see the detailed description of step 301).

[0115] See also Figure 4 ,In some embodiments, the label text description includes a positive label text description;

[0116] Step 103 may include but is not limited to steps 401 to 406:

[0117] Step 401, obtaining an image corresponding to a sample retinal pulse signal to obtain an original image;

[0118] Step 402: Obtain at least two texts for describing the original image to obtain at least two candidate text descriptions; wherein any two candidate text descriptions are different from each other;

[0119] Step 403: extract visual feature keywords from each candidate text description to obtain visual feature keywords;

[0120] Step 404: extract semantic feature keywords from each candidate text description to obtain semantic feature keywords;

[0121] Step 405 , performing keyword data fusion based on the visual feature keywords and semantic feature keywords of at least two candidate text descriptions to obtain image feature keywords;

[0122] Step 406 : generate sentences based on at least two image feature keywords to obtain a positive label text description.

[0123] The advantage of this embodiment is that it can obtain multiple text descriptions of the image corresponding to the sample retinal pulse signal (i.e., candidate positive text descriptions), then extract keywords used to describe the visual features and semantic features of the image, and generate sentences for the fused keywords to integrate the candidate positive text descriptions into a single text description (i.e., positive label text description). This allows the positive label text description to more fully and completely describe the visual information contained in the sample retinal pulse signal, thereby improving the accuracy and reliability of the trained model in extracting visual information from the retinal pulse signal.

[0124] In some embodiments, in step 401, a sample retinal pulse signal and an image corresponding to the signal (i.e., an original image) may be obtained from a retinal pulse-image database. The meaning and function of the original image can be found in the detailed description of step 301 above and will not be repeated here.

[0125] In step 402 of some embodiments, each candidate text description is a text used to describe the original image, and each candidate text description at least partially contains the visual information contained in the original image (i.e., the visual information of the sample retinal pulse signal). For example, assume that the visual information A contained in the original image is as shown in the following sentence: "A white plane is flying in the blue sky." The at least two candidate text descriptions include a candidate text description ST1 and a candidate text description ST2. Then, the candidate text description ST1 can be: "A white plane is flying", and the candidate text description ST2 can be: "The plane is flying in the blue sky."

[0126] In step 403 of some embodiments, the visual feature keywords refer to keywords used to describe the visual features of the image in the candidate text description. Specifically, the visual feature keywords can be keywords used to describe color (such as white, blue) and light intensity (such as bright, dim).

[0127] In step 404 of some embodiments, the semantic feature keywords refer to keywords used to describe the semantic features of the image in the candidate text description. Specifically, the semantic feature keywords can be keywords for object types (such as airplane, sky) and object actions (such as flying, sitting, standing).

[0128] In step 405 of some embodiments, the at least two image feature keywords may include a visual feature keyword of each candidate text description and a semantic feature keyword of each candidate text description.

[0129] In step 406 of some embodiments, a positive label text description can be generated for at least two image feature keywords through a text generation model, so that the positive label text description contains the visual information of each candidate positive text description, thereby improving the sufficiency and completeness of the visual information contained in the positive label text description.

[0130] See also Figure 5 In some embodiments, step 104 may include, but is not limited to, steps 501 to 504:

[0131] Step 501, performing similarity calculation based on the initial retinal pulse feature and the label text feature to obtain the initial pulse text similarity;

[0132] Step 502: determining a similarity weight based on the difference between the visual information of the image corresponding to the sample retinal pulse signal and the label text description;

[0133] Step 503, calculating the contrast loss value based on the product of the initial pulse text similarity and the similarity weight to obtain the contrast loss value;

[0134] Step 504 : Adjust the parameters of the initial pulse encoder according to the contrast loss value to obtain a target pulse encoder.

[0135] The advantage of this embodiment is that, considering that the label text feature may not fully describe the visual features of the initial retinal pulse feature, before calculating the contrast loss value, a similarity weight is first determined based on the visual information of the image corresponding to the sample retinal pulse signal and the difference between the label text description. This can determine whether the label text description completely and accurately describes the visual information in the image corresponding to the sample retinal pulse signal, thereby adjusting the weight of the initial pulse text similarity (i.e., the similarity weight). Then, the product of the initial pulse text similarity and the similarity weight is input into the contrast loss function of the initial pulse encoder to calculate the contrast loss value, thereby more accurately calculating the contrast loss value to adjust the parameters of the initial pulse encoder, thereby improving the accuracy of the target pulse encoder in extracting visual features from the retinal pulse signal.

[0136] In step 501 of some embodiments, the initial pulse text similarity may be the vector cosine similarity between the initial retinal pulse feature and the label text description. In another embodiment, the initial pulse text similarity may also be other types of similarity, not limited thereto.

[0137] In step 502 of some embodiments, the similarity weight is used to reflect the difference between the label text description (specifically, the positive label text description) and the actual visual information of the image corresponding to the sample retinal pulse signal. The definition and calculation process of the similarity weight can be referred to the detailed description of steps 601 to 604 below and will not be repeated here.

[0138] In step 503 of some embodiments, the product of the initial pulse text similarity and the similarity weight may be input into a preset contrast loss function to calculate a contrast loss value.

[0139] In step 504 of some embodiments, the target pulse encoder may be obtained by adjusting parameters of the initial pulse encoder to minimize the contrast loss value.

[0140] See also Figure 6 In some embodiments, step 502 may include, but is not limited to, steps 601 to 604:

[0141] Step 601, obtaining an image corresponding to a sample retinal pulse signal to obtain an original image;

[0142] Step 602: semantically describe the original image using a preset image description generation model to obtain a visual description text;

[0143] Step 603: Calculate the text semantic similarity between the visual description text and the label text description to obtain the text semantic similarity;

[0144] Step 604: normalize the text semantic similarity to obtain a similarity weight.

[0145] The advantage of this embodiment is that by describing the original image corresponding to the sample retinal pulse signal as text and then calculating the textual semantic similarity between the visual description text and the label text description, the textual semantic similarity can reflect whether the label text description completely and accurately describes the visual information in the image corresponding to the sample retinal pulse signal. The textual semantic similarity is then normalized into a similarity weight to adjust the parameters of the initial pulse encoder, thereby improving the accuracy of the target pulse encoder in extracting visual features from the retinal pulse signal.

[0146] In step 601 of some embodiments, the meaning and function of the original image can be referred to the detailed description of step 301 above, which will not be repeated here.

[0147] In step 602 of some embodiments, the image description generation model is a model for describing the visual information in the image as text. The visual description text is a description text of the visual information contained in the original image.

[0148] In step 603 of some embodiments, the visual description text may be encoded into visual text features, and the label text description may be encoded into label text features, and then cosine similarity may be calculated for the visual text features and the label text features to obtain text semantic similarity. In another embodiment, text semantic similarity may also be calculated in other ways, not limited thereto.

[0149] In step 604 of some embodiments, the value range of the similarity weight is [0, 1].

[0150] See also Figure 7 In some embodiments, the label text description includes a positive label text description and a negative label text description; wherein the positive label text description indicates the visual information contained in the sample retinal pulse signal, and the negative label text description indicates the visual information contained in other pulse signals;

[0151] Step 503 may include but is not limited to steps 701 to 707:

[0152] Step 701: determining the initial pulse text similarity corresponding to the positive label text description as the positive pulse text similarity, and determining the initial pulse text similarity corresponding to the negative label text description as the negative pulse text similarity;

[0153] Step 702: obtaining a positive similarity weighted value based on the product of the positive pulse text similarity and the similarity weight; wherein the similarity weight has a value between 0 and 1;

[0154] Step 703, performing a first exponential calculation on the ratio of the positive similarity weighted value to the temperature parameter of the initial pulse encoder to obtain a weighted index value of the positive sample pulse text;

[0155] Step 704: performing a second index calculation on the ratio of the negative pulse text similarity to the temperature parameter to obtain a negative sample pulse text index value;

[0156] Step 705, summing the weighted index value of the positive sample pulse text and the index value of the negative sample pulse text to obtain a target cumulative index value;

[0157] Step 706: performing a logarithmic calculation on the ratio of the weighted index value of the positive sample pulse text to the target accumulated index value to obtain a target logarithmic value;

[0158] Step 707: perform inverse calculation on the target logarithm value to obtain the contrast loss value.

[0159] The advantage of this embodiment is that after determining the similarity weight, the product of the positive pulse text similarity and the similarity weight is input into the contrast loss function of the initial pulse encoder, so that the contrast loss value can be affected by the completeness of the description of the visual information contained in the sample retinal pulse signal by the label text description, thereby more accurately calculating the contrast loss value to adjust the parameters of the initial pulse encoder, thereby improving the accuracy of the target pulse encoder in extracting visual features from the retinal pulse signal.

[0160] In step 701 of some embodiments, it should be noted that the positive label text description is text used to describe the visual information contained in the sample retinal pulse signal, and the negative label text description is text used to describe the visual information contained in other pulse signals other than the sample retinal pulse signal.

[0161] In step 702 of some embodiments, the positive similarity weighted value is the product of the positive pulse text similarity and the similarity weight. For example, in the definition formula of the contrast loss value below, k·sim i,i (H s ,H c ) indicates a positive similarity weight value.

[0162] In step 703 of some embodiments, for example, in the definition formula of the contrast loss value below, exp(k·sim i,i (H s ,H c ) / τ) represents the weighted index value of the positive sample pulse text.

[0163] In step 704 of some embodiments, for example, in the definition formula of the contrast loss value below, exp(sim i,j (H s ,H c ) / τ) represents the negative sample pulse text index value.

[0164] In step 705 of some embodiments, the target cumulative index value is the sum of the positive sample pulse text weighted index value and the negative sample pulse text index value.

[0165] In step 706 of some embodiments, the target logarithmic value is a value obtained by performing a logarithmic calculation on the ratio of the weighted index value of the positive sample pulse text to the target accumulated index value.

[0166] In step 707 of some embodiments, the contrast loss value is defined as follows:

[0167]

[0168] Among them, L contrastive Indicates the contrast loss value; H srepresents the initial retinal pulse characteristics; H c Represents the label text feature; k represents the similarity weight; sim i,i (H s ,H c ) represents the i-th initial retinal pulse feature H s and the i-th label text feature H c The cosine similarity between i,j (H s ,H c ) represents the i-th spike eigenvector H s and the jth text feature vector H c where i∈[1,n] and j∈[1,n], N represents the total number of initial retinal pulse features and the total number of label text features; τ represents the temperature parameter of the initial pulse encoder.

[0169] See also Figure 8 In some embodiments, the process of constructing a target pulse signal decoding model may include: obtaining a sample retinal pulse signal generated based on the original image and a text description corresponding to the sample retinal pulse signal (i.e., the label text description mentioned above). The sample retinal pulse signal corresponds to a preset text prompt word. The sample retinal pulse signal (Spike) is initially pulse-encoded by a trainable pulse encoder to obtain a retinal pulse feature. The pulse encoder is. The text description (Caption) is text-encoded by a pre-trained text encoder to obtain a text feature (i.e., the label text feature mentioned above). Based on the similarity between the retinal pulse feature and the text feature, the initial pulse encoder is trained by comparative learning to align the semantics of the retinal pulse feature and the label text feature. Then, the retinal pulse feature corresponding to the sample retinal pulse signal is encoded by the trained pulse encoder. The retinal pulse feature and the text prompt word (Prompt) are input into a trainable text generator (such as an LLaMA model) to train the text generator. Finally, the combination of the trained pulse encoder and the text generator is used as the target pulse signal decoding model.

[0170] In some embodiments, as Figure 8 As shown, the original image can be an image of an airplane flying in the sky. The sample retinal pulse signal can be an electrical signal generated by the retina after being stimulated by the image. Specifically, the Lora technology can be used to fine-tune the LLaMA model.

[0171] See also Figure 9 The present application also provides a retinal pulse signal decoding device that can implement the above-mentioned retinal pulse signal decoding method. The device includes:

[0172] The pulse acquisition module 801 is used to acquire sample retinal pulse signals;

[0173] An initial pulse encoding module 802 is configured to encode the sample retinal pulse signal using a pre-built initial pulse encoder to obtain an initial retinal pulse feature;

[0174] The text encoding module 803 is used to obtain a label text description and perform text encoding on the label text description to obtain a label text feature; wherein the label text description indicates the visual information contained in the sample retinal pulse signal, or indicates the visual information contained in other pulse signals;

[0175] A contrastive learning training module 804 is configured to perform contrastive learning training on the initial pulse encoder based on the similarity between the initial retinal pulse features and the label text features to obtain a target pulse encoder;

[0176] A target pulse encoding module 805 is configured to perform target pulse signal encoding on the sample retinal pulse signal through a target pulse encoder to obtain a target retinal pulse feature;

[0177] A text generation module 806 is configured to generate text from a preset text prompt word and a target retinal pulse feature using a pre-built initial text generator to obtain a sample text description; wherein the text prompt word indicates that the retinal pulse feature is described as text;

[0178] The text generation training module 807 is used to perform model training on the initial text generator based on the sample text description and the label text description to obtain a target text generator;

[0179] The model combination module 808 is used to connect the target text generator after the target pulse encoder to obtain a target pulse signal decoding model, so as to decode the retinal pulse signal into text through the target pulse signal decoding model.

[0180] The specific implementation of the retinal pulse signal decoding device is basically the same as the specific embodiment of the above-mentioned retinal pulse signal decoding method, and will not be repeated here.

[0181] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-described retinal pulse signal decoding method when executing the computer program. The electronic device may include any intelligent terminal such as a tablet computer or an in-vehicle computer.

[0182] See also Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0183] The processor 901 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0184] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the retinal pulse signal decoding method of the embodiments of this application.

[0185] Input / output interface 903, used to implement information input and output;

[0186] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0187] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0188] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0189] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which implements the above-mentioned retinal pulse signal decoding method when executed by a processor.

[0190] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0191] It should be noted that the non-Company's software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.

[0192] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0193] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0194] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0195] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0196] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0197] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0198] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0199] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0201] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0202] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A retinal pulse signal decoding method, characterized in that: The method comprises: Acquire sample retinal pulse signals; performing initial pulse signal encoding on the sample retinal pulse signal by a pre-built initial pulse encoder to obtain an initial retinal pulse feature; Obtaining a label text description and performing text encoding on the label text description to obtain a label text feature; wherein the label text description indicates visual information contained in the sample retinal pulse signal, or indicates visual information contained in other pulse signals; performing comparative learning training on the initial pulse encoder based on the similarity between the initial retinal pulse feature and the label text feature to obtain a target pulse encoder; Performing target pulse signal encoding on the sample retinal pulse signal by the target pulse encoder to obtain a target retinal pulse feature; Generating text from a preset text prompt word and the target retinal pulse feature using a pre-built initial text generator to obtain a sample text description; wherein the text prompt word indicates that the retinal pulse feature is described as text; Performing model training on the initial text generator according to the sample text description and the label text description to obtain a target text generator; The target text generator is connected after the target pulse encoder to obtain a target pulse signal decoding model, so as to decode the retinal pulse signal into text through the target pulse signal decoding model.

2. The method according to claim 1, characterized in that The step of performing comparative learning training on the initial pulse encoder based on the similarity between the initial retinal pulse feature and the label text feature to obtain a target pulse encoder includes: Calculating similarity based on the initial retinal pulse feature and the label text feature to obtain initial pulse text similarity; determining a similarity weight based on a degree of distinction between visual information of an image corresponding to the sample retinal pulse signal and the label text description; Calculating a contrast loss value based on the product of the initial pulse text similarity and the similarity weight to obtain a contrast loss value; The parameters of the initial pulse encoder are adjusted according to the contrast loss value to obtain the target pulse encoder.

3. The method according to claim 2, characterized in that The determining of the similarity weight according to the difference between the visual information of the image corresponding to the sample retinal pulse signal and the label text description includes: Acquiring an image corresponding to the sample retinal pulse signal to obtain an original image; Performing semantic description on the original image using a preset image description generation model to obtain a visual description text; Calculating text semantic similarity between the visual description text and the label text description to obtain text semantic similarity; The text semantic similarity is normalized to obtain the similarity weight.

4. The method according to claim 2, characterized in that The label text description includes a positive label text description and a negative label text description; wherein the positive label text description indicates the visual information contained in the sample retinal pulse signal, and the negative label text description indicates the visual information contained in other pulse signals; The calculating of the contrast loss value according to the product of the initial pulse text similarity and the similarity weight to obtain the contrast loss value includes: Determining the initial pulse text similarity corresponding to the positive label text description as a positive pulse text similarity, and determining the initial pulse text similarity corresponding to the negative label text description as a negative pulse text similarity; A positive similarity weighted value is obtained according to the product of the positive pulse text similarity and the similarity weight; wherein the value of the similarity weight is between 0 and 1; Performing a first exponential calculation on the ratio of the positive similarity weighted value to the temperature parameter of the initial pulse encoder to obtain a positive sample pulse text weighted index value; Performing a second exponential calculation on the ratio of the negative pulse text similarity to the temperature parameter to obtain a negative sample pulse text index value; Summing the positive sample pulse text weighted index value and the negative sample pulse text index value to obtain a target cumulative index value; Performing logarithmic calculation on the ratio of the positive sample pulse text weighted index value to the target cumulative index value to obtain a target logarithmic value; The target logarithm value is inverted to obtain the contrast loss value.

5. The method according to any one of claims 1 to 4, characterized in that The label text description includes a positive label text description, wherein the positive label text description indicates visual information contained in the sample retinal pulse signal; The step of obtaining the label text description includes: Acquiring an image corresponding to the sample retinal pulse signal to obtain an original image; Performing binary image segmentation on the original image to obtain a foreground area and a background area of ​​the original image; performing foreground image feature extraction on the foreground area to obtain foreground image features; Extracting background image features from the background area to obtain background image features; Performing a first text conversion on the foreground image feature to obtain a foreground text description; Performing a second text conversion on the background image features to obtain a background text description; Perform text fusion on the foreground text description and the background text description to obtain the positive label text description.

6. The method according to any one of claims 1 to 4, characterized in that The label text description includes a positive label text description; The step of obtaining the label text description includes: Acquiring an image corresponding to the sample retinal pulse signal to obtain an original image; Obtain at least two texts for describing the original image to obtain at least two candidate text descriptions; wherein any two candidate text descriptions are different from each other; Extracting visual feature keywords from each candidate text description to obtain visual feature keywords; Extracting semantic feature keywords from each candidate text description to obtain semantic feature keywords; Perform keyword data fusion based on the visual feature keywords and the semantic feature keywords of at least two candidate text descriptions to obtain image feature keywords; Sentence generation is performed based on at least two of the image feature keywords to obtain the positive label text description.

7. The method according to any one of claims 1 to 4, characterized in that The initial pulse encoder comprises at least two hidden layers arranged in sequence, wherein the output data of the hidden layer of the previous layer is the input data of the hidden layer of the next layer; The step of encoding the sample retinal pulse signal with an initial pulse encoder to obtain an initial retinal pulse feature includes: determining the hidden layer of the first layer of the initial pulse encoder as a target layer, and determining the sample retinal pulse signal as a target input value; Performing a linear calculation based on the target input value, the weight matrix and the bias vector of the target layer to obtain a linear value; If the target layer is not the hidden layer of the last layer, performing batch normalization processing on the linear value to obtain a normalized value; Performing activation calculation on the normalized value using a preset activation function to obtain an activation value; Determining the next hidden layer of the target layer as the target layer, determining the activation value as the target input value, and returning to the step of performing linear calculation based on the target input value and the weight matrix and bias vector of the target layer to obtain a linear value; If the target layer is the last hidden layer, the linear value is determined as the initial retinal pulse feature.

8. A retinal pulse signal decoding device, characterized in that: The device comprises: A pulse acquisition module, used for acquiring sample retinal pulse signals; an initial pulse encoding module, configured to perform initial pulse signal encoding on the sample retinal pulse signal using a pre-built initial pulse encoder to obtain an initial retinal pulse feature; a text encoding module, configured to obtain a label text description and perform text encoding on the label text description to obtain a label text feature; wherein the label text description indicates visual information contained in the sample retinal pulse signal, or indicates visual information contained in other pulse signals; a contrastive learning training module, configured to perform contrastive learning training on the initial pulse encoder based on the similarity between the initial retinal pulse feature and the label text feature to obtain a target pulse encoder; a target pulse encoding module, configured to perform target pulse signal encoding on the sample retinal pulse signal through the target pulse encoder to obtain a target retinal pulse feature; a text generation module, configured to generate text from a preset text prompt word and the target retinal pulse feature using a pre-built initial text generator to obtain a sample text description; wherein the text prompt word indicates that the retinal pulse feature is described as text; A text generation training module, configured to perform model training on the initial text generator based on the sample text description and the label text description to obtain a target text generator; A model combination module is used to connect the target text generator after the target pulse encoder to obtain a target pulse signal decoding model, so as to decode the retinal pulse signal into text through the target pulse signal decoding model.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the retinal pulse signal decoding method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the retinal pulse signal decoding method according to any one of claims 1 to 7 is implemented.