A data processing system for AR head-mounted displays

CN122569722APending Publication Date: 2026-08-14GUANGZHOU TCL IND RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,现有的AR设备存在翻译结果不准确的问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122569722A_ABST
    Figure CN122569722A_ABST
Patent Text Reader

Abstract

This application discloses a data processing system for AR head-mounted display devices, which extracts features from the data to be processed to obtain target feature information; and determines target text information based on the target feature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically to a data processing system. Background Technology

[0002] With the advancement of globalization and the development of information technology, cross-cultural communication is becoming increasingly frequent. Augmented Reality (AR) devices, by translating user-generated audio data into text in the target language, can provide real-time language translation and are therefore widely used. However, existing AR devices suffer from inaccurate translation results. Summary of the Invention

[0003] This application provides a data processing system.

[0004] In a first aspect, this application provides a method comprising:

[0005] Obtain the data to be processed;

[0006] Feature extraction is performed on the data to be processed to obtain target feature information;

[0007] Based on target feature information, target text information is determined.

[0008] Secondly, this application provides a system comprising:

[0009] The information acquisition module is used to acquire data information to be processed;

[0010] The feature extraction module is used to extract features from the data to be processed to obtain target feature information;

[0011] The data determination module is used to determine the target text information based on the target feature information.

[0012] Thirdly, this application also provides an apparatus, the apparatus comprising:

[0013] One or more processors;

[0014] Memory; and

[0015] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the methods of any one of the first aspects.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the method in any of the first aspects. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of a data processing system provided in an embodiment of the present invention;

[0019] Figure 2 This is a flowchart of one embodiment of the data processing method provided by the present invention;

[0020] Figure 3 This is a schematic diagram of the specific structure of the third processing model provided in the embodiments of the present invention;

[0021] Figure 4 This is a schematic diagram of the specific structure of the seventh feature extraction unit provided in an embodiment of the present invention;

[0022] Figure 5 This is a flowchart of a specific embodiment of feature extraction of data information to be processed provided by the present invention;

[0023] Figure 6 This is a schematic diagram of the specific structure of the first processing model provided in the embodiment of the present invention;

[0024] Figure 7 This is a schematic diagram of the specific structure of the second processing model provided in the embodiments of the present invention;

[0025] Figure 8 This is a flowchart illustrating a specific embodiment of determining target text information provided in this invention.

[0026] Figure 9 This is a schematic block diagram of the data processing system provided in the embodiments of the present invention;

[0027] Figure 10 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features. In the description of this application, "several" means one or more, unless otherwise explicitly specified.

[0030] This application provides a data processing method and system. Furthermore, the above-mentioned data processing method and system can be used for AR head-mounted display devices, which will be described in detail below.

[0031] Please see Figure 1 , Figure 1 This is a schematic diagram of a data processing system provided in an embodiment of this application. The data processing system may include a computer device 100, which integrates the data processing system, such as... Figure 1 Computer equipment in the country.

[0032] In this embodiment, the computer device 100 can be a standalone server, a server network, or a server cluster. For example, the computer device 100 described in this embodiment includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0033] It is understood that the computer device 100 used in the embodiments of this application can be a device that includes both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device may include: cellular or other communication devices having a single-line display, a multi-line display, or a cellular or other communication device without a multi-line display. Specifically, the computer device 100 may be a desktop terminal or a mobile terminal, and the computer device 100 may also be one of the following: Augmented Reality (AR) based wearable devices (e.g., AR glasses, AR helmets, AR goggles, etc.), mobile phones, tablets, laptops, etc.

[0034] Those skilled in the art will understand that Figure 1 The application environment shown is merely one application scenario of the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include those that are more specific to this application. Figure 1 The number of computer devices shown is more or less, for example Figure 1Only one computer device is shown in the diagram. It is understood that the data processing system may also include one or more other services, which are not limited here.

[0035] In addition, such as Figure 1 As shown, the data processing system may also include a memory 200 for storing data, such as data to be processed, such as first image data information, first audio data information, etc., and feature information, such as first feature information, second feature information, etc.

[0036] It should be noted that, Figure 1 The schematic diagram of the data processing system shown is merely an example. The data processing system and scenario described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of data processing systems and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0037] like Figure 2 The diagram shown is a flowchart of an embodiment of the data processing method in this application. The data processing method may include the following steps S201 to S203, as detailed below:

[0038] S201. Obtain the data information to be processed.

[0039] In this embodiment, the data information to be processed is data information acquired by a computer device that needs to be translated. The data information to be processed can be unimodal data information, for example, it may include first image data information or first audio data information; it can also be multimodal data information, for example, it may include first image data information and first audio data information. This embodiment does not impose any limitations. Unimodal data information refers to data of only one type, such as text, image, audio, video, electromagnetic signals, etc.; multimodal data information refers to data including at least two types of unimodal information, such as at least two types of data including text, image, audio, video, electromagnetic signals, etc.

[0040] Optionally, the first image data information may be image data information obtained through the imaging module configured on the computer device itself, or image data information obtained from other computer devices through means such as network, Bluetooth, and infrared. The first image data information may also be image data information obtained by segmenting the second image data information. This embodiment does not limit the scope of the first image data information.

[0041] Furthermore, the first audio data information can be audio data information collected by the audio sensor configured on the computer device itself, or audio data information obtained from other computer devices through means such as network, Bluetooth, and infrared. The first audio data information can also be audio data information obtained by processing the second audio data information through the third processing model. This embodiment does not limit the scope of the first audio data information.

[0042] For example, the data processing method of this application is applied to wearable devices based on Augmented Reality (AR) (e.g., AR glasses, AR helmets, AR goggles, etc.). The AR device can acquire second image data information and second audio data information through its own configured imaging module and audio sensor, respectively. Then, the second image data information is segmented to obtain first image data information, and the second audio data information is processed by a third processing model to obtain first audio data information.

[0043] In some embodiments, the first image data information is determined based on the following method: acquiring second image data information; segmenting the second image data information to obtain the first image data information. The second image data information is the unprocessed raw image data information. The second image data information can be image data information acquired by an imaging module configured within the computer device itself, image data information acquired by a high-definition camera, or image data information acquired from other computer devices via network, Bluetooth, infrared, etc. This embodiment does not impose any limitations on these methods.

[0044] In this embodiment, segmentation refers to the technique and process of decomposing an image into several regions with unique properties and extracting the target of interest. For example, if the second image data is 3840×1080 pixels, it can be segmented into 1800 image blocks of size 64×36 to obtain the first image data. This embodiment improves the efficiency of data processing and the accuracy of the obtained target text information by segmenting the second image data and then processing the data based on the segmented first image data.

[0045] In some embodiments, the first audio data information is obtained by processing the second audio data information using a third processing model. In this embodiment, processing the second audio data information using the third processing model can remove noise from the second audio data information, improving the accuracy of the obtained target text information. Optionally, the second audio data information is unprocessed raw audio data information. The second audio data information can be audio data information collected by an audio sensor configured on the computer device itself, or audio data information obtained from other computer devices via networks, Bluetooth, infrared, etc. This embodiment does not impose any limitations.

[0046] In some embodiments, refer to Figure 3 As shown, the third processing model includes: a fourth feature extraction module, a fifth feature extraction module, a first feature fusion module, and a decoding module. The input of the fourth feature extraction module is configured to receive second audio data information. The output of the fourth feature extraction module is connected to the input of the fifth feature extraction module. The outputs of the fourth and fifth feature extraction modules are connected to the input of the first feature fusion module, and the output of the first feature fusion module is connected to the input of the decoding module.

[0047] Furthermore, the fourth feature extraction module is configured to extract features from the second audio data information, the fifth feature extraction module is configured to extract features from the data representation output by the fourth feature extraction module, the first feature fusion module is configured to fuse the data representation output by the fourth feature extraction module and the data representation output by the fifth feature extraction module, and the decoding module is configured to decode the data representation output by the first feature fusion module.

[0048] In some embodiments, continue to refer to Figure 3 As shown, the fourth feature extraction module includes a first encoding unit, a seventh feature extraction unit, and a third normalization unit. The input of the first encoding unit is configured to receive second audio data information, and the output of the first encoding unit is connected to the input of the seventh feature extraction unit. The output of the seventh feature extraction unit is connected to the input of the third normalization unit. The first encoding unit is configured to encode the second audio data information; the seventh feature extraction unit is configured to extract features from the data representation output by the first encoding unit; and the third normalization unit is configured to normalize the data representation output by the seventh feature extraction unit.

[0049] In this embodiment, encoding the second audio data information refers to converting the second audio data information into a data representation that can be used by the seventh feature extraction unit for feature extraction. The encoding process of the first encoding unit on the second audio data information can be expressed as: y(t)=H(z(t)U T In this context, y(t) represents the data representation output by the first coding unit, H(·) represents the nonlinear activation function, U represents the encoder matrix, and z(t) represents the second audio data information. Among them, s i (t) represents the noise signal to be removed, M represents the number of noise sources, and x(t) represents the first audio data information. Optionally, the first encoding unit can be built based on a Transformer encoder, an LSTM network, or a GRU network; the embodiments of the present invention are not limited thereto.

[0050] In some embodiments, refer to Figure 4 As shown, the seventh feature extraction unit includes a second convolutional unit and a third convolutional unit. The output of the first encoding unit is connected to the input of the second convolutional unit, and the output of the second convolutional unit is connected to the input of the third convolutional unit.

[0051] Optionally, continue to refer to Figure 4 As shown, the second convolutional unit includes a second convolutional layer, a third normalization layer, and a first activation layer. The output of the first encoding unit is connected to the input of the second convolutional layer, the output of the second convolutional layer is connected to the input of the third normalization layer, and the output of the third normalization layer is connected to the input of the first activation layer.

[0052] Furthermore, continue to refer to Figure 4 As shown, the third convolutional unit includes a third convolutional layer, a fourth normalization layer, and a second activation layer. The output of the second convolutional unit is connected to the input of the third convolutional layer, the output of the third convolutional layer is connected to the input of the fourth normalization layer, and the output of the fourth normalization layer is connected to the input of the second activation layer.

[0053] In one specific embodiment, the second convolutional layer is constructed based on a depthwise separable convolutional layer (DWConv), and the third convolutional layer is constructed based on a regular convolutional layer. The size of the convolutional kernels used in the second and third convolutional layers can be set according to actual needs. For example, the second convolutional layer can use a 3×3 convolutional kernel, and the third convolutional layer can use a 1×1 convolutional kernel, or the second convolutional layer can use a 5×5 convolutional kernel, and the third convolutional layer can use a 3×3 convolutional kernel, etc. This embodiment does not impose any limitations. In this application embodiment, the seventh feature extraction unit extracts features from the data representation output by the first encoding unit, which can transform the data representation output by the first encoding unit into a higher-dimensional data representation, thereby improving the accuracy of the determined first audio data information.

[0054] Optionally, the third normalization unit, the third normalization layer, and the fourth normalization layer can be constructed based on a batch normalization (BN) layer, a layer normalization (LN) layer, or a spectral normalization (SN) layer; this embodiment does not impose any limitations. In one specific embodiment, the third normalization unit is constructed based on an LN layer, and the third and fourth normalization layers are constructed based on a BN layer.

[0055] Furthermore, the first and second activation layers can be constructed based on the Sigmoid activation function, the ReLU activation function, or the Softmax activation function; this embodiment does not impose any limitations. In some embodiments, the first and second activation layers are constructed based on the ReLU activation function.

[0056] In some embodiments, continue to refer to Figure 3 As shown, the fifth feature extraction module includes an eighth feature extraction unit and a first activation unit. The input of the eighth feature extraction unit is configured to receive the data representation output from the fourth feature extraction module, and the output of the eighth feature extraction unit is connected to the input of the first activation unit.

[0057] Optionally, the fifth feature extraction module may include an eighth feature extraction unit, or the fifth feature extraction module may include multiple eighth feature extraction units, for example, continuing to refer to Figure 3As shown, the fifth feature extraction module includes R eighth feature extraction units, where R is a positive integer. In a specific embodiment, the fifth feature extraction module includes 6 eighth feature extraction units.

[0058] Furthermore, continue to refer to Figure 3 As shown, the eighth feature extraction unit includes at least one first convolutional unit, and each first convolutional unit includes: a first convolutional layer, a first normalization layer, and a second normalization layer. The output of the first convolutional layer is connected to the input of the first normalization layer, and the output of the first normalization layer is connected to the input of the second normalization layer.

[0059] Optionally, the first convolutional layer can be constructed based on a 1×1 convolutional kernel, a 3×3 convolutional kernel, or a 5×5 convolutional kernel; this embodiment does not impose any limitation. The first and second normalization layers can be constructed based on batch normalization (BN) layers, layer normalization (LN) layers, or spectral normalization (SN) layers; this embodiment does not impose any limitation. In some embodiments, the first normalization layer is constructed based on an LN layer, and the second normalization layer is constructed based on an SN layer.

[0060] In some embodiments, continue to refer to Figure 3 As shown, the first feature fusion module includes a fourth fusion unit and a ninth feature extraction unit. The outputs of the fourth and fifth feature extraction modules are connected to the input of the fourth fusion unit, and the output of the fourth fusion unit is connected to the input of the ninth feature extraction unit. The fourth fusion unit is configured to fuse the data representations output by the fourth and fifth feature extraction modules, and the ninth feature extraction unit is configured to extract features from the data representation output by the fourth fusion unit.

[0061] Optionally, the process of the fourth fusion unit fusing the data representation output by the fourth feature extraction module and the data representation output by the fifth feature extraction module can be expressed as: d(t) = w(t) ⊙ m(t), where d(t) represents the data representation output by the fourth fusion unit, m(t) represents the data representation output by the fourth feature extraction module, and w(t) represents the data representation output by the fifth feature extraction module.

[0062] Optionally, the structure of the ninth feature extraction unit is the same as that of the seventh feature extraction unit. For details, please refer to the description of the structure of the seventh feature extraction unit described above. To avoid repetition, this embodiment will not repeat the description here. As mentioned in the preceding steps, the seventh feature extraction unit can transform the data representation output by the first encoding unit into a higher-dimensional data representation. Therefore, the data representation output by the fourth fusion unit is also a high-dimensional data representation. By extracting features from the data representation output by the fourth fusion unit through the ninth feature extraction unit, the high-dimensional data representation output by the fourth fusion unit can be mapped back to the original-dimensional data representation.

[0063] In this embodiment of the application, decoding processing refers to the process of converting the data representation output by the first feature fusion module into first audio data information. The decoding processing process can be represented as follows: in, Let p(t) represent the first audio data information, p(t) represent the data representation output by the first feature fusion module, and V represent the decoding module. Optionally, the decoding module can be built based on a Transformer decoder, an LSTM network, or a GRU network; this application embodiment does not limit the specific implementation.

[0064] In some embodiments, the third processing model is obtained by training the first network model using the first training sample set. The first network model and the third processing model have the same model structure. The difference between the first network model and the third processing model is that the model parameters of the first network model are the initial model parameters, while the model parameters of the third processing model are the trained model parameters.

[0065] Optionally, the first training sample set is obtained by labeling noise signals in multiple first sample audio data. The first training sample set can be obtained by manually labeling noise signals in the first sample audio data, automatically labeling noise signals in the first sample audio data, or semi-automatically labeling noise signals in the first sample audio data. This embodiment does not limit the method.

[0066] In some embodiments, the training process of the third processing model is as follows: A first training sample set is input into a first network model, and the first network model outputs first probability information corresponding to the noise signal in each first sample audio data. Based on the first probability information, a first loss value is determined. If the first loss value does not meet a first condition, the model parameters of the first network model are adjusted based on a first learning rate, and the steps of inputting the first training sample set into the first network model and outputting first probability information corresponding to the noise signal in each first sample audio data are continued until the first loss value meets the first condition or the number of updates of the first network model reaches a first threshold. The updated first network model is then determined as the third processing model. The first loss value meeting the first condition can be either the first loss value being less than a first threshold or the difference between two consecutive first loss values ​​being less than a second threshold.

[0067] Optionally, the process of determining the first loss value can be expressed as: Where θ represents the set of model parameters for the first network model, S represents the number of first sample audio data, and m i Let m represent the mask vector corresponding to the i-th first sample audio data. i The value at the location corresponding to the noise signal is 0, and the value at the location corresponding to other signals is 1, p(m=m i |θ) represents the first probability information corresponding to the noise signal in the i-th first sample audio data.

[0068] Furthermore, the initial learning rate threshold can range from 900 to 1100, and optionally, it can be 1000. During model training, a dynamic learning rate adjustment strategy can be used to adjust the initial learning rate. For example, the initial learning rate can be set to 0.001, and every 20 model updates, the initial learning rate is adjusted exponentially at a decay rate of 0.9. This can improve the model's generalization ability, enabling it to converge faster in the early stages of training and avoiding getting stuck in local minima in the later stages.

[0069] Optionally, the first network model can be trained on a computer device with powerful graphics processing unit (GPU) computing power, high system memory (e.g., 32GB or more of system content), and large-capacity high-speed hard disk or solid-state disk (SSD) (e.g., at least 1TB SSD). The GPU can provide efficient parallel computing capabilities, accelerate complex operations during model training, and significantly shorten the training time of the model. The high system memory can ensure that there is no memory shortage during data loading, model parameter storage, and intermediate result calculation. The large-capacity high-speed hard disk or solid-state disk can guarantee data read and write speed and reduce the impact of data loading time on training efficiency.

[0070] S202. Extract features from the data to be processed to obtain target feature information.

[0071] In this embodiment, the target feature information is a data representation extracted from the data information to be processed. The data information to be processed includes first image data information and / or first audio data information. The target feature information includes first feature information and / or second feature information. Specifically, the first feature information is a data representation extracted from the first image data information, and the second feature information is a data representation extracted from the first audio data information.

[0072] In some embodiments, the data information to be processed includes first image data information and / or first audio data information, and the target feature information includes first feature information and / or second feature information. The steps of extracting features from the data information to be processed to obtain target feature information specifically include: extracting features from the first image data information to obtain first feature information; and / or extracting features from the first audio data information to obtain second feature information.

[0073] Optionally, the data to be processed includes first image data and first audio data, and the target feature information includes first feature information and second feature information, as referenced. Figure 5 As shown, step S202 above, which involves extracting features from the data to be processed to obtain target feature information, may include steps S301 to S302, as follows:

[0074] S301. Extract features from the first image data to obtain first feature information.

[0075] In some embodiments, the first feature information is obtained by feature extraction from the first image data information based on a first processing model, referring to... Figure 6As shown, the first processing model includes a first encoding module and several first feature extraction modules connected in series. The input of the first encoding module is configured to receive first image data information, and the output of the first encoding module is connected to the inputs of the several first feature extraction modules. The first encoding module is configured to encode the first image data information, and the several first feature extraction modules are configured to extract features from the data representation output by the first encoding module to obtain first feature information.

[0076] In this embodiment, encoding the first image data information refers to converting the first image data information into a high-dimensional vector representation, thereby transforming the first image data information into a form suitable for processing by several first feature extraction modules. Optionally, the first encoding module can be built based on a Transformer encoder, or it can be built based on an LSTM network; this embodiment of the invention does not impose any limitations.

[0077] In some embodiments, the first processing model may include a first feature extraction module, or the first processing model may include multiple first feature extraction modules, for example, continuing to refer to Figure 6 As shown, the first processing model includes N cascaded first feature extraction modules, where N is a positive integer. Optionally, the first processing model may include 6 cascaded first feature extraction modules.

[0078] Optionally, continue to refer to Figure 6 As shown, each first feature extraction module includes: a first feature extraction unit, a first fusion unit, a second feature extraction unit, a second fusion unit, and a third feature extraction unit. The output of the first feature extraction unit is connected to the input of the first fusion unit; the output of the first fusion unit is connected to the input of the second feature extraction unit; the outputs of the first fusion unit and the second feature extraction unit are connected to the input of the second fusion unit; and the output of the second fusion unit is connected to the input of the third feature extraction unit. The first feature extraction unit is configured to extract features from the data representation input to the first feature extraction module; the first fusion unit is configured to fuse the data representation input to the first feature extraction module and the data representation output by the first feature extraction unit; the second feature extraction unit is configured to extract features from the data representation output by the first fusion unit; the second fusion unit is configured to fuse the data representation output by the first fusion unit and the data representation output by the second feature extraction unit; and the third feature extraction unit is configured to extract features from the data representation output by the second fusion unit.

[0079] In some embodiments, continue to refer to Figure 6As shown, the first feature extraction unit includes a first normalization unit and a first attention unit. The output of the first normalization unit is connected to the input of the first attention unit. The first normalization unit is configured to normalize the data representation input to the first feature extraction module; the first attention unit is configured to perform attention calculation on the data representation output by the first normalization unit.

[0080] Furthermore, the second feature extraction unit includes a second normalization unit and a first feature processing unit. The input of the second normalization unit is connected to the output of the first fusion unit, and the output of the second normalization unit is connected to the input of the first feature processing unit. The second normalization unit is configured to normalize the data representation output by the first fusion unit; the first feature processing unit is configured to perform weighted calculation on the data representation output by the second normalization unit.

[0081] In this embodiment, the first normalization unit and the second normalization unit can be constructed based on a batch normalization (BN) layer, a layer normalization (LN) layer, or a spectral normalization (SN) layer; this embodiment does not impose any limitations. In some embodiments, the first normalization unit and the second normalization unit are constructed based on an LN layer.

[0082] In some embodiments, the first attention unit includes: a plurality of third linear units, a plurality of fifth feature processing units, a fifth fusion unit, and a fourth linear unit. The plurality of third linear units correspond one-to-one with the plurality of fifth feature processing units. The input terminals of the plurality of third linear units are connected to the output terminal of the first normalization unit. The output terminal of each third linear unit is connected to the input terminal of the fifth feature processing unit corresponding to that third linear unit. The output terminals of the plurality of fifth feature processing units are connected to the input terminal of the fifth fusion unit. The output terminal of the fifth fusion unit is connected to the input terminal of the fourth linear unit. Each third linear unit is configured to perform linear transformation processing on the data representation output by the first normalization unit; each fifth feature processing unit is configured to perform attention calculation processing on the data representation output by the third linear unit corresponding to that fifth feature processing unit; the fifth fusion unit is configured to perform fusion processing on the data representation output by the plurality of fifth feature processing units; and the fourth linear unit is configured to perform linear transformation processing on the data representation output by the fifth fusion unit. This embodiment enhances the ability of the first processing model to recognize the relationships between different regions within an image patch in the first image data information by performing attention calculation in parallel by a plurality of fifth feature processing units.

[0083] In some embodiments, the first feature processing unit includes a first fully connected layer, a third activation layer, and a second fully connected layer. The input of the first fully connected layer is connected to the output of the second normalization unit, the output of the first fully connected layer is connected to the input of the third activation layer, and the output of the third activation layer is connected to the input of the second fully connected layer. The first fully connected layer is configured to perform weighted calculation processing on the data representation output by the second normalization unit; the third activation layer is configured to perform nonlinear transformation processing on the data representation output by the first fully connected layer; and the second fully connected layer is configured to perform weighted calculation processing on the data representation output by the third activation layer. This embodiment of the application, by using the first feature processing unit to perform weighted calculation processing on the data representation output by the second normalization unit, can further extract and integrate deep feature information, improving the accuracy of the determined first feature information.

[0084] In this embodiment, the third activation layer can be constructed based on the Sigmoid activation function, the ReLU activation function, or the GeLU activation function; this embodiment does not impose any limitation. In some embodiments, the third activation layer is constructed based on the GeLU activation function.

[0085] In some embodiments, continue to refer to Figure 6As shown, the third feature extraction unit includes a second feature processing unit and a first linear unit. The input of the second feature processing unit is connected to the output of the second fusion unit, and the output of the second feature processing unit is connected to the input of the first linear unit. The second feature processing unit is configured to perform weighted calculations on the data representation output by the second fusion unit; the first linear unit is configured to perform linear transformations on the data representation output by the second feature processing unit. The second feature processing unit can be constructed based on a feedforward neural network (FNN). By processing the data representation output by the second fusion unit through the second feature processing unit and the first linear unit, comprehensive extraction of image feature information can be achieved, thereby improving the accuracy of the obtained first feature information.

[0086] S302. Extract features from the first audio data information to obtain the second feature information.

[0087] In this embodiment, the second feature information is obtained by feature extraction from the first audio data information based on the second processing model, as shown below. Figure 7 As shown, the second processing model includes a second feature extraction module and several third feature extraction modules connected in series. The input of the second feature extraction module is configured to receive first audio data information, and the output of the second feature extraction module is connected to the inputs of the several third feature extraction modules. The second feature extraction module is configured to extract features from the first audio data information; the several third feature extraction modules are configured to extract features from the data representation output by the second feature extraction module to obtain second feature information.

[0088] Optionally, when the second feature extraction module extracts features from the first audio data information, it can perform a series of processes such as Fast Fourier Transform (FFT), Mel filtering, and Discrete Cosine Transform on the first audio data information, thereby extracting Mel Frequency Cepstral Coefficient (MFCC) features from the first audio data information. The MFCC features are the data representation output by the second feature extraction module.

[0089] Furthermore, the second processing model may include a third feature extraction module, or it may include multiple third feature extraction modules; this embodiment does not limit this. For example, continuing to refer to... Figure 7 The second processing model includes L cascaded third feature extraction modules, where L is a positive integer. Optionally, the second processing model may include 6 cascaded third feature extraction modules.

[0090] In some embodiments, continue to refer to Figure 7 As shown, each third feature extraction module includes: a fourth feature extraction unit, a fifth feature extraction unit, a third fusion unit, and a sixth feature extraction unit. The output of the fourth feature extraction unit is connected to the input of the fifth feature extraction unit; the outputs of the fourth and fifth feature extraction units are connected to the input of the third fusion unit; and the output of the third fusion unit is connected to the input of the sixth feature extraction unit. The fourth feature extraction unit is configured to extract features from the data representation input to the third feature extraction module; the fifth feature extraction unit is configured to extract features from the data representation output by the fourth feature extraction unit; the third fusion unit is configured to fuse the data representations output by the fourth and fifth feature extraction units; and the sixth feature extraction unit is configured to extract features from the data representation output by the third fusion unit.

[0091] Optionally, continue to refer to Figure 7 As shown, the fourth feature extraction unit includes a third feature processing unit and a second attention unit. The output of the third feature processing unit is connected to the input of the second attention unit. The third feature processing unit is configured to perform weighted calculation processing on the data representation input to the third feature extraction module; the second attention unit is configured to perform attention calculation processing on the data representation output by the third feature processing unit.

[0092] In this embodiment, the third feature processing unit can be constructed based on a feedforward neural network (FNN). By performing weighted calculation processing on the data representation input to the third feature extraction module through the third feature processing unit, the data representation input to the third feature extraction module can be mapped to a higher-dimensional space, thereby capturing more complex features and improving the accuracy of the determined second feature information.

[0093] Furthermore, the second attention unit has the same structure as the first attention unit, and the specific details can be found in the description of the first attention unit described above. To avoid repetition, this embodiment will not repeat the description here. In this embodiment, the second attention unit performs attention calculation processing on the data representation output by the third feature processing unit, enabling the second processing model to simultaneously capture and integrate multiple interactive information in different subspaces. This enhances the understanding of the contextual information of the audio signal and improves the accuracy of the determined second feature information.

[0094] In some embodiments, the fifth feature extraction unit may be based on a convolutional neural network (CNN). By extracting features from the data representation output by the fourth feature extraction unit, local features of the audio signal can be extracted, thereby improving the accuracy of the determined second feature information.

[0095] In some embodiments, continue to refer to Figure 7 As shown, the sixth feature extraction unit includes a fourth feature processing unit and a second linear unit. The input of the fourth feature processing unit is connected to the output of the third fusion unit, and the output of the fourth feature processing unit is connected to the input of the second linear unit. The fourth feature processing unit is configured to perform weighted calculations on the data representation output by the third fusion unit; the second linear unit is configured to perform linear transformations on the data representation output by the fourth feature processing unit. The fourth feature processing unit can be constructed based on a feedforward neural network (FNN). By processing the data representation output by the third fusion unit through the fourth feature processing unit and the second linear unit, comprehensive extraction of audio feature information can be achieved, thereby improving the accuracy of the obtained second feature information.

[0096] S203. Determine the target text information based on the target feature information.

[0097] In this embodiment, the target text information is the translated text information corresponding to the data information to be processed. For example, if the source text information corresponding to the data information to be processed is "The flat provided a perfect vantage point for observing the surrounding valleys", and the data processing method of steps S201 to S203 is used to process the data information to be processed, the resulting target text information is "The flat highland provided an excellent view for observing the surrounding valleys". In this embodiment, feature extraction is performed on the data information to be processed to obtain target feature information, and the target text information is determined based on the target feature information, which can improve the accuracy of the determined target text information.

[0098] Optionally, the target feature information includes first feature information and / or second feature information. In this embodiment, when determining the target text information based on the target feature information, the first feature information and the second feature information can be directly decoded to obtain the target text information, or the third feature information can be determined based on the first feature information and the second feature information, and then the second feature information and the third feature information can be decoded to obtain the target feature information. This embodiment does not limit the scope of the target feature information.

[0099] In some embodiments, the target feature information includes first feature information and / or second feature information, referring to... Figure 8 As shown, the step S203 above, which determines the target text information based on the target feature information, may include steps S401 to S402, as follows:

[0100] S401. Based on the first feature information and the second feature information, determine the third feature information.

[0101] In this embodiment, the third feature information is a data representation associated with the second feature information, filtered from the first feature information based on the second feature information. Optionally, the first image data information includes several image data information to be processed, which are image block information obtained by segmenting the first image data information. The first feature information includes several feature information to be processed, which are data representations extracted from the several image data information to be processed respectively. The step of determining the third feature information based on the first and second feature information specifically includes: calculating and processing each feature information to be processed and the second feature information to obtain a first feature value corresponding to each feature information to be processed; and filtering the several feature information to be processed based on the first feature value to obtain the third feature information.

[0102] In this embodiment, the first feature value is used to characterize the degree of correlation between each feature information to be processed and the second feature information. When filtering several feature information to be processed based on the first feature value, feature information with a first feature value greater than a first threshold can be selected from several feature information to be processed and determined as the third feature information. Alternatively, multiple feature information with the highest first feature value can be selected from several feature information to be processed and determined as the third feature information. This embodiment does not limit this.

[0103] In some embodiments, the step of calculating and processing each feature information to be processed and the second feature information to obtain the first feature value corresponding to each feature information to be processed specifically includes: performing a transpose operation on each feature information to be processed to obtain the fourth feature information corresponding to each feature information to be processed; multiplying the fourth feature information and the second feature information to obtain the fifth feature information corresponding to each feature information to be processed; dividing the fifth feature information and the second feature value corresponding to the feature information to be processed to obtain the third feature value corresponding to each feature information to be processed; and performing feature transformation processing on the third feature value to obtain the first feature value corresponding to each feature information to be processed.

[0104] Optionally, the second eigenvalue is used to characterize the feature dimension corresponding to the feature information to be processed, and the process of determining the third eigenvalue can be expressed as follows: Among them, Atti F represents the third feature value corresponding to the i-th feature information to be processed. voice E represents the second feature information. img-i Let d represent the i-th feature information to be processed. n Let T denote the first eigenvalue and T denote the transpose operation.

[0105] Furthermore, when performing feature transformation on the third eigenvalue, the Sigmoid activation function can be used to perform feature transformation on the third eigenvalue, thereby compressing and mapping the third eigenvalue to a value between 0 and 1. Optionally, the larger the first eigenvalue, the higher the correlation between the feature information to be processed corresponding to the first eigenvalue and the second feature information; the smaller the first eigenvalue, the lower the correlation between the feature information to be processed corresponding to the first eigenvalue and the second feature information.

[0106] S402. Decode the second and third feature information to obtain the target text information.

[0107] In this embodiment, a third feature is determined based on the first feature information and the second feature information. The target text information is obtained by decoding the second feature information and the third feature information. Feature information that is not related to the second feature information can be filtered out from the first feature information, thereby improving data processing efficiency and the accuracy of the obtained target text information.

[0108] In some embodiments, the step of decoding the second feature information and the third feature information to obtain the target text information specifically includes: inputting the second feature information and the third feature information into a decoder, and decoding the second feature information and the third feature information through the decoder to obtain the target text information. The decoder can be built based on a Transformer decoder, an LSTM network, or a GRU network; this embodiment of the invention does not limit the scope of the invention.

[0109] Optionally, the process of determining the target text information can be represented as: Y = Dec(F voice E img ), where Y represents the target text information, F voice E represents the second feature information. img Representing the third feature information, Dec(·) indicates that F is processed by the decoder. voice and E img Perform decoding.

[0110] In some embodiments, the first processing model, the second processing model, and the decoder are obtained by jointly training the second network model, the third network model, and the fourth network model using a second training sample set. The model structure of the second network model is the same as that of the first processing model, the model structure of the third network model is the same as that of the second processing model, and the model structure of the fourth network model is the same as that of the decoder. Specific details can be found in the aforementioned structures of the first processing model, the second processing model, and the decoder; these will not be repeated here.

[0111] Optionally, the second training sample set includes second sample audio data, first sample image data corresponding to the second sample audio data, and target word information for multiple time steps corresponding to the second sample audio data. The first processing model, the second processing model, and the decoder are trained through the following steps: inputting the first sample image data corresponding to the second sample audio data into the second network model, and outputting the first sample feature information corresponding to the second sample audio data through the second network model; inputting the second sample audio data into the third network model, and outputting the second sample feature information corresponding to the second sample audio data through the third network model; determining the third sample feature information corresponding to the second sample audio data based on the first sample feature information and the second sample feature information; inputting the third sample feature information and the second sample feature information into the fourth network model, and outputting each target word information for multiple time steps through the fourth network model. The second probability information corresponding to the target word information at each time step; the target loss value is determined based on the first sample feature information, the second sample feature information, the third sample feature information, and the second probability value; if the target loss value does not meet the second condition, the model parameters of the second network model, the third network model, and the fourth network model are adjusted based on the second learning rate, and the steps of inputting the first sample image data corresponding to the second sample audio data into the second network model and outputting the first sample feature information corresponding to the second sample audio data through the second network model are continued until the target loss value meets the second condition or the number of model updates reaches the second threshold; the updated second network model is determined as the first processing model, the updated third network model is determined as the second processing model, and the updated fourth network model is determined as the decoder. The target loss value meeting the second condition can be defined as the target loss value being less than the third threshold or the difference between the two obtained target loss values ​​being less than the fourth threshold.

[0112] In some embodiments, the step of determining the target loss value based on the first sample feature information, the second sample feature information, the third sample feature information, and the second probability value specifically includes: determining the second loss value based on the first sample feature information, the second sample feature information, and the third sample feature information; determining the third loss value based on the second probability value; and performing a weighted summation of the second loss value and the third loss value to obtain the target loss value.

[0113] Optionally, the process of determining the target loss value is: L=α*L2(ξ)+β*L3(η), where L represents the target loss value, L2(ξ) represents the second loss value, L3(η) represents the third loss value, ξ represents the set of model parameters of the second and third network models, η represents the set of model parameters of the fourth network model, and α and β represent hyperparameters.

[0114] Furthermore, the process of determining the third sample feature information corresponding to the second sample audio data based on the first sample feature information and the second sample feature information is the same as the step of determining the third feature information based on the first and second feature information in step S401. Specifically, the process of determining the third feature information based on the first and second feature information can be referred to. To avoid repetition, this embodiment will not repeat the steps here.

[0115] In some embodiments, the first sample feature information includes a plurality of first sample features, and the third sample feature information is a feature set composed of sample features selected from the plurality of first sample features. The step of determining the second loss value based on the first sample feature information, the second sample feature information, and the third sample feature information specifically includes: determining a fourth feature value corresponding to each first sample feature based on the first sample feature information and the third sample feature information; the fourth feature value is used to characterize whether the first sample feature belongs to the features in the third sample feature information; performing similarity calculation on each first sample feature and the second sample feature information to obtain first similarity information; and determining the second loss value based on the first similarity information and the fourth feature value.

[0116] Alternatively, the process of determining the second loss value can be expressed as: Where L2(ξ) represents the second loss value, ξ represents the set of model parameters for the second and third network models, R represents the number of first sample features in the first sample feature information, and D i represents the similarity between the i-th feature of the first sample and the feature information of the second sample, k represents the fourth feature value, and m represents the margin parameter.

[0117] Optionally, the similarity between the i-th first sample feature and the second sample feature can be the Euclidean distance between the i-th first sample feature and the second sample feature. Among them, A img-i Let B represent the feature of the i-th first sample. voiceThis represents the second sample feature information. Further, if the i-th first sample feature belongs to the features in the third sample feature information, then the fourth feature value corresponding to the i-th first sample feature is 1; if the i-th first sample feature does not belong to the features in the third sample feature information, then the fourth feature value corresponding to the i-th first sample feature is 0.

[0118] Furthermore, the process of determining the third loss value can be expressed as: Where L3(η) represents the third loss value, η represents the set of model parameters for the fourth network model, T represents the number of time steps, and z t The target word representing time step t, p(z) t |z1,…,z t-1 η) represents the second probability information of the target word at time step t.

[0119] In some embodiments, the second learning rate threshold can range from 900 to 1100, and optionally, it can be 1000. During model training, a dynamic learning rate adjustment strategy can be used to adjust the second learning rate. For example, the initial second learning rate can be set to 0.001, and every 20 model updates, the second learning rate is adjusted exponentially at a decay rate of 0.9. This can improve the model's generalization ability, enabling the model to converge faster in the early stages of training and avoiding getting stuck in local minima in the later stages.

[0120] Furthermore, the second, third, and fourth network models can be trained on computer devices with powerful Graphics Processing Unit (GPU) computing capabilities, high system memory (e.g., 32GB or more), and large-capacity high-speed hard disks or solid-state drives (SSDs) (e.g., at least 1TB SSDs). GPUs provide efficient parallel computing capabilities, accelerating complex calculations during model training and significantly shortening training time. High system memory ensures that there will be no memory shortage during data loading, model parameter storage, and intermediate result calculation. Large-capacity high-speed hard disks or SSDs guarantee data read and write speeds, reducing the impact of data loading time on training efficiency.

[0121] In some embodiments, after determining the target text information based on the target feature information in step S203 above, the steps may include: obtaining target display location information; and displaying the first text information corresponding to the target text information and the data information to be processed based on the target display location information. This embodiment displays the target text information and the first text information based on the target display location information, allowing users to intuitively obtain the target text information and the first text information through the interface, enabling users of different languages ​​to communicate without barriers.

[0122] In this embodiment, the first text information corresponding to the data information to be processed is the source text information corresponding to the data information to be processed. The first text information and the target text information are text information in different languages, and the semantics of the first text information and the target text information are the same or similar. For example, the first text information is "The flat provided a perfect vantage point for observing the surrounding valleys", and the target text information is "The flat highland provided a perfect vantage point for observing the surrounding valleys".

[0123] Furthermore, the target display position information refers to the display positions of the target text information and the first text information in the corresponding screen. For example, when the data processing method is applied to an AR device, the target display position information refers to the display positions of the target text information and the first text information on the display interface of the AR device. The target display position information can be pre-set position information, or it can be position information determined based on other information. This embodiment does not limit this. For example, if the target display position information is consistent with the user's line of sight and viewing angle focus, then the target display position information can be determined based on the user's line of sight and viewing angle focus.

[0124] In summary, the data processing method provided in this implementation scheme obtains target feature information by acquiring the data to be processed, extracting features from the data to be processed, and determining target text information based on the target feature information. This scheme improves the accuracy of the determined target text information by extracting features from the data to be processed and determining target text information based on the target feature information. Furthermore, feature extraction is performed on the first image data to obtain first feature information, and feature extraction is performed on the first audio data to obtain second feature information. This can be combined with the first image data for translation, further improving the accuracy of the determined target text information. Even further, a third feature information is determined based on the first and second feature information. Decoding the second and third feature information to obtain the target text information allows irrelevant feature information to be filtered out from the first feature information, improving data processing efficiency and the accuracy of the target text information. Furthermore, processing the second audio data information using a third processing model to obtain the first audio data information, and then processing the first audio data information, can filter out noise signals in the second audio data information, further improving the accuracy of the determined target text information.

[0125] To better implement the data processing method in the embodiments of this application, a data processing system is also provided in the embodiments of this application, based on the data processing method, with reference to... Figure 9 As shown, the data processing system 600 includes:

[0126] Information acquisition module 610 is used to acquire data information to be processed;

[0127] The data processing module 620 is used to extract features from the data to be processed to obtain target feature information;

[0128] The data determination module 630 is used to determine the target text information based on the target feature information.

[0129] In this embodiment of the application, by extracting features from the data information to be processed to obtain target feature information, and determining target text information based on the target feature information, the accuracy of the determined target text information can be improved.

[0130] In some embodiments of this application, the data information to be processed includes first image data information and / or first audio data information, the target feature information includes first feature information and / or second feature information, and the data processing module 620 performs feature extraction on the data information to be processed to obtain target feature information, including:

[0131] Feature extraction is performed on the first image data to obtain first feature information; and / or,

[0132] Feature extraction is performed on the first audio data to obtain the second feature information.

[0133] In some embodiments of this application, the first feature information is obtained by feature extraction of the first image data information based on a first processing model. The first processing model includes: a first encoding module and a plurality of serially connected first feature extraction modules.

[0134] The input end of the first encoding module is configured to receive first image data information, and the output end of the first encoding module is connected to the input ends of several first feature extraction modules.

[0135] The first encoding module is configured to encode the first image data information;

[0136] Several first feature extraction modules are configured to extract features from the data representation output by the first encoding module to obtain first feature information.

[0137] In some embodiments of this application, each first feature extraction module includes: a first feature extraction unit, a first fusion unit, a second feature extraction unit, a second fusion unit, and a third feature extraction unit;

[0138] The output of the first feature extraction unit is connected to the input of the first fusion unit, the output of the first fusion unit is connected to the input of the second feature extraction unit, the output of the first fusion unit and the output of the second feature extraction unit are connected to the input of the second fusion unit, and the output of the second fusion unit is connected to the input of the third feature extraction unit.

[0139] The first feature extraction unit is configured to extract features from the data representation input to the first feature extraction module;

[0140] The first fusion unit is configured to perform fusion processing on the data representation input to the first feature extraction module and the data representation output by the first feature extraction unit;

[0141] The second feature extraction unit is configured to extract features from the data representation output by the first fusion unit;

[0142] The second fusion unit is configured to perform fusion processing on the data representation output by the first fusion unit and the data representation output by the second feature extraction unit;

[0143] The third feature extraction unit is configured to extract features from the data representation output by the second fusion unit.

[0144] In some embodiments of this application, the first feature extraction unit includes: a first normalization unit and a first attention unit;

[0145] The output of the first normalization unit is connected to the input of the first attention unit.

[0146] The first normalization unit is configured to normalize the data representation input to the first feature extraction module;

[0147] The first attention unit is configured to perform attention computation processing on the data representation output by the first normalization unit; and / or,

[0148] The second feature extraction unit includes: a second normalization unit and a first feature processing unit;

[0149] The input of the second normalization unit is connected to the output of the first fusion unit, and the output of the second normalization unit is connected to the input of the first feature processing unit.

[0150] The second normalization unit is configured to normalize the data representation output by the first fusion unit;

[0151] The first feature processing unit is configured to perform weighted calculation processing on the data representation output by the second normalization unit; and / or,

[0152] The third feature extraction unit includes: a second feature processing unit and a first linear unit;

[0153] The input of the second feature processing unit is connected to the output of the second fusion unit, and the output of the second feature processing unit is connected to the input of the first linear unit.

[0154] The second feature processing unit is configured to perform weighted calculation processing on the data representation output by the second fusion unit;

[0155] The first linear unit is configured to perform linear transformation processing on the data representation output by the second feature processing unit.

[0156] In some embodiments of this application, the second feature information is obtained by feature extraction of the first audio data information based on a second processing model. The second processing model includes: a second feature extraction module and several third feature extraction modules connected in series.

[0157] The input of the second feature extraction module is configured to receive the first audio data information, and the output of the second feature extraction module is connected to the input of several third feature extraction modules.

[0158] The second feature extraction module is configured to extract features from the first audio data information;

[0159] Several third feature extraction modules are configured to extract features from the data representation output by the second feature extraction module to obtain second feature information.

[0160] In some embodiments of this application, each third feature extraction module includes: a fourth feature extraction unit, a fifth feature extraction unit, a third fusion unit, and a sixth feature extraction unit;

[0161] The output of the fourth feature extraction unit is connected to the input of the fifth feature extraction unit, the outputs of the fourth and fifth feature extraction units are connected to the input of the third fusion unit, and the output of the third fusion unit is connected to the input of the sixth feature extraction unit.

[0162] The fourth feature extraction unit is configured to extract features from the data representation input to the third feature extraction module;

[0163] The fifth feature extraction unit is configured to extract features from the data representation output by the fourth feature extraction unit;

[0164] The third fusion unit is configured to perform fusion processing on the data representation output by the fourth feature extraction unit and the data representation output by the fifth feature extraction unit;

[0165] The sixth feature extraction unit is configured to extract features from the data representation output by the third fusion unit.

[0166] In some embodiments of this application, the fourth feature extraction unit includes: a third feature processing unit and a second attention unit;

[0167] The output of the third feature processing unit is connected to the input of the second attention unit.

[0168] The third feature processing unit is configured to perform weighted calculation processing on the data representation input to the third feature extraction module;

[0169] The second attention unit is configured to perform attention computation processing on the data representation output by the third feature processing unit; and / or,

[0170] The sixth feature extraction unit includes: the fourth feature processing unit and the second linear unit;

[0171] The input of the fourth feature processing unit is connected to the output of the third fusion unit, and the output of the fourth feature processing unit is connected to the input of the second linear unit.

[0172] The fourth feature processing unit is configured to perform weighted calculation processing on the data representation output by the third fusion unit;

[0173] The second linear unit is configured to perform linear transformation processing on the data representation output by the fourth feature processing unit.

[0174] In some embodiments of this application, the first audio data information is obtained by processing the second audio data information through a third processing model. The third processing model includes: a fourth feature extraction module, a fifth feature extraction module, a first feature fusion module, and a decoding module.

[0175] The input of the fourth feature extraction module is configured to receive second audio data information. The output of the fourth feature extraction module is connected to the input of the fifth feature extraction module. The outputs of the fourth and fifth feature extraction modules are connected to the input of the first feature fusion module. The output of the first feature fusion module is connected to the input of the decoding module.

[0176] The fourth feature extraction module is configured to extract features from the second audio data.

[0177] The fifth feature extraction module is configured to extract features from the data representation output by the fourth feature extraction module;

[0178] The first feature fusion module is configured to fuse the data representations output by the fourth feature extraction module and the data representations output by the fifth feature extraction module.

[0179] The decoding module is configured to decode the data representation output by the first feature fusion module.

[0180] In some embodiments of this application, the fourth feature extraction module includes: a first encoding unit, a seventh feature extraction unit, and a third normalization unit;

[0181] The input of the first encoding unit is configured to receive the second audio data information, the output of the first encoding unit is connected to the input of the seventh feature extraction unit, and the output of the seventh feature extraction unit is connected to the input of the third normalization unit.

[0182] The first encoding unit is configured to encode the second audio data information;

[0183] The seventh feature extraction unit is configured to extract features from the data representation output by the first encoding unit;

[0184] The third normalization unit is configured to normalize the data representation output by the seventh feature extraction unit; and / or,

[0185] The fifth feature extraction module includes: an eighth feature extraction unit and a first activation unit;

[0186] The input of the eighth feature extraction unit is configured to receive the data representation output by the fourth feature extraction module, and the output of the eighth feature extraction unit is connected to the input of the first activation unit.

[0187] The eighth feature extraction unit includes at least one first convolutional unit, and each first convolutional unit includes: a first convolutional layer, a first normalization layer, and a second normalization layer;

[0188] The output of the first convolutional layer is connected to the input of the first normalized layer, and the output of the first normalized layer is connected to the input of the second normalized layer.

[0189] In some embodiments of this application, the seventh feature extraction unit includes: a second convolution unit and a third convolution unit;

[0190] The output of the first encoding unit is connected to the input of the second convolutional unit, and the output of the second convolutional unit is connected to the input of the third convolutional unit.

[0191] The second convolutional unit includes: a second convolutional layer, a third normalization layer, and a first activation layer;

[0192] Wherein, the output of the first coding unit is connected to the input of the second convolutional layer, the output of the second convolutional layer is connected to the input of the third normalization layer, and the output of the third normalization layer is connected to the input of the first activation layer.

[0193] The third convolutional unit includes: a third convolutional layer, a fourth normalization layer, and a second activation layer;

[0194] The output of the second convolutional unit is connected to the input of the third convolutional layer, the output of the third convolutional layer is connected to the input of the fourth normalization layer, and the output of the fourth normalization layer is connected to the input of the second activation layer.

[0195] In some embodiments of this application, the target feature information includes first feature information and / or second feature information. The data determination module 630 determines target text information based on the target feature information, including:

[0196] Based on the first and second feature information, the third feature information is determined;

[0197] The second and third feature information are decoded to obtain the target text information.

[0198] In some embodiments of this application, the first feature information includes several feature information to be processed. The data determination module 630 determines the third feature information based on the first feature information and the second feature information, including:

[0199] Each feature information to be processed and the second feature information are calculated and processed to obtain the first feature value corresponding to each feature information to be processed.

[0200] Based on the first feature value, several feature information to be processed are filtered to obtain the third feature information.

[0201] In some embodiments of this application, the data determination module 630 performs calculations on each feature information to be processed and the second feature information to obtain a first feature value corresponding to each feature information to be processed, including:

[0202] Perform a transpose operation on each feature information to be processed to obtain the fourth feature information corresponding to each feature information;

[0203] The fourth feature information and the second feature information are multiplied together to obtain the fifth feature information corresponding to each feature information to be processed.

[0204] Divide the second feature value corresponding to the fifth feature information and the feature information to be processed to obtain the third feature value corresponding to each feature information to be processed.

[0205] The third feature value is subjected to feature transformation processing to obtain the first feature value corresponding to each feature information to be processed.

[0206] In some embodiments of this application, after the data determination module 630 determines the target text information based on the target feature information, the data determination module 630 is further configured to:

[0207] Obtain the target display location information;

[0208] The first text information corresponding to the target text information and the data information to be processed is displayed based on the target display location information.

[0209] This application embodiment also provides a computer device, the computer device including:

[0210] One or more processors;

[0211] Memory; and

[0212] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the data processing method in any of the embodiments described above.

[0213] This application also provides a computer device, such as... Figure 10 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:

[0214] The computer device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that... Figure 10 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0215] The processor 801 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802, thereby providing overall monitoring of the computer device. Optionally, the processor 801 may include one or more processing cores; optionally, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 801.

[0216] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0217] The computer device also includes a power supply 803 that supplies power to the various components. Optionally, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0218] The computer device may also include an input unit 804, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0219] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 runs the application programs stored in the memory 802 to realize various functions, as follows:

[0220] Obtain the data to be processed;

[0221] Feature extraction is performed on the data to be processed to obtain target feature information;

[0222] Based on target feature information, target text information is determined.

[0223] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0224] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the data processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps:

[0225] Obtain the data to be processed;

[0226] Feature extraction is performed on the data to be processed to obtain target feature information;

[0227] Based on target feature information, target text information is determined.

[0228] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0229] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0230] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0231] The data processing method and system provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method, characterized in that, include: Obtain the data to be processed; Feature extraction is performed on the data to be processed to obtain target feature information; Based on the target feature information, the target text information is determined.

2. The method according to claim 1, characterized in that, The data information to be processed includes first image data information and / or first audio data information; the target feature information includes first feature information and / or second feature information. The step of extracting features from the data to be processed to obtain target feature information includes: Feature extraction is performed on the first image data information to obtain first feature information; And / or, Feature extraction is performed on the first audio data information to obtain the second feature information.

3. The method according to claim 2, characterized in that, The first feature information is obtained by feature extraction of the first image data information based on the first processing model. The first processing model includes: a first encoding module and a plurality of cascaded first feature extraction modules. The input end of the first encoding module is configured to receive the first image data information, and the output end of the first encoding module is connected to the input ends of several first feature extraction modules. The first encoding module is configured to encode the first image data information; Several of the first feature extraction modules are configured to extract features from the data representation output by the first encoding module to obtain first feature information.

4. The method according to claim 3, characterized in that, Each of the first feature extraction modules includes: a first feature extraction unit, a first fusion unit, a second feature extraction unit, a second fusion unit, and a third feature extraction unit; Wherein, the output end of the first feature extraction unit is connected to the input end of the first fusion unit, the output end of the first fusion unit is connected to the input end of the second feature extraction unit, the output ends of the first fusion unit and the second feature extraction unit are connected to the input end of the second fusion unit, and the output end of the second fusion unit is connected to the input end of the third feature extraction unit; The first feature extraction unit is configured to extract features from the data representation input to the first feature extraction module; The first fusion unit is configured to perform fusion processing on the data representation input to the first feature extraction module and the data representation output by the first feature extraction unit; The second feature extraction unit is configured to extract features from the data representation output by the first fusion unit; The second fusion unit is configured to perform fusion processing on the data representation output by the first fusion unit and the data representation output by the second feature extraction unit; The third feature extraction unit is configured to extract features from the data representation output by the second fusion unit.

5. The method according to claim 4, characterized in that, The first feature extraction unit includes: a first normalization unit and a first attention unit; Wherein, the output of the first normalization unit is connected to the input of the first attention unit; The first normalization unit is configured to normalize the data representation input to the first feature extraction module; The first attention unit is configured to perform attention computation processing on the data representation output by the first normalization unit; and / or, The second feature extraction unit includes: a second normalization unit and a first feature processing unit; Wherein, the input terminal of the second normalization unit is connected to the output terminal of the first fusion unit, and the output terminal of the second normalization unit is connected to the input terminal of the first feature processing unit; The second normalization unit is configured to normalize the data representation output by the first fusion unit; The first feature processing unit is configured to perform weighted calculation processing on the data representation output by the second normalization unit; and / or, The third feature extraction unit includes: a second feature processing unit and a first linear unit; Wherein, the input end of the second feature processing unit is connected to the output end of the second fusion unit, and the output end of the second feature processing unit is connected to the input end of the first linear unit; The second feature processing unit is configured to perform weighted calculation processing on the data representation output by the second fusion unit; The first linear unit is configured to perform linear transformation processing on the data representation output by the second feature processing unit.

6. The method according to claim 2, characterized in that, The second feature information is obtained by extracting features from the first audio data information based on the second processing model. The second processing model includes: a second feature extraction module and several cascaded third feature extraction modules. The input terminal of the second feature extraction module is configured to receive the first audio data information, and the output terminal of the second feature extraction module is connected to the input terminals of several third feature extraction modules. The second feature extraction module is configured to extract features from the first audio data information; Several of the third feature extraction modules are configured to extract features from the data representation output by the second feature extraction module to obtain second feature information.

7. The method according to claim 6, characterized in that, Each of the third feature extraction modules includes: a fourth feature extraction unit, a fifth feature extraction unit, a third fusion unit, and a sixth feature extraction unit; The output of the fourth feature extraction unit is connected to the input of the fifth feature extraction unit; the outputs of the fourth and fifth feature extraction units are connected to the input of the third fusion unit; and the output of the third fusion unit is connected to the input of the sixth feature extraction unit. The fourth feature extraction unit is configured to extract features from the data representation input to the third feature extraction module; The fifth feature extraction unit is configured to extract features from the data representation output by the fourth feature extraction unit; The third fusion unit is configured to perform fusion processing on the data representation output by the fourth feature extraction unit and the data representation output by the fifth feature extraction unit; The sixth feature extraction unit is configured to extract features from the data representation output by the third fusion unit.

8. The method according to claim 7, characterized in that, The fourth feature extraction unit includes: a third feature processing unit and a second attention unit; The output of the third feature processing unit is connected to the input of the second attention unit. The third feature processing unit is configured to perform weighted calculation processing on the data representation input to the third feature extraction module; The second attention unit is configured to perform attention calculation processing on the data representation output by the third feature processing unit; and / or, The sixth feature extraction unit includes: a fourth feature processing unit and a second linear unit; The input terminal of the fourth feature processing unit is connected to the output terminal of the third fusion unit, and the output terminal of the fourth feature processing unit is connected to the input terminal of the second linear unit. The fourth feature processing unit is configured to perform weighted calculation processing on the data representation output by the third fusion unit; The second linear unit is configured to perform linear transformation processing on the data representation output by the fourth feature processing unit.

9. The method according to claim 2, characterized in that, The first audio data information is obtained by processing the second audio data information through a third processing model. The third processing model includes: a fourth feature extraction module, a fifth feature extraction module, a first feature fusion module, and a decoding module. The input terminal of the fourth feature extraction module is configured to receive the second audio data information. The output terminal of the fourth feature extraction module is connected to the input terminal of the fifth feature extraction module. The output terminals of the fourth and fifth feature extraction modules are connected to the input terminal of the first feature fusion module. The output terminal of the first feature fusion module is connected to the input terminal of the decoding module. The fourth feature extraction module is configured to extract features from the second audio data information; The fifth feature extraction module is configured to extract features from the data representation output by the fourth feature extraction module; The first feature fusion module is configured to perform fusion processing on the data representation output by the fourth feature extraction module and the data representation output by the fifth feature extraction module; The decoding module is configured to decode the data representation output by the first feature fusion module.

10. The method according to claim 9, characterized in that, The fourth feature extraction module includes: a first encoding unit, a seventh feature extraction unit, and a third normalization unit; The input terminal of the first encoding unit is configured to receive the second audio data information, the output terminal of the first encoding unit is connected to the input terminal of the seventh feature extraction unit, and the output terminal of the seventh feature extraction unit is connected to the input terminal of the third normalization unit. The first encoding unit is configured to encode the second audio data information; The seventh feature extraction unit is configured to extract features from the data representation output by the first encoding unit; The third normalization unit is configured to normalize the data representation output by the seventh feature extraction unit; and / or, The fifth feature extraction module includes: an eighth feature extraction unit and a first activation unit; The input terminal of the eighth feature extraction unit is configured to receive the data representation output by the fourth feature extraction module, and the output terminal of the eighth feature extraction unit is connected to the input terminal of the first activation unit. The eighth feature extraction unit includes at least one first convolutional unit, and each first convolutional unit includes: a first convolutional layer, a first normalization layer, and a second normalization layer; The output of the first convolutional layer is connected to the input of the first normalization layer, and the output of the first normalization layer is connected to the input of the second normalization layer.

11. The method according to claim 10, characterized in that, The seventh feature extraction unit includes: a second convolutional unit and a third convolutional unit; Wherein, the output of the first encoding unit is connected to the input of the second convolution unit, and the output of the second convolution unit is connected to the input of the third convolution unit; The second convolutional unit includes: a second convolutional layer, a third normalization layer, and a first activation layer; Wherein, the output of the first encoding unit is connected to the input of the second convolutional layer, the output of the second convolutional layer is connected to the input of the third normalization layer, and the output of the third normalization layer is connected to the input of the first activation layer; The third convolutional unit includes: a third convolutional layer, a fourth normalization layer, and a second activation layer; The output of the second convolutional unit is connected to the input of the third convolutional layer, the output of the third convolutional layer is connected to the input of the fourth normalization layer, and the output of the fourth normalization layer is connected to the input of the second activation layer.

12. The method according to claim 1, characterized in that, The target feature information includes first feature information and / or, second feature information; The step of determining the target text information based on the target feature information includes: Based on the first feature information and the second feature information, the third feature information is determined; The second feature information and the third feature information are decoded to obtain the target text information.

13. The method according to claim 12, characterized in that, The first feature information includes several feature information to be processed; The step of determining the third feature information based on the first feature information and the second feature information includes: Each of the features to be processed and the second feature information are calculated and processed to obtain a first feature value corresponding to each feature to be processed; Based on the first feature value, several of the features to be processed are filtered to obtain the third feature information.

14. The method according to claim 13, characterized in that, The step of calculating and processing each of the unprocessed feature information and the second feature information to obtain a first feature value corresponding to each unprocessed feature information includes: Perform a transpose operation on each of the features to be processed to obtain the fourth feature information corresponding to each of the features to be processed. The fourth feature information and the second feature information are multiplied together to obtain the fifth feature information corresponding to each feature information to be processed. Divide the fifth feature information and the second feature value corresponding to the feature information to be processed to obtain the third feature value corresponding to each feature information to be processed. The third feature value is subjected to feature transformation processing to obtain the first feature value corresponding to each feature information to be processed.

15. The method according to any one of claims 1 to 14, characterized in that, After determining the target text information based on the target feature information, the process includes: Obtain the target display location information; Based on the target display location information, the first text information corresponding to the target text information and the data information to be processed is displayed.

16. A system, characterized in that, include: The information acquisition module is used to acquire data information to be processed; The data processing module is used to extract features from the data to be processed to obtain target feature information; The data determination module is used to determine target text information based on the target feature information; Optionally, the data information to be processed includes first image data information and / or first audio data information, the target feature information includes first feature information and / or second feature information, and the data processing module performs feature extraction on the data information to be processed to obtain target feature information, including: Feature extraction is performed on the first image data information to obtain first feature information; and / or, The first audio data information is subjected to feature extraction to obtain the second feature information; Optionally, the first feature information is obtained by feature extraction of the first image data information based on a first processing model, wherein the first processing model includes: a first encoding module and a plurality of cascaded first feature extraction modules; The input end of the first encoding module is configured to receive the first image data information, and the output end of the first encoding module is connected to the input ends of several first feature extraction modules. The first encoding module is configured to encode the first image data information; Several of the first feature extraction modules are configured to extract features from the data representation output by the first encoding module to obtain first feature information; Optionally, each of the first feature extraction modules includes: a first feature extraction unit, a first fusion unit, a second feature extraction unit, a second fusion unit, and a third feature extraction unit; Wherein, the output end of the first feature extraction unit is connected to the input end of the first fusion unit, the output end of the first fusion unit is connected to the input end of the second feature extraction unit, the output ends of the first fusion unit and the second feature extraction unit are connected to the input end of the second fusion unit, and the output end of the second fusion unit is connected to the input end of the third feature extraction unit; The first feature extraction unit is configured to extract features from the data representation input to the first feature extraction module; The first fusion unit is configured to perform fusion processing on the data representation input to the first feature extraction module and the data representation output by the first feature extraction unit; The second feature extraction unit is configured to extract features from the data representation output by the first fusion unit; The second fusion unit is configured to perform fusion processing on the data representation output by the first fusion unit and the data representation output by the second feature extraction unit; The third feature extraction unit is configured to extract features from the data representation output by the second fusion unit; Optionally, the first feature extraction unit includes: a first normalization unit and a first attention unit; Wherein, the output of the first normalization unit is connected to the input of the first attention unit; The first normalization unit is configured to normalize the data representation input to the first feature extraction module; The first attention unit is configured to perform attention computation processing on the data representation output by the first normalization unit; and / or, The second feature extraction unit includes: a second normalization unit and a first feature processing unit; Wherein, the input terminal of the second normalization unit is connected to the output terminal of the first fusion unit, and the output terminal of the second normalization unit is connected to the input terminal of the first feature processing unit; The second normalization unit is configured to normalize the data representation output by the first fusion unit; The first feature processing unit is configured to perform weighted calculation processing on the data representation output by the second normalization unit; and / or, The third feature extraction unit includes: a second feature processing unit and a first linear unit; Wherein, the input end of the second feature processing unit is connected to the output end of the second fusion unit, and the output end of the second feature processing unit is connected to the input end of the first linear unit; The second feature processing unit is configured to perform weighted calculation processing on the data representation output by the second fusion unit; The first linear unit is configured to perform a linear transformation on the data representation output by the second feature processing unit; Optionally, the second feature information is obtained by feature extraction of the first audio data information based on the second processing model, and the second processing model includes: a second feature extraction module and several cascaded third feature extraction modules; The input terminal of the second feature extraction module is configured to receive the first audio data information, and the output terminal of the second feature extraction module is connected to the input terminals of several third feature extraction modules. The second feature extraction module is configured to extract features from the first audio data information; Several of the third feature extraction modules are configured to extract features from the data representation output by the second feature extraction module to obtain second feature information; Optionally, each of the third feature extraction modules includes: a fourth feature extraction unit, a fifth feature extraction unit, a third fusion unit, and a sixth feature extraction unit; The output of the fourth feature extraction unit is connected to the input of the fifth feature extraction unit; the outputs of the fourth and fifth feature extraction units are connected to the input of the third fusion unit; and the output of the third fusion unit is connected to the input of the sixth feature extraction unit. The fourth feature extraction unit is configured to extract features from the data representation input to the third feature extraction module; The fifth feature extraction unit is configured to extract features from the data representation output by the fourth feature extraction unit; The third fusion unit is configured to perform fusion processing on the data representation output by the fourth feature extraction unit and the data representation output by the fifth feature extraction unit; The sixth feature extraction unit is configured to extract features from the data representation output by the third fusion unit; Optionally, the fourth feature extraction unit includes: a third feature processing unit and a second attention unit; The output of the third feature processing unit is connected to the input of the second attention unit. The third feature processing unit is configured to perform weighted calculation processing on the data representation input to the third feature extraction module; The second attention unit is configured to perform attention calculation processing on the data representation output by the third feature processing unit; and / or, The sixth feature extraction unit includes: a fourth feature processing unit and a second linear unit; The input terminal of the fourth feature processing unit is connected to the output terminal of the third fusion unit, and the output terminal of the fourth feature processing unit is connected to the input terminal of the second linear unit. The fourth feature processing unit is configured to perform weighted calculation processing on the data representation output by the third fusion unit; The second linear unit is configured to perform a linear transformation on the data representation output by the fourth feature processing unit: Optionally, the first audio data information is obtained by processing the second audio data information through a third processing model, the third processing model including: a fourth feature extraction module, a fifth feature extraction module, a first feature fusion module and a decoding module; The input terminal of the fourth feature extraction module is configured to receive the second audio data information. The output terminal of the fourth feature extraction module is connected to the input terminal of the fifth feature extraction module. The output terminals of the fourth and fifth feature extraction modules are connected to the input terminal of the first feature fusion module. The output terminal of the first feature fusion module is connected to the input terminal of the decoding module. The fourth feature extraction module is configured to extract features from the second audio data information; The fifth feature extraction module is configured to extract features from the data representation output by the fourth feature extraction module; The first feature fusion module is configured to perform fusion processing on the data representation output by the fourth feature extraction module and the data representation output by the fifth feature extraction module; The decoding module is configured to decode the data representation output by the first feature fusion module. Optionally, the fourth feature extraction module includes: a first encoding unit, a seventh feature extraction unit, and a third normalization unit; The input terminal of the first encoding unit is configured to receive the second audio data information, the output terminal of the first encoding unit is connected to the input terminal of the seventh feature extraction unit, and the output terminal of the seventh feature extraction unit is connected to the input terminal of the third normalization unit. The first encoding unit is configured to encode the second audio data information; The seventh feature extraction unit is configured to extract features from the data representation output by the first encoding unit; The third normalization unit is configured to normalize the data representation output by the seventh feature extraction unit; and / or, The fifth feature extraction module includes: an eighth feature extraction unit and a first activation unit; The input terminal of the eighth feature extraction unit is configured to receive the data representation output by the fourth feature extraction module, and the output terminal of the eighth feature extraction unit is connected to the input terminal of the first activation unit. The eighth feature extraction unit includes at least one first convolutional unit, and each first convolutional unit includes: a first convolutional layer, a first normalization layer, and a second normalization layer; Wherein, the output of the first convolutional layer is connected to the input of the first normalization layer, and the output of the first normalization layer is connected to the input of the second normalization layer; Optionally, the seventh feature extraction unit includes: a second convolutional unit and a third convolutional unit; Wherein, the output of the first encoding unit is connected to the input of the second convolution unit, and the output of the second convolution unit is connected to the input of the third convolution unit; The second convolutional unit includes: a second convolutional layer, a third normalization layer, and a first activation layer; Wherein, the output of the first encoding unit is connected to the input of the second convolutional layer, the output of the second convolutional layer is connected to the input of the third normalization layer, and the output of the third normalization layer is connected to the input of the first activation layer; The third convolutional unit includes: a third convolutional layer, a fourth normalization layer, and a second activation layer; Wherein, the output of the second convolutional unit is connected to the input of the third convolutional layer, the output of the third convolutional layer is connected to the input of the fourth normalization layer, and the output of the fourth normalization layer is connected to the input of the second activation layer; Optionally, the target feature information includes first feature information and / or second feature information, and the data determination module determines the target text information based on the target feature information, including: Based on the first feature information and the second feature information, the third feature information is determined; The second feature information and the third feature information are decoded to obtain the target text information; Optionally, the first feature information includes several feature information to be processed, and the data determination module determines the third feature information based on the first feature information and the second feature information, including: Each of the features to be processed and the second feature information are calculated and processed to obtain a first feature value corresponding to each feature to be processed; Based on the first feature value, several of the feature information to be processed are filtered to obtain the third feature information; Optionally, the data determination module performs calculations on each of the features to be processed and the second feature to obtain a first feature value corresponding to each of the features to be processed, including: Perform a transpose operation on each of the features to be processed to obtain the fourth feature information corresponding to each of the features to be processed. The fourth feature information and the second feature information are multiplied together to obtain the fifth feature information corresponding to each feature information to be processed. Divide the fifth feature information and the second feature value corresponding to the feature information to be processed to obtain the third feature value corresponding to each feature information to be processed. The third feature value is subjected to feature transformation processing to obtain the first feature value corresponding to each feature information to be processed; Optionally, after the data determination module determines the target text information based on the target feature information, the data determination module is further configured to: Obtain the target display location information; Based on the target display location information, the first text information corresponding to the target text information and the data information to be processed is displayed.

17. A device, characterized in that, The device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 15.

18. A computer-readable storage medium, characterized in that, It contains a computer program that is loaded by a processor to perform the steps of the method according to any one of claims 1 to 15.