Data processing system for AR head display device

By acquiring images through AR headsets and automatically translating target text using image recognition technology, the problem of translating unfamiliar languages ​​is solved, and translation efficiency is improved.

CN121963008APending Publication Date: 2026-05-01SHENZHEN TCL HIGH TECH DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TCL HIGH TECH DEVELOPMENT CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing language translation methods require users to manually input text, which is not suitable for translating unfamiliar languages.

Method used

Image information is acquired through AR headsets, and image recognition technology is used to automatically identify and translate target text, avoiding manual input.

Benefits of technology

It improves the efficiency of text translation, especially for translating unfamiliar languages, and simplifies the process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963008A_ABST
    Figure CN121963008A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing system for AR head-mounted display equipment, and the system determines target text information based on image recognition information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically to a data processing system for AR head-mounted display devices. Background Technology

[0002] In daily life, users often encounter scenarios that require translation between different languages. For example, when traveling abroad, users need to translate unfamiliar languages ​​on billboards, menus, and other advertisements. Existing language translation methods typically require users to manually input the text to be translated into the terminal device using a mouse and keyboard, and then the terminal device translates the input text.

[0003] However, the inability to use input methods such as a mouse and keyboard when translating unfamiliar languages ​​makes existing language translation methods unsuitable for translating unfamiliar languages. Summary of the Invention

[0004] This application provides a data processing system for AR head-mounted display devices.

[0005] In a first aspect, this application provides a method comprising:

[0006] Obtain the first image information;

[0007] The first image information is processed to obtain image recognition information;

[0008] The target text information is determined based on image recognition information.

[0009] Secondly, this application provides a system comprising:

[0010] The image acquisition module is used to acquire the first image information;

[0011] The image recognition module is used to process the first image information to obtain image recognition information;

[0012] The text determination module is used to determine the target text information based on image recognition information.

[0013] Thirdly, this application also provides a computer device, which includes:

[0014] One or more processors;

[0015] Memory; and

[0016] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the methods of any of the first aspects.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the method in any of the first aspects. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of a data processing system provided in an embodiment of the present invention;

[0020] Figure 2 This is a flowchart of one embodiment of the data processing method provided by the present invention;

[0021] Figure 3 This is a flowchart of a specific embodiment of the recognition and processing of first image information provided by the present invention;

[0022] Figure 4 This is a schematic diagram of the object detection model provided in an embodiment of the present invention;

[0023] Figure 5 This is a schematic diagram of the structure of the object recognition model provided in an embodiment of the present invention;

[0024] Figure 6 This is a schematic diagram of the structure of the third feature extraction unit provided in an embodiment of the present invention;

[0025] Figure 7 This is a schematic diagram of the structure of the text detection model provided in an embodiment of the present invention;

[0026] Figure 8 This is a schematic diagram illustrating a specific embodiment of the text detection module provided in this invention performing text detection processing;

[0027] Figure 9 This is a flowchart illustrating a specific embodiment of determining target text information provided in this invention.

[0028] Figure 10 This is a schematic diagram of a finger pointing gesture scenario provided in an embodiment of the present invention;

[0029] Figure 11 This is a schematic diagram of a scenario for a two-finger pointing gesture provided in an embodiment of the present invention;

[0030] Figure 12This is a schematic diagram of a scenario for a two-handed framing gesture provided in an embodiment of the present invention;

[0031] Figure 13 This is a schematic block diagram of the data processing system provided in the embodiments of the present invention;

[0032] Figure 14 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation

[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first," "second," "third," "fourth," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," "third," "fourth," etc., may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0035] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0036] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.

[0037] This application provides a data processing method and system for AR head-mounted display devices, which will be described in detail below.

[0038] Please see Figure 1 , Figure 1 This is a schematic diagram of a data processing system provided in an embodiment of this application. The data processing system may include a computer device 100, which integrates the data processing system, such as... Figure 1 Computer equipment in the country.

[0039] In this embodiment, the computer device 100 is mainly used to acquire first image information; perform recognition processing on the first image information to obtain image recognition information; and determine the target text information based on the image recognition information. This eliminates the need to manually input the text to be translated using a mouse and keyboard, thereby improving the efficiency of text translation.

[0040] In this embodiment, the computer device 100 can be a standalone server, a server network, or a server cluster. For example, the computer device 100 described in this embodiment includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0041] It is understood that the computer device 100 used in the embodiments of this application can be a device that includes both receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device may include: cellular or other communication devices having a single-line display, a multi-line display, or a cellular or other communication device without a multi-line display. Specifically, the computer device 100 may be a desktop terminal or a mobile terminal, and the computer device 100 may also be one of the following: Augmented Reality (AR) based wearable devices (e.g., AR glasses, AR helmets, AR goggles, etc.), mobile phones, tablets, laptops, etc.

[0042] Those skilled in the art will understand that Figure 1The application environment shown is merely one application scenario of the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include those that are more specific to this application. Figure 1 The number of computer devices shown is more or less, for example Figure 1 Only one computer device is shown in the diagram. It is understood that the data processing system may also include one or more other services, which are not limited here.

[0043] In addition, such as Figure 1 As shown, the data processing system may also include a memory 200 for storing data, such as image information, such as first image information, second image information, etc., and text information, such as target text information, candidate text information, etc.

[0044] It should be noted that, Figure 1 The schematic diagram of the data processing system shown is merely an example. The data processing system and scenario described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of data processing systems and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0045] First, this application provides a data processing method for an AR head-mounted display device. The execution subject of the data processing method is a data processing system, which is applied to a computer device. The data processing method includes: acquiring first image information; performing recognition processing on the first image information to obtain image recognition information; and determining target text information based on the image recognition information.

[0046] like Figure 2 The diagram shown is a flowchart of an embodiment of the data processing method in this application. The data processing method may include the following steps S201 to S203, as detailed below:

[0047] S201, Obtain the first image information.

[0048] Optionally, the first image information can be image information acquired through the imaging module configured in the computer device itself, image information captured by a high-definition camera, or image information acquired through the imaging module of other computer devices via networks, Bluetooth, infrared, etc. This embodiment of the invention does not impose limitations. For example, when the data processing method of this application is applied to a smartphone, the smartphone can directly acquire the first image information through its own configured imaging module. When the data processing method of this application is applied to a server, the server can acquire the first image information through the imaging module of the smartphone and obtain the first image information from the smartphone via networks, Bluetooth, infrared, etc.

[0049] For example, the data processing method of this application can be applied to wearable devices based on Augmented Reality (AR) (e.g., AR glasses, AR helmets, AR goggles, etc.). When a user needs to translate a piece of text, they can select the text to be translated through a specified gesture. After detecting the user's specified gesture, the AR-based wearable device acquires the first image information through its own configured imaging module. In this way, the user can obtain the translated text corresponding to the text to be translated through a specified gesture, without having to manually input the text to be translated using a mouse or keyboard, which can improve the efficiency of text translation.

[0050] For example, when the data processing method of this application is applied to a server, a user can select the text to be translated by specifying a gesture. After detecting the user's specified gesture, the AR-based wearable device acquires the first image information through its own configured imaging module and sends the first image information to the server through the network, Bluetooth, and infrared. The server processes the first image information to obtain the translated text corresponding to the text to be translated and returns the translated text to the AR-based wearable device through the network, Bluetooth, and infrared. In this case, the user can also obtain the translated text corresponding to the text to be translated by specifying a gesture, without having to manually input the text to be translated using a mouse and keyboard, which can improve the efficiency of text translation.

[0051] S202. The first image information is processed to obtain image recognition information.

[0052] In one specific embodiment, the image recognition information is information identified from the first image information, which includes objects and text. The object can be a hand or a physical object; this embodiment does not limit this. For example, the object can be a deformable object capable of selecting or bounding over text in the first image information. The image recognition information may include one or more of the following: the object's position information in the first image information, the category information corresponding to the object in the first image information, and the text region information corresponding to the text in the first image information. For example, when the object is a hand, the image recognition information may include the hand's position information in the first image information and the category information of the hand's gesture in the first image information. In one specific embodiment, the image recognition information includes the object's position information in the first image information. Based on the object's position information in the first image information, the text region to be translated is determined, and then the text region to be translated is translated. Compared to directly translating all text in the first image information, this improves the efficiency of text translation.

[0053] In another specific embodiment, the image recognition information includes category information corresponding to the object in the first image information and text region information corresponding to the text in the first image information. Based on the category information corresponding to the object in the first image information and the text region information corresponding to the text in the first image information, the text region to be translated is determined, and then the text region to be translated is translated. This can reduce information loss and make the determined text region to be translated more accurate, thereby further improving the efficiency of text translation.

[0054] In one specific implementation, the image recognition information includes first recognition information and second recognition information, as referred to... Figure 3 As shown, step S202 involves recognizing the first image information to obtain image recognition information, which may include the following steps S301 to S302, as detailed below:

[0055] S301. Perform object recognition on the first image information to obtain first recognition information.

[0056] In one specific embodiment, the first image information includes an object, which can be a human hand or an item; this embodiment does not limit this. For example, the object is a deformable object capable of selecting or bounding text in the first image information. The image recognition information includes first recognition information, which is the category information corresponding to the object in the first image information. For example, if the object is a human hand, the first recognition information can be pointing with one finger, pointing with two fingers, or bounding with both hands, etc.

[0057] In one specific embodiment, the step of performing object recognition on the first image information to obtain first recognition information specifically includes: performing object detection on the first image information to obtain object detection result information; cropping the first image information based on the object detection result information to obtain second image information; and performing object recognition on the second image information to obtain the first recognition information.

[0058] The object detection result information represents the position information of the object in the first image information. In this embodiment, the first image information is cropped based on the object detection result information to obtain the second image information, and then the object recognition is performed on the second image information. This can avoid the interference of redundant information in the first image information on object recognition and improve the accuracy of the obtained first recognition information.

[0059] In one specific embodiment, the object detection result information is obtained by performing object detection on the first image information using an object detection model, referring to... Figure 4 As shown, the object detection model includes a first feature extraction module, a second feature extraction module, and a first feature fusion module. The input terminals of the first and second feature extraction modules are configured to receive first image information; the output terminals of the first and second feature extraction modules are connected to the input terminal of the first feature fusion module.

[0060] Furthermore, the first feature extraction module is configured to extract features from the first image information to obtain first feature information; the second feature extraction module is configured to extract features from the first image information to obtain second feature information; and the first feature fusion module is configured to fuse the first feature information and the second feature information to obtain object detection result information.

[0061] In one specific embodiment, reference continues to be made to... Figure 4 As shown, the first feature extraction module includes: a first feature extraction unit, a first deconvolution unit, a second deconvolution unit, a first convolution unit, and a first normalization unit. The input of the first feature extraction unit is configured to receive first image information, and its output is connected to the input of the first deconvolution unit. The output of the first deconvolution unit is connected to the input of the second deconvolution unit, and the output of the second deconvolution unit is connected to the input of the first convolution unit. The output of the first convolution unit is connected to the input of the first normalization unit. In this embodiment, two deconvolution units are set after the first feature extraction unit, which can transform the features output by the first feature extraction unit into features of the same size as the features output by the second feature extraction unit.

[0062] The first feature extraction unit can be built based on ResNet-101 or VGG-16; this embodiment does not impose any limitations. The first deconvolution unit, the second deconvolution unit, and the first convolution unit can be built based on a 3×3 convolution kernel, a 5×5 convolution kernel, or a 7×7 convolution kernel; this embodiment does not impose any limitations.

[0063] In one specific embodiment, the second feature extraction module includes: a second feature extraction unit, a second convolution unit, and a second normalization unit. The input of the second feature extraction unit is configured to receive first image information, and the output of the second feature extraction unit is connected to the input of the second convolution unit; the output of the second convolution unit is connected to the input of the second normalization unit.

[0064] In one specific embodiment, the second feature extraction unit can be built based on ResNet-101 or VGG-16; this embodiment does not impose any limitations. The second convolutional unit can be built based on a 3×3 convolutional kernel, a 5×5 convolutional kernel, or a 7×7 convolutional kernel; this embodiment does not impose any limitations.

[0065] In one specific embodiment, the step of the first feature fusion module to fuse the first feature information and the second feature information specifically includes: performing element-wise multiplication on the first feature information and the second feature information to obtain object detection result information.

[0066] In one specific embodiment, the first identification information is obtained by performing object recognition on the second image information using an object recognition model, referring to... Figure 5 As shown, the object recognition model includes a third feature extraction module, a fully connected module, and an activation module. The input of the third feature extraction module is configured to receive second image information, and its output is connected to the input of the fully connected module; the output of the fully connected module is connected to the input of the activation module.

[0067] The activation module can be built based on the Softmax activation function, the Sigmoid activation function, or the ReLU activation function; this embodiment does not impose any limitations.

[0068] In one specific embodiment, reference continues to be made to... Figure 5 and Figure 6As shown, the third feature extraction module includes multiple third feature extraction units, each of which includes: a first convolutional layer, a second convolutional layer, and a pooling layer; the output of the first convolutional layer is connected to the input of the second convolutional layer, and the output of the second convolutional layer is connected to the input of the pooling layer.

[0069] The pooling layer can be constructed based on an average pooling layer or a max pooling layer; this embodiment of the invention is not limited to either. It should be noted that... Figure 5 The object recognition model shown includes three third feature extraction units. The actual object recognition model may also include more or fewer third feature extraction units. For example, the object recognition model may include two third feature extraction units, four third feature extraction units, five third feature extraction units, etc. This embodiment does not limit the scope.

[0070] In one specific embodiment, before cropping the first image information based on object detection result information to obtain the second image information, the process includes: filtering the object detection result information to obtain filtered object detection result information; and determining the filtered object detection result information as the new object detection result information. This embodiment utilizes an object detection model to perform object detection on the first image information, and then filters the object detection result information to improve the accuracy of the obtained object detection result information, thereby improving the accuracy of the obtained first recognition information. For example, the minimum bounding rectangle (MBR) can be used to filter the object detection result information to obtain filtered object detection result information.

[0071] S302. Perform region recognition on the first image information to obtain the second recognition information.

[0072] In this embodiment, the first image information includes text. Region recognition refers to identifying the region where the text is located in the first image information. A region represents a set of pixels in the first image information that are interconnected and have consistent attributes. The image recognition information also includes second recognition information, which characterizes the text region information corresponding to the text in the first image information. For example, the second recognition information includes the first text region information corresponding to a word in the first image information, the second text region information corresponding to a sentence in the first image information, and the third text region information corresponding to a paragraph in the first image information. This embodiment identifies the first image information to obtain the first and second recognition information, and then performs text translation based on the first and second recognition information. This can reduce information loss and make the determined text region to be translated more accurate, thereby further improving the efficiency of text translation.

[0073] In one specific embodiment, the step of performing region recognition on the first image information to obtain the second recognition information specifically includes: performing text detection on the first image information to obtain character position information; and dividing the first image information into text regions based on the character position information to obtain the second recognition information.

[0074] As mentioned in the preceding steps, the first image information includes text, which is composed of a series of characters, such as letters, numbers, punctuation marks, etc. The character position information is the position information of the character in the first image information. For example, the character position information includes the bounding box position information of the character in the first image information and the geometric center position information of the bounding box corresponding to the character in the first image information.

[0075] In one specific embodiment, the aforementioned character position information is obtained by performing text detection on the first image information using a text detection model, referring to... Figure 7 As shown, the text detection model includes: a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a second feature fusion module, and multiple text detection modules. The input of the fourth feature extraction module is configured to receive first image information; the output of the fourth feature extraction module is connected to the inputs of the fifth feature extraction module and the second feature fusion module; the output of the fifth feature extraction module is connected to the inputs of the sixth feature extraction module and the second feature fusion module; and the output of the second feature fusion module is connected to the inputs of the multiple text detection modules.

[0076] Furthermore, the fourth feature extraction module is configured to extract features from the first image information to obtain the third feature information; the fifth feature extraction module is configured to extract features from the third feature information to obtain the fourth feature information; the sixth feature extraction module is configured to extract features from the fourth feature information to obtain the fifth feature information; the second feature fusion module is configured to fuse the third, fourth, and fifth feature information to obtain multiple fused feature information; the multiple fused feature information corresponds one-to-one with multiple text detection modules; the multiple text detection modules are configured to perform text detection processing on the fused feature information corresponding to each text detection module to obtain character position information.

[0077] In one specific embodiment, the sizes of the third, fourth, and fifth feature information gradually decrease; that is, the fourth feature information is obtained by downsampling the third feature information, and the fifth feature information is obtained by downsampling the fourth feature information. The text detection model can be built based on a Feature Pyramid Network (FPN) model. The fourth, fifth, and sixth feature extraction modules constitute the backbone network of the text detection model. The backbone network of the text detection model can be built based on a ResNet model, or it can be built based on an EfficientNet model; this embodiment does not impose any limitations.

[0078] In one specific embodiment, reference continues to be made to... Figure 7 As shown, the second feature fusion module includes: a fourth feature extraction unit, a first feature fusion unit, a second feature fusion unit, a fifth feature extraction unit, and a sixth feature extraction unit. The input of the fourth feature extraction unit is connected to the output of the fourth feature extraction module; the outputs of the fourth and fifth feature extraction modules are respectively connected to the input of the first feature fusion unit; the outputs of the first and sixth feature extraction modules are respectively connected to the input of the second feature fusion unit; the output of the second feature fusion unit is connected to the input of the fifth feature extraction unit; and the output of the fifth feature extraction unit is connected to the input of the sixth feature extraction unit.

[0079] In one specific embodiment, the size of the fused feature information output by the fourth feature extraction unit, the first feature fusion unit, the second feature fusion unit, the fifth feature extraction unit, and the sixth feature extraction unit decreases in that order. By inputting multiple fused feature information of different sizes into multiple text detection modules, text detection can be performed using multiple fused feature information of different sizes, thereby improving the accuracy of the obtained character position information.

[0080] In one specific embodiment, reference is made to Figure 8 As shown, each text detection module includes a first branch and a second branch. The first branch is used to classify the text in the fused image information and predict the geometric center position of the bounding box corresponding to the character in the first image information. The second branch is used to perform regression processing on the fused image information to obtain the bounding box position information of the character in the first image information.

[0081] Furthermore, the first branch includes multiple third convolutional layers, a classification unit, and a location prediction unit. The outputs of the multiple third convolutional layers are connected to the inputs of the classification unit and the location prediction unit. These multiple third convolutional layers can be constructed using 3×3 convolutional kernels, 5×5 convolutional kernels, or 7×7 convolutional kernels; this embodiment does not impose any limitations. For example, continuing to refer to… Figure 8 As shown, the first branch includes four third convolutional layers. The fused feature information is passed through the four third convolutional layers in sequence to obtain the first convolutional feature map of H*W*256. The first convolutional feature map is input into the classification unit and the location prediction unit respectively to obtain the geometric center position information of the bounding box of the character in the first image information.

[0082] In one specific embodiment, the second branch includes multiple fourth convolutional layers and a regression processing unit, with the outputs of the multiple fourth convolutional layers connected to the inputs of the regression processing unit. These multiple fourth convolutional layers can be constructed using 3×3 convolutional kernels, 5×5 convolutional kernels, or 7×7 convolutional kernels; this embodiment does not impose any limitation. For example, continuing to refer to... Figure 8 As shown, the second branch includes four fourth convolutional layers. The fused feature information is sequentially passed through the four fourth convolutional layers to obtain a second convolutional feature map of H*W*256. The second convolutional feature map is input into the regression processing unit for regression processing to obtain the bounding box position information of the character in the first image information.

[0083] It should be noted that when performing text detection processing through multiple text detection modules, a certain location point in the first image information may belong to multiple bounding boxes at the same time. In this case, the bounding box with the smallest area can be selected as the bounding box to which the location point belongs.

[0084] In one specific embodiment, the second identification information includes at least one of the first text region information, the second text region information, and the third text region information. The text in the first image information includes several paragraphs, each paragraph consists of several sentences, and each sentence consists of several words. The first text region information represents the position information of each word corresponding to the text in the first image information, the second text region information represents the position information of each sentence corresponding to the text in the first image information, and the third text region information represents the position information of each paragraph corresponding to the text in the first image information.

[0085] Optionally, the above-mentioned step of dividing the first image information into text regions based on character position information to obtain the second recognition information specifically includes: determining the first text region information based on character position information; and / or performing clustering processing on the first text region information to obtain the second text region information; and / or performing fusion processing on the first text region information to obtain the third text region information.

[0086] In one specific embodiment, the step of determining the first text region information based on character position information specifically includes: calculating and processing the character position information to obtain multiple first distance information; performing clustering processing on the multiple first distance information to obtain first target clustering information; and performing fusion processing on the character position information corresponding to the first distance information in the first target clustering information to obtain the first text region information.

[0087] In this embodiment, the character position information includes the geometric center position information corresponding to the bounding box of the character in the first image information, and the first distance information is the distance between the geometric centers corresponding to the bounding boxes of the characters in the first image information, which can be determined based on the geometric center position information corresponding to the bounding boxes of the characters in the first image information. For example, the geometric center position information corresponding to the bounding box of character A in the first image information is (x a1 ,y a1 The geometric center position information of the bounding box corresponding to character B in the first image information is (x b1 ,y b1 Then, the first distance information determined by characters A and B can be represented as:

[0088] Furthermore, clustering multiple first distance information sets allows for the grouping of first distance information sets with similar distance values ​​into the same cluster, resulting in multiple first clusters. The first target cluster is the cluster with the most first distance information sets among these multiple first clusters. For example, clustering multiple first distance information sets yields first cluster A, first cluster B, and first cluster C. The order of the number of first distance information sets in first cluster A, first cluster B, and first cluster C from most to least is: first cluster B > first cluster A > first cluster C. Therefore, first cluster B is the first target cluster.

[0089] It should be noted that multiple first distance information can be clustered based on HAC clustering, K-means clustering, or K-Medians clustering. This embodiment does not limit the specific clustering method.

[0090] In one specific embodiment, the character position information further includes the bounding box position information of the character in the first image information. The first target clustering information includes several first distance information, each first distance information corresponding to two characters. The fusion processing of the character position information corresponding to the first distance information in the first target clustering information refers to merging the bounding box position information of the two characters corresponding to each first distance information in the first target clustering information. For example, the first target clustering information includes first distance information d... 1a First distance information d 1a For characters A and B, the first distance information d 1a The corresponding character position information fusion processing refers to merging the bounding box position information of character A and the bounding box position information of character B.

[0091] In one specific embodiment, the step of clustering the first text region information to obtain the second text region information specifically includes: calculating and processing the first text region information to obtain multiple second distance information; clustering the multiple second distance information to obtain second target clustering information; and fusing the first text region information corresponding to the second distance information in the second target clustering information to obtain the second text region information.

[0092] In this embodiment, the first text region information includes the geometric center position information corresponding to the bounding box of a word in the first image information, and the second distance information is the distance between the geometric centers corresponding to the bounding boxes of words in the first image information, which can be determined based on the geometric center position information corresponding to the bounding boxes of words in the first image information. For example, the geometric center position information corresponding to the bounding box of word A in the first image information is (x a2 ,y a2 ), the geometric center location information (x) of the bounding box corresponding to word B in the first image information. b2 ,y b2 Then, the second distance information determined by word A and word B can be represented as:

[0093] Furthermore, clustering multiple second distance information sets allows for the grouping of second distance information sets with similar distance values ​​into the same cluster, resulting in multiple second clusters. The second target cluster is the cluster with the most second distance information sets among these multiple clusters. For example, clustering multiple second distance information sets can yield second cluster A, second cluster B, and second cluster C. The order of the number of second distance information sets in second cluster A, second cluster B, and second cluster C from most to least is second cluster B > second cluster A > second cluster C. Therefore, second cluster B is the second target cluster.

[0094] It should be noted that multiple second distance information can be clustered based on HAC clustering, K-means clustering, or K-Medians clustering. This embodiment does not limit the specific clustering method.

[0095] In one specific embodiment, the first text region information further includes the bounding box position information of words in the first image information. The second target clustering information includes several second distance information, each second distance information corresponding to two words. The fusion processing of the first text region information corresponding to the second distance information in the second target clustering information refers to merging the bounding box position information of the two words corresponding to each second distance information in the second target clustering information. For example, the second target clustering information includes second distance information d. 1a Second distance information d 1a For words A and B, the second distance information d 1a The corresponding first text region information fusion processing refers to merging the bounding box position information of word A and the bounding box position information of word B.

[0096] In one specific embodiment, the step of fusing the first text region information to obtain the third text region information specifically includes: adjusting the text region corresponding to the first text region information to obtain the fourth text region information; and fusing the fourth text region information based on the overlapping area of ​​the text regions corresponding to the fourth text region information to obtain the third text region information.

[0097] Optionally, adjusting the text region corresponding to the first text region information means expanding the text region corresponding to the first text region information according to a preset ratio. By adjusting the text region corresponding to the first text region information, the blank spaces between words in a line and between lines in a paragraph can be filled according to a preset ratio, thereby improving the accuracy of recognizing the third text region information.

[0098] In one specific embodiment, the step of fusing the fourth text region information based on the overlapping area of ​​the text regions corresponding to the fourth text region information specifically includes: for any two fourth text region information, calculating the overlapping area of ​​the text regions corresponding to the two fourth text region information; comparing the overlapping area of ​​the text regions corresponding to the two fourth text region information with a preset area threshold; if the overlapping area of ​​the text regions corresponding to the two fourth text region information is greater than the area threshold, merging the two fourth text region information.

[0099] S203. Based on image recognition information, determine the target text information.

[0100] In this embodiment, the target text information is the translated text information corresponding to the text to be translated in the first image information. In this embodiment, the first image information is processed to obtain image recognition information, and the target text information is determined based on the image recognition information. The text to be translated in the first image information can be automatically translated, thereby improving the efficiency of text translation.

[0101] In one specific implementation, the image recognition information includes first recognition information and second recognition information, as referred to... Figure 9 As shown, the determination of target text information based on image recognition information in step S203 above may include the following steps S401 to S402, as detailed below:

[0102] S401. Based on the first identification information and the second identification information, determine the target text region information.

[0103] In this embodiment, the first identification information is the category information corresponding to the object in the first image information, the second identification information is the text region information corresponding to the text in the first image information, and the target text region information is the region information where the text to be translated is located in the first image information. This embodiment determines the target text region information based on the category information corresponding to the object in the first image information and the text region information corresponding to the text in the first image information, which can reduce information loss and make the determined target text region information more accurate, thereby further improving the efficiency of text translation and the accuracy of the obtained target text information.

[0104] In one specific embodiment, the second identification information includes at least one of first text region information, second text region information, and third text region information. The step of determining the target text region information based on the first and second identification information specifically includes: if the object category corresponding to the first identification information is a first category, filtering the first text region information based on the third distance information between the first reference point in the first image information and the text region corresponding to the first text region information to obtain the target text region information; and / or, if the object category corresponding to the first identification information is a second category, filtering the second text region information based on the fourth distance information between the second reference point in the first image information and the text region corresponding to the second text region information to obtain the target text region information; and / or, if the object category corresponding to the first identification information is a third category, filtering the third text region information based on the similarity information between the reference region in the first image information and the text region corresponding to the third text region information to obtain the target text region information.

[0105] In one specific embodiment, reference is made to Figure 10 As shown, the first category is finger pointing, and the first reference point is the position point corresponding to the fingertip in the first image information. For example, the first reference point is... Figure 10 The intersection of the crosshair cursor in the image. The third distance information is the distance between the position point corresponding to the fingertip in the first image information and the text area corresponding to the first text area information. When filtering the first text area information based on the third distance information between the first reference point in the first image information and the text area corresponding to the first text area information, the first text area information with the smallest distance value corresponding to the third distance information is determined as the target text area information.

[0106] Furthermore, referring to Figure 11 As shown, the second category is two-finger pointing, and the second reference point is the midpoint of the line segment formed by the two fingertips in the first image information. For example, the second reference point is... Figure 11 The intersection of the crosshair cursor in the image. The fourth distance information is the distance between the midpoint of the line segment formed by the two fingertips in the first image information and the text region corresponding to the second text region information. When filtering the second text region information based on the fourth distance information between the second reference point in the first image information and the text region corresponding to the second text region information, the second text region information with the smallest distance value corresponding to the fourth distance information is determined as the target text region information.

[0107] In one specific embodiment, reference is made to Figure 12 As shown, the third category is defined by two hands, with the reference area being the region enclosed by both hands in the first image information. For example, the reference area is... Figure 12The similarity information is the intersection-union ratio (IOU) of the base region and the text regions corresponding to the third text region, which is enclosed by four fingers. The determination of the similarity information can be expressed as: IOU = A∩B / A∪B, where IOU represents the similarity information, A represents the area of ​​the base region, B represents the area of ​​the text region corresponding to the third text region, A∩B represents the intersection area of ​​the base region and the text regions corresponding to the third text region, and A∪B represents the union area of ​​the base region and the text regions corresponding to the third text region.

[0108] Furthermore, when filtering the third text region information based on the similarity information between the reference region in the first image information and the text region corresponding to the third text region information, the third text region information with the highest similarity can be determined as the target text region information.

[0109] S402. Based on the target text region information, determine the target text information.

[0110] This embodiment determines the target text region information based on the first identification information and the second identification information, and then determines the target text information based on the target text region information, which can improve the efficiency of text translation and the accuracy of the obtained target text information.

[0111] In one specific embodiment, the step of determining the target text information based on the target text region information specifically includes: extracting text information from the target text region corresponding to the target text region information to obtain candidate text information; and translating the candidate text information to obtain the target text information.

[0112] In this embodiment, the candidate text information is the text information directly extracted from the first image information based on the target text region information; that is, the candidate text information is the text information that the user needs to translate. When translating the candidate text information, it can be input into an offline translator to obtain the translated target text information.

[0113] Optionally, text information can be extracted from the target text region based on Optical Character Recognition (OCR) technology. OCR technology can not only achieve rapid extraction of text information, but also support the extraction of text information in multiple languages.

[0114] In one specific embodiment, considering that the determined target text region information may be inaccurate, when the object category corresponding to the first identification information is the second category and the target text information includes multiple sentences, the fifth distance information between the geometric center of each sentence and the second reference point can be calculated, and the sentence with the smallest distance value corresponding to the fifth distance information can be taken as the final target text information, thereby improving the accuracy of the obtained target text information.

[0115] In one specific embodiment, after determining the target text information based on image recognition information in step S203, the process may include outputting the target text information. For example, when the data processing method of this application is applied to an AR-based wearable device, the AR-based wearable device can display the target text information on an AR interface, and the user can obtain the translated target text information through the content displayed on the AR interface.

[0116] In summary, the data processing method provided in this implementation scheme acquires first image information, performs recognition processing on the first image information to obtain image recognition information, and determines the target text information based on the image recognition information. This scheme obtains image recognition information by recognizing the first image information and then determines the target text information based on the image recognition information, eliminating the need for manual input of the text to be translated using a mouse and keyboard, thus improving the efficiency of text translation. Furthermore, object recognition is performed on the first image information to obtain first recognition information, and region recognition is performed on the first image information to obtain second recognition information. The target text region information is determined based on the first and second recognition information, which reduces information loss and makes the determined target text region information more accurate, thereby further improving the efficiency of text translation and the accuracy of the obtained target text information. Moreover, compared with existing text detectors, the text detection model used to detect the character position information in the first image information has higher accuracy and robustness for long-distance target detection, and has lower latency and computational complexity, making it convenient for use on AR devices.

[0117] To better implement the data processing method in the embodiments of this application, a data processing system is also provided in the embodiments of this application, such as... Figure 13 As shown, the data processing system 600 includes:

[0118] Image acquisition module 610 is used to acquire first image information;

[0119] Image recognition module 620 is used to process the first image information to obtain image recognition information;

[0120] The text determination module 630 is used to determine the target text information based on image recognition information.

[0121] In this embodiment, image recognition information is obtained by recognizing the first image information, and then the target text information is determined based on the image recognition information. This eliminates the need to manually input the text to be translated using a mouse and keyboard, thus improving the efficiency of text translation.

[0122] In some embodiments of this application, the image recognition information includes first recognition information and second recognition information. The image recognition module 620 performs recognition processing on the first image information to obtain image recognition information, including:

[0123] Object recognition is performed on the first image information to obtain first recognition information;

[0124] Region recognition is performed on the first image information to obtain the second recognition information.

[0125] In some embodiments of this application, the image recognition module 620 performs object recognition on the first image information to obtain first recognition information, including:

[0126] Object detection is performed on the first image information to obtain object detection result information;

[0127] The first image information is cropped based on the object detection results to obtain the second image information;

[0128] Object recognition is performed on the second image information to obtain the first recognition information.

[0129] In some embodiments of this application, the object detection result information is obtained by performing object detection on the first image information using an object detection model. The object detection model includes: a first feature extraction module, a second feature extraction module, and a first feature fusion module.

[0130] The input terminals of the first feature extraction module and the second feature extraction module are respectively configured to receive first image information; the output terminals of the first feature extraction module and the second feature extraction module are respectively connected to the input terminal of the first feature fusion module.

[0131] The first feature extraction module is configured to extract features from the first image information to obtain first feature information;

[0132] The second feature extraction module is configured to extract features from the first image information to obtain second feature information;

[0133] The first feature fusion module is configured to fuse the first feature information and the second feature information to obtain object detection result information.

[0134] In some embodiments of this application, the first feature extraction module includes: a first feature extraction unit, a first deconvolution unit, a second deconvolution unit, a first convolution unit, and a first normalization unit; the second feature extraction module includes: a second feature extraction unit, a second convolution unit, and a second normalization unit.

[0135] The input of the first feature extraction unit is configured to receive first image information, and the output of the first feature extraction unit is connected to the input of the first deconvolution unit; the output of the first deconvolution unit is connected to the input of the second deconvolution unit, and the output of the second deconvolution unit is connected to the input of the first convolution unit; the output of the first convolution unit is connected to the input of the first normalization unit.

[0136] The input of the second feature extraction unit is configured to receive the first image information, and the output of the second feature extraction unit is connected to the input of the second convolution unit; the output of the second convolution unit is connected to the input of the second normalization unit.

[0137] In some embodiments of this application, the first identification information is obtained by performing object recognition on the second image information through an object recognition model. The object recognition model includes: a third feature extraction module, a fully connected module, and an activation module.

[0138] The input of the third feature extraction module is configured to receive the second image information, and the output of the third feature extraction module is connected to the input of the fully connected module; the output of the fully connected module is connected to the input of the activation module.

[0139] The third feature extraction module includes multiple third feature extraction units, each of which includes: a first convolutional layer, a second convolutional layer, and a pooling layer; the output of the first convolutional layer is connected to the input of the second convolutional layer, and the output of the second convolutional layer is connected to the input of the pooling layer.

[0140] In some embodiments of this application, the image recognition module 620 performs region recognition on the first image information to obtain second recognition information, including:

[0141] Text detection is performed on the first image information to obtain character position information;

[0142] The first image information is divided into text regions based on character position information to obtain the second recognition information.

[0143] In some embodiments of this application, character position information is obtained by performing text detection on the first image information using a text detection model. The text detection model includes: a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a second feature fusion module, and multiple text detection modules.

[0144] The fourth feature extraction module is configured to receive first image information at its input end, and its output end is connected to the input end of the fifth feature extraction module and the input end of the second feature fusion module. The output end of the fifth feature extraction module is connected to the input end of the sixth feature extraction module and the input end of the second feature fusion module. The output end of the second feature fusion module is connected to the input ends of multiple text detection modules.

[0145] The fourth feature extraction module is configured to extract features from the first image information to obtain the third feature information;

[0146] The fifth feature extraction module is configured to extract features from the third feature information to obtain the fourth feature information;

[0147] The sixth feature extraction module is configured to extract features from the fourth feature information to obtain the fifth feature information;

[0148] The second feature fusion module is configured to fuse the third, fourth, and fifth feature information to obtain multiple fused feature information; each of the multiple fused feature information corresponds one-to-one with a multiple text detection module.

[0149] Multiple text detection modules are configured to perform text detection processing on the fused feature information corresponding to each text detection module to obtain character position information.

[0150] In some embodiments of this application, the second feature fusion module includes: a fourth feature extraction unit, a first feature fusion unit, a second feature fusion unit, a fifth feature extraction unit, and a sixth feature extraction unit;

[0151] The input of the fourth feature extraction unit is connected to the output of the fourth feature extraction module. The outputs of the fourth and fifth feature extraction modules are connected to the inputs of the first feature fusion unit. The outputs of the first and sixth feature extraction modules are connected to the inputs of the second feature fusion unit. The output of the second feature fusion unit is connected to the input of the fifth feature extraction unit. The output of the fifth feature extraction unit is connected to the input of the sixth feature extraction unit.

[0152] In some embodiments of this application, the second identification information includes at least one of first text region information, second text region information, and third text region information. The image recognition module 620 divides the first image information into text regions based on character position information to obtain the second identification information, including:

[0153] Based on character position information, determine the first text region information; and / or,

[0154] Clustering is performed on the information from the first text region to obtain the information from the second text region; and / or,

[0155] The information in the first text region is fused to obtain the information in the third text region.

[0156] In some embodiments of this application, the image recognition module 620 determines first text region information based on character position information, including:

[0157] The character position information is calculated and processed to obtain multiple first distance information;

[0158] Clustering is performed on multiple first distance information to obtain first target cluster information;

[0159] The character position information corresponding to the first distance information in the first target cluster information is fused to obtain the first text region information.

[0160] In some embodiments of this application, the image recognition module 620 performs clustering processing on the first text region information to obtain second text region information, including:

[0161] The information in the first text region is processed to obtain multiple second distance information;

[0162] Clustering is performed on multiple second distance information to obtain second target clustering information;

[0163] The first text region information corresponding to the second distance information in the second target clustering information is fused to obtain the second text region information.

[0164] In some embodiments of this application, the image recognition module 620 performs fusion processing on the first text region information to obtain third text region information, including:

[0165] The text region corresponding to the first text region information is adjusted to obtain the fourth text region information;

[0166] The fourth text region information is fused based on the overlapping area of ​​the text regions corresponding to the fourth text region information to obtain the third text region information.

[0167] In some embodiments of this application, the text determination module 630 determines the target text information based on image recognition information, including:

[0168] Based on the first and second identification information, the target text region information is determined;

[0169] The target text information is determined based on the target text region information.

[0170] In some embodiments of this application, the text determination module 630 determines target text region information based on first identification information and second identification information, including:

[0171] If the object category corresponding to the first identification information is the first category, the first text region information is filtered based on the third distance information between the first reference point in the first image information and the text region corresponding to the first text region information to obtain the target text region information; and / or,

[0172] If the object category corresponding to the first identification information is the second category, the second text region information is filtered based on the fourth distance information between the second reference point in the first image information and the text region corresponding to the second text region information to obtain the target text region information; and / or,

[0173] If the object category corresponding to the first identification information is the third category, the third text region information is filtered based on the similarity information between the reference region in the first image information and the text region corresponding to the third text region information to obtain the target text region information.

[0174] In some embodiments of this application, the text determination module 630 determines target text information based on target text region information, including:

[0175] Extract text information from the target text region corresponding to the target text region information to obtain candidate text information;

[0176] The candidate text information is translated to obtain the target text information.

[0177] This application also provides a computer device that integrates any of the data processing systems provided in this application. The computer device includes:

[0178] One or more processors;

[0179] Memory; and

[0180] One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the data processing method in any of the embodiments described above.

[0181] This application also provides a computer device that integrates any of the data processing systems provided in this application. For example... Figure 14 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically:

[0182] The computer device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that... Figure 14 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0183] The processor 801 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 801.

[0184] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.

[0185] The computer device also includes a power supply 803 that supplies power to the various components. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0186] The computer device may also include an input unit 804, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0187] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 runs the application programs stored in the memory 802 to realize various functions, as follows:

[0188] Obtain the first image information;

[0189] The first image information is processed to obtain image recognition information;

[0190] The target text information is determined based on image recognition information.

[0191] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0192] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the data processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps:

[0193] Obtain the first image information;

[0194] The first image information is processed to obtain image recognition information;

[0195] The target text information is determined based on image recognition information.

[0196] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0197] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0198] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0199] The above provides a detailed description of a data processing method and system for an AR head-mounted display device provided by the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method, characterized in that, include: Obtain the first image information; The first image information is processed to obtain image recognition information; Based on the image recognition information, the target text information is determined.

2. The method according to claim 1, characterized in that, The image recognition information includes first recognition information and second recognition information; The step of performing recognition processing on the first image information to obtain image recognition information includes: Object recognition is performed on the first image information to obtain the first recognition information; The first image information is used to perform region recognition to obtain the second recognition information.

3. The method according to claim 2, characterized in that, The step of performing object recognition on the first image information to obtain the first recognition information includes: Object detection is performed on the first image information to obtain object detection result information; Based on the object detection result information, the first image information is cropped to obtain the second image information; The second image information is used to perform object recognition to obtain the first recognition information.

4. The method according to claim 3, characterized in that, The object detection result information is obtained by performing object detection on the first image information through an object detection model. The object detection model includes: a first feature extraction module, a second feature extraction module, and a first feature fusion module. The input terminals of the first feature extraction module and the second feature extraction module are respectively configured to receive the first image information; the output terminals of the first feature extraction module and the second feature extraction module are respectively connected to the input terminal of the first feature fusion module. The first feature extraction module is configured to extract features from the first image information to obtain first feature information; The second feature extraction module is configured to extract features from the first image information to obtain second feature information; The first feature fusion module is configured to fuse the first feature information and the second feature information to obtain the object detection result information.

5. The method according to claim 4, characterized in that, The first feature extraction module includes: a first feature extraction unit, a first deconvolution unit, a second deconvolution unit, a first convolution unit, and a first normalization unit; the second feature extraction module includes: a second feature extraction unit, a second convolution unit, and a second normalization unit. Wherein, the input end of the first feature extraction unit is configured to receive the first image information, and the output end of the first feature extraction unit is connected to the input end of the first deconvolution unit; the output end of the first deconvolution unit is connected to the input end of the second deconvolution unit, and the output end of the second deconvolution unit is connected to the input end of the first convolution unit; the output end of the first convolution unit is connected to the input end of the first normalization unit. The input of the second feature extraction unit is configured to receive the first image information, and the output of the second feature extraction unit is connected to the input of the second convolution unit; the output of the second convolution unit is connected to the input of the second normalization unit.

6. The method according to claim 3, characterized in that, The first identification information is obtained by performing object recognition on the second image information through an object recognition model. The object recognition model includes: a third feature extraction module, a fully connected module, and an activation module. The input of the third feature extraction module is configured to receive the second image information, and the output of the third feature extraction module is connected to the input of the fully connected module; the output of the fully connected module is connected to the input of the activation module. The third feature extraction module includes multiple third feature extraction units, each of which includes a first convolutional layer, a second convolutional layer, and a pooling layer; the output of the first convolutional layer is connected to the input of the second convolutional layer, and the output of the second convolutional layer is connected to the input of the pooling layer.

7. The method according to claim 2, characterized in that, The step of performing region recognition on the first image information to obtain the second recognition information includes: Text detection is performed on the first image information to obtain character position information; The first image information is divided into text regions based on the character position information to obtain the second recognition information.

8. The method according to claim 7, characterized in that, The character position information is obtained by performing text detection on the first image information using a text detection model. The text detection model includes: a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a second feature fusion module, and multiple text detection modules. The fourth feature extraction module has its input terminal configured to receive the first image information, and its output terminal connected to the input terminal of the fifth feature extraction module and the input terminal of the second feature fusion module; the output terminal of the fifth feature extraction module is connected to the input terminal of the sixth feature extraction module and the input terminal of the second feature fusion module; and the output terminal of the second feature fusion module is connected to the input terminals of multiple text detection modules. The fourth feature extraction module is configured to extract features from the first image information to obtain third feature information; The fifth feature extraction module is configured to extract features from the third feature information to obtain the fourth feature information; The sixth feature extraction module is configured to extract features from the fourth feature information to obtain the fifth feature information; The second feature fusion module is configured to fuse the third feature information, the fourth feature information, and the fifth feature information to obtain multiple fused feature information; the multiple fused feature information correspond one-to-one with the multiple text detection modules respectively; The multiple text detection modules are configured to perform text detection processing on the fused feature information corresponding to each text detection module to obtain the character position information.

9. The method according to claim 8, characterized in that, The second feature fusion module includes: a fourth feature extraction unit, a first feature fusion unit, a second feature fusion unit, a fifth feature extraction unit, and a sixth feature extraction unit; The input terminal of the fourth feature extraction unit is connected to the output terminal of the fourth feature extraction module; the output terminals of the fourth feature extraction unit and the fifth feature extraction module are respectively connected to the input terminal of the first feature fusion unit; the output terminals of the first feature fusion unit and the sixth feature extraction module are respectively connected to the input terminal of the second feature fusion unit; the output terminal of the second feature fusion unit is connected to the input terminal of the fifth feature extraction unit; and the output terminal of the fifth feature extraction unit is connected to the input terminal of the sixth feature extraction unit.

10. The method according to claim 7, characterized in that, The second identification information includes at least one of the first text region information, the second text region information, and the third text region information; The step of dividing the first image information into text regions based on the character position information to obtain the second recognition information includes: Based on the character position information, determine the first text region information; and / or, Clustering is performed on the first text region information to obtain the second text region information; and / or, The first text region information is fused to obtain the third text region information.

11. The method according to claim 10, characterized in that, Determining the first text region information based on the character position information includes: The character position information is calculated and processed to obtain multiple first distance information; Clustering is performed on multiple sets of the first distance information to obtain the first target cluster information; The character position information corresponding to the first distance information in the first target clustering information is fused to obtain the first text region information.

12. The method according to claim 10, characterized in that, The step of clustering the first text region information to obtain the second text region information includes: The information of the first text region is processed to obtain multiple second distance information; Clustering is performed on multiple sets of the second distance information to obtain the second target cluster information; The first text region information corresponding to the second distance information in the second target clustering information is fused to obtain the second text region information.

13. The method according to claim 10, characterized in that, The process of fusing the first text region information to obtain the third text region information includes: The text region corresponding to the first text region information is adjusted to obtain the fourth text region information. The fourth text region information is fused based on the overlapping area of ​​the text regions corresponding to the fourth text region information to obtain the third text region information.

14. The method according to claim 1, characterized in that, The image recognition information includes first recognition information and second recognition information. The step of determining the target text information based on the image recognition information includes: Based on the first identification information and the second identification information, the target text region information is determined; Based on the target text region information, the target text information is determined.

15. The method according to claim 14, characterized in that, The second identification information includes at least one of the first text region information, the second text region information, and the third text region information; The step of determining the target text region information based on the first identification information and the second identification information includes: If the object category corresponding to the first identification information is the first category, the first text region information is filtered based on the third distance information between the first reference point in the first image information and the text region corresponding to the first text region information to obtain the target text region information; and / or, If the object category corresponding to the first identification information is the second category, the second text region information is filtered based on the fourth distance information between the second reference point in the first image information and the text region corresponding to the second text region information to obtain the target text region information; and / or, If the object category corresponding to the first identification information is the third category, the third text region information is filtered based on the similarity information between the reference region in the first image information and the text region corresponding to the third text region information to obtain the target text region information.

16. The method according to claim 14, characterized in that, The step of determining the target text information based on the target text region information includes: Text information is extracted from the target text region corresponding to the target text region information to obtain candidate text information; The candidate text information is translated to obtain the target text information.

17. A system, characterized in that, include: The image acquisition module is used to acquire the first image information; The image recognition module is used to perform recognition processing on the first image information to obtain image recognition information; The text determination module is used to determine the target text information based on the image recognition information; Optionally, the image recognition information includes first recognition information and second recognition information. The image recognition module performs recognition processing on the first image information to obtain image recognition information, including: Object recognition is performed on the first image information to obtain the first recognition information; The first image information is used to perform region recognition to obtain the second recognition information; Optionally, the image recognition module performs object recognition on the first image information to obtain the first recognition information, including: Object detection is performed on the first image information to obtain object detection result information; Based on the object detection result information, the first image information is cropped to obtain the second image information; Perform object recognition on the second image information to obtain the first recognition information; Optionally, the object detection result information is obtained by performing object detection on the first image information through an object detection model, and the object detection model includes: a first feature extraction module, a second feature extraction module, and a first feature fusion module; The input terminals of the first feature extraction module and the second feature extraction module are respectively configured to receive the first image information; the output terminals of the first feature extraction module and the second feature extraction module are respectively connected to the input terminal of the first feature fusion module. The first feature extraction module is configured to extract features from the first image information to obtain first feature information; The second feature extraction module is configured to extract features from the first image information to obtain second feature information; The first feature fusion module is configured to fuse the first feature information and the second feature information to obtain the object detection result information; Optionally, the first feature extraction module includes: a first feature extraction unit, a first deconvolution unit, a second deconvolution unit, a first convolution unit, and a first normalization unit; the second feature extraction module includes: a second feature extraction unit, a second convolution unit, and a second normalization unit. Wherein, the input end of the first feature extraction unit is configured to receive the first image information, and the output end of the first feature extraction unit is connected to the input end of the first deconvolution unit; the output end of the first deconvolution unit is connected to the input end of the second deconvolution unit, and the output end of the second deconvolution unit is connected to the input end of the first convolution unit; the output end of the first convolution unit is connected to the input end of the first normalization unit. The input of the second feature extraction unit is configured to receive the first image information, and the output of the second feature extraction unit is connected to the input of the second convolution unit; the output of the second convolution unit is connected to the input of the second normalization unit. Optionally, the first identification information is obtained by performing object recognition on the second image information through an object recognition model, wherein the object recognition model includes: a third feature extraction module, a fully connected module, and an activation module; The input of the third feature extraction module is configured to receive the second image information, and the output of the third feature extraction module is connected to the input of the fully connected module; the output of the fully connected module is connected to the input of the activation module. The third feature extraction module includes multiple third feature extraction units, each of which includes: a first convolutional layer, a second convolutional layer, and a pooling layer; the output of the first convolutional layer is connected to the input of the second convolutional layer, and the output of the second convolutional layer is connected to the input of the pooling layer. Optionally, the image recognition module performs region recognition on the first image information to obtain the second recognition information, including: Text detection is performed on the first image information to obtain character position information; Based on the character position information, the first image information is divided into text regions to obtain the second recognition information; Optionally, the character position information is obtained by performing text detection on the first image information using a text detection model. The text detection model includes: a fourth feature extraction module, a fifth feature extraction module, a sixth feature extraction module, a second feature fusion module, and multiple text detection modules. The fourth feature extraction module has its input terminal configured to receive the first image information, and its output terminal connected to the input terminal of the fifth feature extraction module and the input terminal of the second feature fusion module; the output terminal of the fifth feature extraction module is connected to the input terminal of the sixth feature extraction module and the input terminal of the second feature fusion module; and the output terminal of the second feature fusion module is connected to the input terminals of multiple text detection modules. The fourth feature extraction module is configured to extract features from the first image information to obtain third feature information; The fifth feature extraction module is configured to extract features from the third feature information to obtain the fourth feature information; The sixth feature extraction module is configured to extract features from the fourth feature information to obtain the fifth feature information; The second feature fusion module is configured to fuse the third feature information, the fourth feature information, and the fifth feature information to obtain multiple fused feature information; the multiple fused feature information correspond one-to-one with the multiple text detection modules respectively; Multiple text detection modules are configured to perform text detection processing on the fused feature information corresponding to each text detection module to obtain the character position information; Optionally, the second feature fusion module includes: a fourth feature extraction unit, a first feature fusion unit, a second feature fusion unit, a fifth feature extraction unit, and a sixth feature extraction unit; The input terminal of the fourth feature extraction unit is connected to the output terminal of the fourth feature extraction module; the output terminals of the fourth feature extraction unit and the fifth feature extraction module are respectively connected to the input terminal of the first feature fusion unit; the output terminals of the first feature fusion unit and the sixth feature extraction module are respectively connected to the input terminal of the second feature fusion unit; the output terminal of the second feature fusion unit is connected to the input terminal of the fifth feature extraction unit; and the output terminal of the fifth feature extraction unit is connected to the input terminal of the sixth feature extraction unit. Optionally, the second recognition information includes at least one of a first text region information, a second text region information, and a third text region information. The image recognition module divides the first image information into text regions based on the character position information to obtain the second recognition information, including: Based on the character position information, determine the first text region information; and / or, Clustering is performed on the first text region information to obtain the second text region information; and / or, The first text region information is fused to obtain the third text region information; Optionally, the image recognition module determines the first text region information based on the character position information, including: The character position information is calculated and processed to obtain multiple first distance information; Clustering is performed on multiple sets of the first distance information to obtain the first target cluster information; The character position information corresponding to the first distance information in the first target clustering information is fused to obtain the first text region information; Optionally, the image recognition module performs clustering processing on the first text region information to obtain the second text region information, including: The information of the first text region is processed to obtain multiple second distance information; Clustering is performed on multiple sets of the second distance information to obtain the second target cluster information; The first text region information corresponding to the second distance information in the second target clustering information is fused to obtain the second text region information; Optionally, the image recognition module performs fusion processing on the first text region information to obtain the third text region information, including: The text region corresponding to the first text region information is adjusted to obtain the fourth text region information. The fourth text region information is fused based on the overlapping area of ​​the text regions corresponding to the fourth text region information to obtain the third text region information. Optionally, the text determination module determines the target text information based on the image recognition information, including: Based on the first identification information and the second identification information, the target text region information is determined; Based on the target text region information, the target text information is determined; Optionally, the text determination module determines target text region information based on the first identification information and the second identification information, including: If the object category corresponding to the first identification information is the first category, the first text region information is filtered based on the third distance information between the first reference point in the first image information and the text region corresponding to the first text region information to obtain the target text region information; and / or, If the object category corresponding to the first identification information is the second category, the second text region information is filtered based on the fourth distance information between the second reference point in the first image information and the text region corresponding to the second text region information to obtain the target text region information; and / or, If the object category corresponding to the first identification information is the third category, the third text region information is filtered based on the similarity information between the reference region in the first image information and the text region corresponding to the third text region information to obtain the target text region information; Optionally, the text determination module determines the target text information based on the target text region information, including: Text information is extracted from the target text region corresponding to the target text region information to obtain candidate text information; The candidate text information is translated to obtain the target text information.

18. A computer device, characterized in that, The computer device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 16.

19. A computer-readable storage medium, characterized in that, It contains a computer program that is loaded by a processor to perform the steps of the method according to any one of claims 1 to 16.