A method for transmitting screen content images based on semantic communication

By using an end-to-end encoding/decoding framework and weight allocation network based on semantic communication, combined with OFDM transmission and channel state information feedback, the transmission of text-based screen content images is optimized, solving the problems of high resource consumption and insufficient robustness in traditional methods, and achieving efficient and reliable image transmission.

CN119484857BActive Publication Date: 2025-10-28CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411644143.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-10-28
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Existing methods for transmitting screen content images are limited by the Shannon limit, making it difficult to efficiently transmit text-based screen content images in low-latency and high-speed wireless communication environments. Furthermore, traditional methods fail to effectively utilize the semantic features of information, resulting in excessive resource consumption and insufficient robustness.

Method used

An end-to-end encoding and decoding framework based on semantic communication is adopted, which combines a weighted allocation network and OFDM transmission. The encoding strategy is optimized through channel state information feedback, important semantic features are transmitted first, and a loss function is constructed using OCR and scene text detection models to protect key text information and achieve differentiated transmission.

Benefits of technology

It improves the transmission efficiency and robustness of text-based screen content images, especially performing well under low signal-to-noise ratio and low bit rate conditions, and enhances the system's reliability and transmission quality in complex wireless channel environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484857B_ABST
    Figure CN119484857B_ABST
Patent Text Reader

Abstract

This invention relates to a method for transmitting screen content images based on semantic communication, belonging to the field of wireless communication and multimedia transmission technology. The method includes: for an input text-based screen content image, source encoding is performed using an encoder; for the source-encoded codewords, a weight allocation network is used to assign weights, prioritizing the transmission of important semantic features on high-SNR subcarriers; channel encoding is performed on the weighted codewords, allocating more transmission power to high-SNR subcarriers in the OFDM transmitter; after receiving the image, the OFDM receiver detects and recognizes the text in the image using OCR and a scene text detection model, calculating text-level confidence; a loss function is designed based on the text-level confidence to protect key text semantic features. This invention improves image transmission quality, saves transmission bandwidth, and solves the problem of traditional screen content image encoding methods being limited by the Shannon limit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication and multimedia transmission technology, and relates to application scenarios that require efficient image transmission, including remote conferencing, cloud gaming, online education, and real-time remote collaboration. In particular, it relates to a method for transmitting screen content images based on semantic communication. Background Technology

[0002] According to Google research, 90% of online media interaction is screen-based, making screen content a crucial type of information on the internet. Generally, screen content refers to computer-generated or rendered data, characterized by noise-free images, sharp edges, repetitive patterns, and high contrast compared to natural scene images. Text-based screen content images and videos refer to image and video data of text-centric screen displays; they are static images capturing text content displayed on the screen. They typically contain a large amount of text and may also include a small amount of graphics and charts. This type of data usually includes documents, web pages, presentations, code editors, and other text-based screen content. Currently, China is vigorously promoting economic development led by the information industry. With the rapid development of IoT-based intelligent communication systems, including multimedia and cloud technologies, online media interaction has been widely applied in scenarios such as cloud video conferencing, remote screen sharing, online games, and online education. While providing convenience, these processes also generate massive amounts of data. The explosive growth of screen content data poses considerable challenges to image and video encoding technologies. Therefore, researching efficient screen content encoding and transmission schemes is imperative when facing these diverse application scenarios.

[0003] Traditional image compression uses standards such as JPEG, JPEG2000, H.265 / HEVC, and H.266 / VVC, while recent Learned Image Compression (LIC) algorithms have surpassed H.266 / VVC. However, these methods are primarily designed for natural scene images, which limits their applicability to the specific properties of screen content and poses challenges to the encoding and transmission of screen content images. Many traditional coding standards have integrated screen content coding (SCC), achieving some progress, but a comprehensive solution for encoding and transmitting screen content images is still lacking. Shannon's separation theorem states that, in the case of infinitely long codewords, designing source coding and channel coding separately is optimal. Separate modular design has achieved great success in practical communication systems. From the development of first-generation to fifth-generation mobile communication systems, traditional communication system design has adopted a modular approach, that is, designing the source coding, channel coding, and modulation modules of the encoder separately. However, in various wireless environments, such as emerging wireless applications like autonomous driving, smart manufacturing, and telemedicine, the implementation of communication systems is impractical due to latency considerations. Separate modular designs are insufficient to meet the complex communication demands of today. Therefore, joint optimization of source-channel coding for shorter information lengths is preferable; this method is called Joint Source-Channel Coding (JSCC). Traditional screen content coding methods in communication systems approach the Shannon limit in terms of transmission rate, posing a significant challenge to the current requirements for low latency and high transmission rates. Due to the explosive growth in data volume, existing source-channel coding methods have significant problems in communication systems with limited bandwidth and low latency requirements. Moreover, traditional communication systems do not focus on the semantic aspects of information, only pursuing the precise transmission of information symbols, which exposes their limitations in the era of artificial intelligence. These issues have prompted researchers to explore paradigms for next-generation communication systems.

[0004] Semantic communication, as an emerging communication technology, can significantly improve communication efficiency by extracting, encoding, and transmitting the semantics of information. Compared with traditional communication, semantic communication does not require the accurate transmission of data or communication symbols, but focuses on the matching between the semantic information input at the sending end and the semantic information recovered at the receiving end. By filtering out redundant information and extracting the meaning of effective information, it transmits truly useful information, thereby significantly reducing the consumption of communication resources. Moreover, unlike traditional communication which uses bit error rate (BER) and symbol error rate (SER) as evaluation results, semantic communication systems recover the source information at the decoder end by minimizing the semantic loss between the input information and the reconstructed information. Furthermore, with the rapid development of artificial intelligence, its ability to grasp human knowledge in limited scenarios is gradually improving. Deep learning, as one of the most important technologies in artificial intelligence, has achieved great success in understanding sources such as language, audio, images, and video. Because deep learning can efficiently extract and transmit semantic information contained in a source, and has the advantages of good feature extraction and learning capabilities, it has been widely used in semantic communication models. Deep learning-based semantic communication technology can not only optimize traditional source, channel coding and modulation / demodulation modules, but also effectively address bottleneck problems in traditional communication systems by establishing end-to-end (E2E) JSCC. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a method for transmitting screen content images based on semantic communication, which improves the transmission efficiency and robustness of text-based screen content images, as well as the completion efficiency of downstream tasks.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for transmitting screen content images based on semantic communication, the method comprising:

[0008] For input text-based screen content images, source encoding is performed using an encoder to reduce redundancy in the source.

[0009] For the source-encoded codewords, a weight allocation network is used to assign weights to them, so that important semantic features are transmitted preferentially on high SNR subcarriers;

[0010] Channel coding is performed on the weighted codewords to allocate more transmission power to high SNR subcarriers in the OFDM transmitter;

[0011] After receiving an image, the OFDM receiver uses OCR and scene text detection models to detect and recognize the text in the image and calculate the text-level confidence score.

[0012] A loss function is designed based on text-level confidence to protect key text semantic features.

[0013] Furthermore, weight allocation of the source-encoded codewords through a weight allocation network includes: for a set of codewords obtained after source encoding... Weights are assigned to the codeword Y using a weighting network. Specifically, the weights are assigned element-wise by multiplying the weight matrix generated by the weighting network with the codeword Y; the weight matrix is ​​represented as W = [w1, w2, ..., w...]. n ], The weight matrix is ​​multiplied element-wise by the codeword Y to obtain the symbol sequence. Symbol sequence This refers to the OFDM symbols that will ultimately be transmitted to the channel. The OFDM transmitter will... Mapped onto multiple subcarriers, parallel transmission improves bandwidth utilization and anti-interference capabilities.

[0014] Furthermore, for the weight allocation network, its training process includes: using the source-encoded codeword Y and the channel frequency response H as inputs to the weight allocation network, and generating a weight matrix W for weight classification after training, wherein the weight matrix W has the same shape as Y, that is, each codeword of Y corresponds to a weight.

[0015] Furthermore, the weight allocation network includes two convolutional layers, two batch normalization layers, and one activated ReLU layer.

[0016] In this network, the input data undergoes feature extraction through the first convolutional layer, generating a set of feature maps:

[0017] Y1 = W1*X + b1

[0018] In the formula, W1 represents the convolution kernel, "*" represents the convolution operation, X represents the input data, and b1 represents the bias.

[0019] These feature maps are then processed through the first normalization layer to standardize the feature map data:

[0020]

[0021] In the formula, μ and σ represent the mean and variance of the mini-batch data, and γ and β represent the learnable parameters;

[0022] The standardized feature map is then subjected to a nonlinear transformation through a ReLU layer:

[0023] Y3 = max(0, Y2)

[0024] The feature map after nonlinear transformation is further processed by a second convolutional layer to extract features:

[0025] Y4=W2*Y3+b2

[0026] In the formula, W2 represents the convolution kernel and b2 represents the bias.

[0027] These feature maps are then normalized again through a second batch of normalization layers:

[0028]

[0029] In the formula, μ' and σ' represent the mean and variance of the mini-batch data, and γ' and β' represent the learnable parameters.

[0030] Furthermore, text in the image is detected and recognized using OCR and scene text detection models. The text-level confidence score is calculated by preprocessing the received image and then detecting text regions in the image using the DB detection model. Specifically, the TextRegion() function is used to detect text regions in the image and returns a list of bounding boxes for all individual text regions in the image.

[0031] B = {B1, B2, ..., B} n = TextRegion(x)

[0032] In the formula, B i The text box represents the detected text box, and x represents the input image;

[0033] For each B i Define coordinates, and crop the text region from the input image x based on the coordinates:

[0034] b i =B i (x)

[0035] In the formula, b i This represents the i-th cropped text region of x;

[0036] The cropped text region is input into the CRNN model, and the Recognize() function is used to recognize each character within the text region:

[0037] c j =Recognize(b i )

[0038] In the formula, c j This represents the confidence score for each character; the average confidence score for all characters is then used as the text-level confidence score. n represents the total number of characters in the recognition result.

[0039] Construct the text loss based on the calculated text-level confidence:

[0040]

[0041] In the formula, The image is represented as reconstructed; the total loss function is constructed by combining text loss, reconstruction loss, and channel estimation loss:

[0042]

[0043] In the formula, L rec L represents the loss function between the original image and the reconstructed image. ce λ represents the channel estimation loss at the receiver, and λ represents the weighting parameter that controls the importance of the text loss function.

[0044] The beneficial effects of this invention are as follows:

[0045] (1) This invention effectively improves the transmission quality of images and saves transmission bandwidth through end-to-end semantic encoding and decoding and OFDM transmission framework, and solves the problem that traditional screen content image encoding methods are limited by the Shannon limit.

[0046] (2) This invention utilizes a dynamic coding strategy based on Channel State Information (CSI) feedback, along with an additional loss function to protect key text semantic features, to achieve differentiated transmission of transmission parts of varying importance. Compared to traditional screen content encoding and transmission methods, this invention not only improves transmission efficiency but also enhances the robustness and reliability of the system under heavy wireless channel conditions. Compared to traditional screen content image transmission schemes, this invention demonstrates particularly significant advantages under low SNR and low bit rate conditions.

[0047] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0049] Figure 1 This is a block diagram of the method described in this statement;

[0050] Figure 2 For source channel coding with CSI feedback;

[0051] Figure 3 This is a diagram illustrating the weight allocation. Detailed Implementation

[0052] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0053] An embodiment of the present invention provides a method for transmitting screen content images based on semantic communication, the method comprising:

[0054] S1. Semantic Feature Extraction and Encoding

[0055] This paper describes a method for extracting semantic features from input text-based screen content images. Based on a deep learning model, an end-to-end semantic framework is used to transform the text images into semantic feature representations. The semantic encoder analyzes the image content, extracts important textual information, and compresses or discards irrelevant data. By reducing unnecessary data transmission, it achieves efficient encoding of key semantic features.

[0056] Specifically as follows:

[0057] Suppose we have a screen content image x. After converting it into a semantic feature representation, an encoder first performs source encoding S(·) on it to reduce redundancy in the source. Given an input image of size H×W, the encoder output size is... Where d is the downsampling factor and C is the output channel.

[0058] At this point, the encoder's output is a real value. This output is converted to a complex number by mapping half to the real part and the other half to the imaginary part. Then, its dimensions are reshaped. Where L fft The number of subcarriers in the OFDM symbol is represented by Y, which represents the frequency-domain complex symbol to be fed to the OFDM encoder. Pilot symbol Both the encoder and decoder are known, and it is assumed that N p All pilot symbols are the same.

[0059] Subsequently, an inverse discrete Fourier transform (IDFT) is performed on each OFDM symbol to transform the frequency domain signal into a time domain signal for transmission in the time domain, and a cyclic prefix (CP) is added. Next, OFDM power allocation is performed. In an orthogonal frequency division multiplexing (OFDM) system, the transmission performance of frequency domain symbols is directly affected by the channel frequency response. In an OFDM system, the channel state of each subcarrier can be different, and variations in the signal-to-noise ratio (SNR) affect transmission performance. Resource allocation is performed based on the SNR of each subcarrier to ensure that important semantic features are transmitted on subcarriers with high SNR, thereby reducing the bit error rate.

[0060] Each source image x is passed through a sequence containing N p pilot symbols and N s OFDM information symbols are transmitted in n subcarriers, where pilot symbols are used for channel estimation and synchronization. These are known symbols that help the decoder understand the channel state, thus enabling better reception and decoding of the information symbols. The information symbols are the actual transmitted data. In a typical OFDM system, the encoded OFDM information symbol y is modulated onto n subcarriers, i.e., y = [s1, s2, ..., sn]. n ].

[0061] S2. Optimize coding parameters through real-time Channel State Information (CSI) feedback. In OFDM systems, the channel state often changes dynamically over time. Therefore, this invention designs a dynamic coding algorithm that adjusts the coding power allocation based on the CSI feedback mechanism, enabling important semantic feature data to be transmitted preferentially on high SNR subcarriers, thus ensuring the quality and stability of image transmission even under low SNR or complex channel conditions. This algorithm combines the instantaneous channel state and feedback information, allowing subcarriers with good channel conditions to prioritize the transmission of key semantic features, thereby improving data transmission quality. The specific algorithm is shown in Table 1 below:

[0062] Table 1

[0063]

[0064] like Figure 2 and 3 As shown, the source text screen content image x is processed by the source encoder to obtain a set of codewords Y = [y1, y2, ..., y]. n ], These codewords Y are the encoded representation of image information, a vectorized representation of the original image data, used for subsequent transmission. The weight matrix W = [w1, w2, ..., w...] has the same shape as the codewords Y. n ], It is generated through a weighted allocation network. During training, this network learns an optimization strategy: assigning weights to the encoded codeword Y based on channel characteristics and the optimization objective of the downstream text recognition task in semantic communication. This prioritizes data that benefits the downstream task, thus influencing various parts of the transmitted image information. The codeword Y and the weights W are multiplied element-wise, i.e., s i =y i ·w i This formed a new symbol sequence. Through modulation with weight W, the encoded codeword Y is adjusted to form the OFDM symbol to be transmitted to the channel. OFDM transmitters will display OFDM symbols Mapped onto multiple subcarriers, parallel transmission improves bandwidth utilization and anti-interference capabilities.

[0065] After performing channel estimation, the receiver feeds back the estimated CSI (Channel Indicator) for each subcarrier to the transmitter in real time, including parameters such as SNR (Short-Range Response) and channel gain. The entire CSI feedback process can be described by the following probabilistic model:

[0066]

[0067] Where v represents the output of the compression layer, p φ (v|h'), p(z|v), These are defined for compression, quantization, feedback channel, and decoder, respectively.

[0068] The transmitter allocates power based on this feedback information. During channel coding, it allocates more transmission power to high-SNR subcarriers in the OFDM system to transmit images at the maximum transmission rate and minimize the bit error rate. Furthermore, to ensure that important textual semantic features are prioritized for transmission, the system can allocate high-SNR subcarriers to these features based on their importance.

[0069] The proposed weighted encoding network consists of two convolutional layers, two batch normalization layers, and a ReLU activation function layer. It extracts and normalizes features from the input data, using convolutional layers to extract features, batch normalization layers to stabilize the training process, and ReLU to introduce non-linearity. This entire process helps the network better learn the complex patterns of the input data. First, the input data passes through the first convolutional layer for feature extraction, generating a set of feature maps. This process can be represented by the following formula:

[0070] Y1 = W1*X + b1

[0071] Where W1 is the convolution kernel, "*" indicates the convolution operation, X is the input data, and b1 is the bias. Next, these feature maps are processed through a batch normalization layer to standardize the feature map data.

[0072]

[0073] Where μ and σ are the mean and variance of the mini-batch data, and γ and β are learnable parameters. Then, the standardized feature maps are activated by the ReLU function, applying a non-linear transformation:

[0074] Y3 = max(0, Y2)

[0075] Next, the feature map after nonlinear transformation is passed through a second convolutional layer to further extract features:

[0076] Y4=W2*Y3+b2

[0077] Here, W2 is the second convolutional kernel, and b2 is the second bias. Finally, these feature maps are again processed through a batch normalization layer for standardization.

[0078]

[0079] Here, μ' and σ' are the mean and variance of the new mini-batch data, and γ' and β' are new learnable parameters. This sub-network extracts and normalizes features from the input data, extracts features through convolutional layers, stabilizes the training process through batch normalization layers, and introduces nonlinearity through ReLU. The entire process helps the network better learn the complex patterns of the input data.

[0080] S3. At the receiving end, the transmitted data is decoded and reconstructed using an end-to-end semantic decoder. By combining CSI feedback and text protection mechanisms, the receiving end can effectively recover the semantic content of the image, especially textual information, ensuring the image reconstruction effect at the receiving end. Through a precise semantic decoding process, the image recovered by the receiving end is visually highly consistent with the original image, thus ensuring both transmission efficiency and semantic accuracy of the image content.

[0081] During image transmission, determining the main semantic information of the transmitted data and filtering and optimizing it according to the needs of downstream tasks are crucial. Identifying key objects and features in the image and prioritizing the transmission of this information to enable the receiving end to complete its tasks quickly and accurately is a key consideration in screen content image transmission. Furthermore, images containing a large amount of text, especially screen content images, often suffer from text distortion during encoding and transmission. To address the issues of text reconstruction quality and semantic communication in downstream tasks during image encoding and transmission, this method proposes a loss function for text recovery, designed to measure the difference in text fidelity between the original and reconstructed images.

[0082] Specifically, Optical Character Recognition (OCR) and scene text detection models are used to detect and recognize text in images, and a proposed loss function is constructed as an addition to the total loss. Minimizing this loss enhances the perceptual similarity of text between two images. The total loss function consists of several parts, including a loss L that measures the difference between the original and reconstructed images. rec And the channel estimation loss L at the decoder end ce In addition, a text loss function L is introduced. t As part of the overall loss function, it is used to characterize the difference in text fidelity.

[0083] Among them, text loss L t It is based on the reconstructed image The calculated text confidence list is extracted from the image by recognizing, cropping, and identifying text regions. Text recognition involves four steps: image preprocessing, text detection, text region cropping, and text recognition. First, image preprocessing performs operations such as grayscale conversion and binarization on the input image to improve text recognition accuracy. Then, text detection uses a pre-trained DB model to detect text regions in the image. The function TextRegion() detects text regions in the image and returns a list of bounding boxes for all individual text regions in the image.

[0084] B = {B1, B2, ..., B} n = TextRegion(x)

[0085] Next, based on the detected text boxes, the text regions in the image are cropped out, for each B i This can be defined using the coordinates of the top-left and bottom-right corners. Then, let B... i As a function, the input image is cropped based on coordinates to obtain the text region.

[0086] b i =B i (x)

[0087] in This represents the i-th cropped text region of x, with height h. i Width is w i Finally, the cropped text region is input into the pre-trained CRNN model. The function Recognize(·) can recognize each character within the given text region.

[0088] c i =Recognize(b i )

[0089] Where c i This represents the confidence level for each character, typically ranging from 0 to 1, indicating the model's level of confidence in the recognition result. A higher confidence level indicates greater confidence in the recognition outcome.

[0090] The average confidence score for all characters is taken as the confidence score of the entire text recognition result. Assume the recognition result contains n characters, and the character-level confidence scores are {c1, c2, ..., cn}. n Then, the text-level confidence score is calculated as follows:

[0091]

[0092] After obtaining the text recognition confidence score, it can be applied to the design of the loss function, defining the text loss as:

[0093]

[0094] Text loss function After adding the weights, the total loss function becomes:

[0095]

[0096] Among them, L rec L is the loss function between the original image and the reconstructed image. ce λ is the channel estimation loss at the receiver end, and λ is a weighting parameter that controls the importance of the text loss function. λ can be set to 0.8.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for transmitting screen content images based on semantic communication, characterized in that, The method includes: For input text-based screen content images, source encoding is performed using an encoder to reduce redundancy in the source. For the source-coded codewords, a weight allocation network is used to assign weights to them, so that important semantic features are transmitted preferentially on high SNR subcarriers; first, the source-coded codewords are... Y and channel frequency response H As input to the weight allocation network, a weight matrix for weight classification is generated after training. , where the weight matrix and Y Having the same shape, that is Y Each codeword corresponds to a weight; then, for the codewords obtained after source encoding , The weights are assigned through the weight allocation network, wherein the weight matrix generated by the weight allocation network and the codewords are... Weight allocation is achieved through element-wise multiplication; the weight matrix is ​​represented as follows: , , The number of information symbols in OFDM information symbols. This represents the number of subcarriers in an OFDM symbol; the weight matrix and codeword Y Element-wise multiplication yields a symbolic sequence : OFDM transmitters will transmit symbol sequences Mapped onto multiple subcarriers for parallel transmission; Channel coding is performed on the weighted codewords to allocate more transmission power to high SNR subcarriers in the OFDM transmitter; After receiving an image, the OFDM receiver uses OCR and scene text detection models to detect and recognize the text in the image and calculate the text-level confidence score. A loss function is designed based on text-level confidence, and the text loss is constructed based on the calculated text-level confidence: , The image represents the reconstructed image; key textual semantic features are preserved through text loss.

2. The method according to claim 1, characterized in that: The weight allocation network includes two convolutional layers, two batch normalization layers, and one activated ReLU layer; The input data is processed through the first convolutional layer to extract features, producing a set of feature maps: In the formula, "*" represents the convolution kernel, and "*" represents the convolution operation. X Indicates input data, Indicates bias; These feature maps are then processed through the first normalization layer to standardize the feature map data: In the formula, and This represents the mean and variance of a small batch of data. and Indicates learnable parameters; The standardized feature map is then subjected to a nonlinear transformation through a ReLU layer: The feature map after nonlinear transformation is further processed by a second convolutional layer to extract features: In the formula, Represents the convolution kernel. Indicates bias; These feature maps are then normalized again through a second batch of normalization layers: In the formula, and This represents the mean and variance of a small batch of data. and This represents the learnable parameters.

3. The method according to claim 1, characterized in that: Text in images is detected and recognized using OCR and scene text detection models. The text-level confidence score is calculated by preprocessing the received image and then detecting text regions in the image using a DB detection model. The function detects text regions in an image and returns a list of bounding boxes for all individual text regions in the image. In the formula, This indicates the detected text box. Indicates the input image; For each Define coordinates, and apply them to the input image. The text area is obtained by cropping: In the formula, express x The i A cropped text area; The cropped text region is input into the CRNN model, and then... The function identifies each character within a text region: In the formula, This represents the confidence score for each character; the average confidence score for all characters is then used as the text-level confidence score. , This indicates the total number of characters in the recognition result.

4. The method according to claim 1 or 3, characterized in that: The total loss function is constructed by combining text loss, reconstruction loss, and channel estimation loss: In the formula, This represents the loss function between the original image and the reconstructed image. This represents the channel estimation loss at the receiver. λ The weight parameters represent the weights that control the importance of the text loss function. x This represents the input image.

Citation Information

Patent Citations

  • Scene character recognition method and system based on semantic enhancement encoder decoder framework

    CN111753827A

  • Channel charting in wireless systems

    CN111869291A