Data Processing Method, Apparatus, Electronic Device, and Storage Medium

By extracting and utilizing text information and metadata in the image sequence in the video conferencing system, the problem of degradation in screen sharing image quality in video conferencing is solved, and the clarity of text content and the overall efficiency and quality of video conferencing are improved.

CN114154457BActive Publication Date: 2025-06-03ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010931058.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-07
Publication Date
2025-06-03
Estimated Expiration
2040-09-07

AI Technical Summary

Technical Problem

In video conferencing scenarios, the image/video quality shared by the screen is easily affected by factors such as network bandwidth fluctuations and image compression, resulting in blurred text information and reducing the efficiency and quality of video conferencing.

Method used

By extracting text information and metadata in the image sequence, using machine learning models to detect text information and extract metadata, and encoding the image sequence based on this information, improving the encoding efficiency and clarity of the area where the text content is located.

Benefits of technology

By extracting and utilizing text information metadata, the text content in video conferences can be more effectively encoded, image quality, reduced blur, and improved the efficiency and quality of video conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114154457B_ABST
    Figure CN114154457B_ABST
Patent Text Reader

Abstract

An embodiment of the present disclosure discloses a data processing method, apparatus, electronic device, and storage medium. The method includes: obtaining a first image sequence; the first image sequence includes at least one image; extracting text information from the first image sequence; the text information includes metadata of the text content included in the first image sequence; encoding the first image sequence based on the text information. This technical solution can use the text information extracted from the image sequence as prior information for the encoding device, enabling the encoding device to more effectively encode the text content in the image sequence, and improving the encoding efficiency, accuracy, and clarity of the region where the text content is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a data processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In a video conferencing scenario, screen sharing occurs frequently. Most of the screen sharing content contains text content. In a video conferencing system, the screen sharing content needs to be processed through multiple links such as image / video compression, network transmission, and even image scaling before it can be delivered to the receiving end. Image / video compression is to reduce the amount of information transmitted over the network. Usually, encoding and decoding standards such as H.264 and H.265 are used, which is a lossy compression method. The effect of network transmission is affected by network bandwidth fluctuations, etc., causing fluctuations in the image quality received by the receiving end. Image scaling is to adapt to devices with different performances. Usually, BI or deep learning, etc. is used for upsampling and downsampling of images, which also introduces corresponding losses of image content. The above links will all cause a decrease in image quality, especially the text information in the conference scenario. Once the picture becomes blurred, the receiving party cannot obtain effective information, which will further reduce the efficiency and quality of the video conference. Summary of the Invention

[0003] Embodiments of the present disclosure provide a data processing method, apparatus, electronic device, and computer-readable storage medium.

[0004] In a first aspect, an embodiment of the present disclosure provides a data processing method, including:

[0005] Obtain a first image sequence; the first image sequence includes at least one image;

[0006] Extract text information from the first image sequence; the text information includes metadata of the text content included in the first image sequence;

[0007] Encode the first image sequence based on the text information.

[0008] Further, extracting the text information from the first image sequence includes:

[0009] Extract the text information from the first image sequence by using a machine learning model.

[0010] Further, the metadata includes position information of the area where the text content is located, position change information of the area where the text content is located in consecutive multiple images, and / or information of new text content appearing in the current image.

[0011] Further, encoding the first image sequence based on the text information includes:

[0012] Provide the first image sequence and the text information to an encoding device so that the encoding device encodes the text content included in the first image sequence based on the text information.

[0013] Further, it further includes:

[0014] Receive encoded data;

[0015] Decode the encoded data to obtain a second image sequence;

[0016] Perform information enhancement processing on the text content included in the second image sequence.

[0017] In a second aspect, an embodiment of the present disclosure provides a data processing method, including:

[0018] Receive a first image sequence and the text information in the first image sequence; the first image sequence includes at least one image; the text information includes metadata of the text content included in the first image sequence;

[0019] Encode the first image sequence based on the text information.

[0020] Further, the metadata includes position information of the area where the text content is located, position change information of the area where the text content is located in consecutive multiple images, and / or information of new text content appearing in the current image.

[0021] Further, the metadata includes the area range where the text content is located. Encoding the first image sequence based on the text information includes:

[0022] Encode the images within the area where the text content is located according to preset encoding parameters; the preset encoding parameters include quantization step size.

[0023] Further, the metadata includes position change information of the text content in consecutive multiple images in the first image sequence. Encoding the first image sequence based on the text information includes:

[0024] During the inter-frame prediction process, perform motion estimation on the images within the area where the text content is located based on the position change information.

[0025] Further, the metadata includes information of new text content appearing in the current image. Encoding the first image sequence based on the text information includes:

[0026] When there is new text content in the current image, skip the motion estimation for the images within the area where the new text content is located during the inter-frame prediction process.

[0027] In a third aspect, an embodiment of the present disclosure provides a data processing method, including:

[0028] Obtaining a first image sequence; the first image sequence includes at least one image;

[0029] Invoking a first preset service interface to extract text information from the first image sequence by the first preset service interface and encode the first image sequence based on the text information; the text information includes metadata of the text content included in the first image sequence;

[0030] Returning the encoded data of the first image sequence.

[0031] Further, it further includes:

[0032] Receiving encoded data;

[0033] Invoking a second preset service interface to decode the encoded data by the second preset service interface to obtain a second image sequence and perform information enhancement processing on the text content included in the second image sequence;

[0034] Returning the second image sequence after the enhancement processing.

[0035] In a fourth aspect, an embodiment of the present disclosure provides a data processing device, including:

[0036] A first acquisition module configured to acquire a first image sequence; the first image sequence includes at least one image;

[0037] An extraction module configured to extract text information from the first image sequence; the text information includes metadata of the text content included in the first image sequence;

[0038] A first encoding module configured to encode the first image sequence based on the text information.

[0039] In a fifth aspect, an embodiment of the present disclosure provides a data processing device, including:

[0040] A second receiving module configured to receive a first image sequence and the text information in the first image sequence; the first image sequence includes at least one image; the text information includes metadata of the text content included in the first image sequence;

[0041] A second encoding module configured to encode the first image sequence based on the text information.

[0042] Sixth aspect, an embodiment of the present disclosure provides a data processing device, including:

[0043] A second acquisition module, configured to acquire a first image sequence; the first image sequence includes at least one image;

[0044] A first call module, configured to call a first preset service interface, so that the first preset service interface extracts text information from the first image sequence, and encodes the first image sequence based on the text information; the text information includes metadata of the text content included in the first image sequence;

[0045] A first return module, configured to return the encoded data of the first image sequence.

[0046] The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0047] In a possible design, the structure of the above device includes a memory and a processor. The memory is used to store one or more computer instructions for supporting the above device to execute the corresponding method, and the processor is configured to execute the computer instructions stored in the memory. The above device may further include a communication interface for the above device to communicate with other devices or communication networks.

[0048] Seventh aspect, an embodiment of the present disclosure provides an electronic device, including a memory and a processor; wherein, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method described in any of the above aspects.

[0049] Eighth aspect, an embodiment of the present disclosure provides a computer-readable storage medium for storing computer instructions used by any of the above devices, which includes computer instructions involved in executing the method described in any of the above aspects.

[0050] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0051] For the first image sequence to be encoded, the embodiments of the present disclosure first extract the text information from the first image sequence. The text information includes metadata of the text content included in the first image sequence, and then encodes the area where the text content is located in the first image sequence based on the text information. In this way, the text information extracted from the image sequence is used as the prior information of the encoding device, so that the encoding device can encode the text content in the image sequence more effectively, improving the encoding efficiency, accuracy, and clarity of the area where the text content is located.

[0052] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In conjunction with the accompanying drawings, other features, objects, and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments. In the drawings:

[0054] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure;

[0055] Figure 2 A flowchart showing another data processing method according to an embodiment of the present disclosure;

[0056] Figure 3 A flowchart showing another data processing method according to an embodiment of the present disclosure;

[0057] Figure 4 A schematic diagram of an application scenario of the data processing method according to an embodiment of the present disclosure in screen sharing;

[0058] Figure 5 A schematic diagram of the structure of an electronic device suitable for implementing the data processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts unrelated to the description of the exemplary embodiments are omitted in the drawings.

[0060] In the present disclosure, it should be understood that terms such as "including" or "having" are intended to indicate the presence of the features, numbers, steps, actions, components, parts, or combinations thereof disclosed in this specification, and are not intended to exclude the possibility of the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0061] It should also be noted that, without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0062] The details of the embodiments of the present disclosure will be introduced in detail below through specific examples.

[0063] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure. As Figure 1 shown, the data processing method includes the following steps:

[0064] In step S101, a first image sequence is acquired; the first image sequence includes at least one image;

[0065] In step S102, text information in the first image sequence is extracted; the text information includes metadata of text content contained in the first image sequence;

[0066] In step S103, the first image sequence is encoded based on the text information.

[0067] In this embodiment, the first image sequence may include one or more images, that is, the first image sequence may be a single image or a continuous video frame. In some embodiments, the first image sequence may be an image sequence for screen sharing. After the screen sharing requester initiates a screen sharing request on the terminal device, the terminal device may obtain a screenshot to be shared or record the screen to obtain an image sequence according to the request.

[0068] For the acquired first image sequence, text information can be extracted from each image in the first image sequence, and the text information may include but is not limited to metadata of the text content in each image, which may also be referred to as attribute information of the text content. In some embodiments, the metadata of the text content may include but is not limited to position information of the area where the text content is located, position change information of the area where the text content is located in multiple consecutive images, and / or new text content information appearing in the current image, etc.

[0069] The position information of the area where the text content is located may include the relative position information of the area where the text content is located in the current image, and the position change information may include the change information of the position information of the area where the text content is located in a plurality of consecutive images in the first image sequence in the two previous images, etc. For example, the displacement of the position of the area where the text content is located in the current image compared with the position in the previous image, etc. The new text content information appearing in the current image may include but is not limited to whether new text content appears in the current image and the position information of the area where the new text content appears. Whether new text content appears in the current image can be understood as all or part of the text content appearing in the current image not appearing in the previous image.

[0070] After extracting the text information in the first image sequence, the first image sequence can be encoded based on the text information. In some embodiments, an image encoding device or a video encoding device can be used to encode the first image sequence. When the first image sequence includes only one image, that is, the current screen sharing is a static image, an image encoder is used for encoding. When the first image sequence is a series of consecutive images, that is, the current screen sharing is a dynamic video sequence, a video encoder is used for encoding. Usually, the encoding device encodes image blocks in the image, and different encoding parameters may be selected for different image blocks, and matching encoding may also be performed by referring to corresponding image blocks in the reference frame. The encoding device in the embodiments of the present disclosure can encode the image blocks corresponding to the region where the text content is located based on the text information extracted from the first image sequence. This text information can be used as prior information for the encoding device, enabling the text information to guide the encoding device to more effectively encode the region where the text content is located in the first image sequence, ultimately achieving more accurate and efficient encoding and decoding of the text content in the first image sequence, and ensuring that the decoded text content is not blurred.

[0071] In the embodiments of the present disclosure, for the first image sequence to be encoded, the text information in the first image sequence is first extracted. The text information includes the metadata of the text content included in the first image sequence, and then the region where the text content is located in the first image sequence is encoded based on the text information. In this way, the text information extracted from the image sequence is used as prior information for the encoding device, enabling the encoding device to more effectively encode the text content in the image sequence, and improving the encoding efficiency, accuracy, and clarity of the region where the text content is located.

[0072] In an alternative implementation of this embodiment, step S102, that is, the step of extracting the text information in the first image sequence, further includes the following steps:

[0073] Use a machine learning model to extract the text information from the first image sequence.

[0074] In this alternative implementation, a machine learning model can be used to extract text information from the first image sequence, and the machine learning model can be pre-trained. The machine learning model can be constructed using a neural network and can include, but is not limited to, CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), etc.

[0075] The input of the machine learning model can be a first image sequence, that is, a single frame of image or a continuous multi-frame image, and the output can be text information related to the text content in the input image, including but not limited to the position information of the text content in the entire image, the motion information of the text content in the front and back two frames of images (that is, the position change information of the area where the text content is located in the front and back two frames of images), etc. Among them, the position information can be obtained by detecting the text content through a neural network, and the position change information can be obtained by extracting the motion information through an optical flow network or an adaptive convolutional network, etc. The machine learning model can also detect whether the text content in each image is newly emerging text content, and if so, can also identify which part of the text content in the image is newly emerging text content, etc.

[0076] In an alternative implementation of this embodiment, step S103, that is, the step of encoding the first image sequence based on the text information, further includes the following steps:

[0077] Provide the first image sequence and the text information to an encoding device, so that the encoding device encodes the text content included in the first image sequence based on the text information.

[0078] In this alternative implementation, the encoding device can be an encoding device that complies with domestic and foreign encoding and decoding standards such as H.264, H.265, and AVS. After extracting the text information from the first image sequence to be encoded, the first image sequence and the text information are provided to the encoding device, so that the encoding device can encode the first image sequence according to the text information. In some embodiments, the encoding device can encode the area where the text content in the first image sequence is located according to the text information.

[0079] The text information extracted from the first image sequence in the embodiments of the present disclosure can guide the encoding device to a certain extent during the encoding process. For example, for the area where the text content is located, the encoding device can be guided to perform more refined encoding. For example, a smaller quantization parameter QP can be used or a coding block with a smaller size can be used to encode the image area where the text content is located, so that the image quality of the image area where the text content is located is higher, that is, the compression ratio of the image area where the text content is located is reduced and the encoding quality is improved. In addition, through the position change information of the area where the text part is located in the first image sequence (such as image content translation caused by mouse sliding, etc.), the motion estimation efficiency and accuracy of the encoding device during the inter-frame prediction process can be assisted. This is because neural networks and the like used to build machine learning models are based on pixel motion calculations, and this method itself is more refined than the motion estimation based on image blocks in the encoding device; in addition, the motion estimation of the encoding device during the inter-frame prediction process has a certain search range limit, while machine learning models and the like do not have this limitation. Therefore, the machine learning model will be more accurate in extracting the motion information of the image area where the text content is located. Therefore, the text information extracted from the first image sequence can effectively guide the subsequent encoding device to perform more effective encoding.

[0080] In an alternative implementation of this embodiment, the method further includes:

[0081] Receiving encoded data;

[0082] Decoding the encoded data to obtain a second image sequence;

[0083] Performing information enhancement processing on the text content included in the second image sequence.

[0084] In this alternative implementation, the data processing method can be executed on a terminal device for screen sharing. The terminal device can also receive and display the screen sharing content of other devices. The screen sharing content received by the terminal device can be the encoded data of the screen sharing content on other devices. After receiving the encoded data, the terminal device first decodes the encoded data to obtain the corresponding second image sequence. The second image sequence can include one or more images, that is, the second image sequence can include static images or can include a dynamic video frame sequence. In addition, the second image sequence can also include text content.

[0085] Since the encoded data received from other devices is lossy compressed encoded data, the second image sequence obtained by decoding may have problems such as loss of texture details, for example, the text becomes blurred. Therefore, in the embodiments of the present disclosure, information enhancement processing is performed on the decoded second image sequence so that the text content part in the second image sequence can be made clearer by means of sharpening or the like. In some embodiments, another machine learning model can be constructed through a neural network or the like. The machine learning model can adopt, but is not limited to, the Unet neural network structure. Its input is the second image sequence obtained by decoding, and the output is the image after information enhancement processing. In other embodiments, the machine learning model can also adopt GAN network technology to restore the texture information in the image, especially in the area where the text content is located, to a certain extent.

[0086] Figure 2 FIG. shows a flowchart of another data processing method according to an embodiment of the present disclosure. As Figure 2 shown, the data processing method includes the following steps:

[0087] In step S201, a first image sequence and the text information in the first image sequence are received; the first image sequence includes at least one image; the text information includes metadata of the text content included in the first image sequence;

[0088] In step S202, the first image sequence is encoded based on the text information.

[0089] In this embodiment, the data processing method is executed on an encoding device. The encoding device can be an image encoding device or a video encoding device. For example, it can include, but is not limited to, encoding devices that follow domestic and foreign video coding standards such as H.264, H.265, and AVS. The data processing device that executes Figure 1 the data processing method in the illustrated embodiment and related embodiments can first extract the text information from the first image sequence, and then transmit the extracted text information and the first image sequence to the encoding device. After receiving the first image sequence and the text information, the encoding device uses the text information as prior information to encode the image area where the text content is located in the first image sequence, so that the text content in the first image sequence can be encoded with high quality and high efficiency.

[0090] For details of extracting the text information from the first image sequence, reference can be made to the description in the above Figure 1 illustrated embodiment and related embodiments, which will not be elaborated here.

[0091] In some embodiments, the metadata in the text information may include, but is not limited to, the regional range where the text content is located, the position change information of the text content in multiple consecutive images in the first image sequence, and / or the information of new text content appearing in the current image. The text information extracted from the first image sequence can guide the encoding device during the encoding process. For example, for the image region where the text content is located, it can guide the encoding device to perform more refined encoding. For instance, a smaller quantization parameter QP or a smaller-sized coding block can be used to encode the image region where the text content is located, so that the image quality of the image region where the text content is located is higher, that is, the compression ratio of the image region where the text content is located is reduced, and the encoding quality is improved. In addition, through the position change information of the image region where the text part extracted from the first image sequence is located (such as the translation of the image content caused by mouse sliding, etc.), the efficiency and accuracy of the inter-frame motion estimation of the encoding device can be assisted. This is because the neural network and the like used to construct the machine learning model are a kind of calculation based on pixel motion, and this method itself is more refined than the motion estimation based on image blocks in the encoding device; in addition, there is a certain search range limit for the motion estimation in the inter-frame prediction process of the encoding device, while the machine learning model and the like do not have such a limit. Therefore, the machine learning model will be more accurate in extracting the motion information of the image region where the text content is located. Therefore, the text information extracted from the first image sequence can effectively guide the subsequent encoding device to perform more effective encoding.

[0092] After extracting the text information in the image sequence in the embodiments of the present disclosure, it is sent to the encoding device together with the image sequence to guide the encoding process of the encoding device, so that the encoding device can perform high-quality and high-efficiency encoding on the image region where the text content in the image sequence is located according to the text information.

[0093] In an alternative implementation manner of this embodiment, the metadata includes the regional range where the text content is located. Step S202, that is, the step of encoding the first image sequence based on the text information, further includes the following steps:

[0094] Encode the images within the regional range in the first image sequence according to preset encoding parameters; the preset encoding parameters include quantization step size.

[0095] In this optional implementation, the text information includes metadata of the text content included in the first image sequence. The metadata may include the region range of the text content in the first image sequence, and the region range may be the position information of the image region where the text content is located. After obtaining the position information of the region range where the text content is located in the first image sequence, the encoding device may perform high-quality encoding on the image region where the text content is located. For example, the image region where the text content is located may be encoded according to preset encoding parameters during the encoding process.

[0096] In some embodiments, the preset encoding parameters may include, but are not limited to, the quantization step QP. When the encoding process of the encoding device enters the region range where the text content is located, the quantization step parameter QP at the current moment may be adjusted to be smaller (the smaller the QP, the better the encoding quality; the larger the QP, the worse the encoding quality). The process of adjusting QP guides the original bitrate control process of the encoding device through the region range of the text content extracted from the first image sequence, making it more accurately allocate the bitrate. That is, in this way, more codewords are allocated to key regions in the image sequence, such as the text content region, to ensure the quality of the key region images.

[0097] In an optional implementation of this embodiment, the metadata includes the position change information of the text content in multiple consecutive images in the first image sequence. Step S202, that is, the step of encoding the first image sequence based on the text information, further includes the following steps:

[0098] During the inter-frame prediction process, motion estimation is performed on the images within the region where the text content is located based on the position change information.

[0099] In this optional implementation, during the encoding process, inter-frame prediction mainly includes motion estimation and motion compensation. Motion estimation is the process of searching for a matching block within an allowable search range (generally one sub-block size in each of the up, down, left, and right directions).

[0100] During the search process, the encoding device generally starts from the corresponding position in the reference frame and searches within a certain range. However, this search range is usually fixed. If the region range where the text content is located in the current frame moves relatively violently compared to the reference frame, it is possible that the corresponding motion relationship cannot be captured during the motion search process, resulting in a decline in encoding performance. In contrast, the embodiments of the present disclosure can perform global motion relationship matching using a machine learning model. That is, instead of limiting the search range, it matches with image blocks within the entire range of the reference frame, without omission. Therefore, using the position encoding information of the text content extracted from the first image sequence in the present disclosure, that is, the motion relationship of the image region where the text content part is located, can avoid the occurrence of the above situation.

[0101] In addition, during the search process, the encoding device first determines the search starting point, and then determines whether the starting point meets the requirements of the matching block and whether further search can be continued. According to the search rule, the surrounding points of the starting point are used as the new starting point for recursive search. After the machine learning model extracts the position information of the area where the text content is located in the first image sequence in the embodiment of the present disclosure, since the position information is already accurate and determined position information, the encoding device can directly determine the starting search point based on this position information, enabling the encoding device to only perform fine search within a small range, saving encoding time and improving encoding efficiency.

[0102] In an alternative implementation manner of this embodiment, the metadata includes the information of new text content appearing in the current image. Step S202, that is, the step of encoding the first image sequence based on the text information, further includes the following steps:

[0103] When there is new text content in the current image, the motion estimation step for the image within the area where the new text content is located is skipped during the inter-frame prediction process.

[0104] In this alternative implementation manner, since during the inter-frame prediction process, the encoding device matches the current image block in the current image frame with the reference image block at the corresponding position in the reference image frame. Through the metadata included in the text information in the embodiment of the present disclosure, that is, the information of new text content appearing in the current image, it can guide whether the encoding device needs to perform the above matching process. The information of new text content appearing in the current image may include, but is not limited to, whether new text content appears in the current image and the position information of the area where the new text content appears, etc. If the metadata in the text information indicates that new text content appears in the current image, the encoding device can determine the area where the new text content is located according to the position information, and then directly skip the motion estimation step when encoding this area. This is because the text content in this area is newly appeared and does not need to be encoded according to the information in the reference frame, so there is no need to perform the motion estimation step. In this way, encoding time can be saved and encoding efficiency can be further improved.

[0105] Figure 3 The flowchart showing another data processing method according to an embodiment of the present disclosure is as follows Figure 3 As shown, this data processing method includes the following steps:

[0106] In step S301, a first image sequence is obtained; the first image sequence includes at least one image;

[0107] In step S302, a first preset service interface is called so that the first preset service interface extracts text information in the first image sequence and encodes the first image sequence based on the text information; the text information includes metadata of text content contained in the first image sequence;

[0108] In step S303, the encoded data of the first image sequence is returned.

[0109] In this embodiment, the data processing method can be executed on a terminal device. A first preset service interface can be pre-deployed in the cloud, and the first preset service interface can be a SaaS (Software-as-a-service) interface. The demander can obtain the right to use the first preset service interface in advance, and when necessary, can encode the first image sequence with text content by calling the first preset service interface and obtain the encoded data. The first preset service interface extracts the text information from the acquired first image sequence, and the text information may include but is not limited to the metadata of the text content contained in the first image sequence; the first preset service interface also sends the text information and the first image sequence to the encoding device for encoding, so that the encoding device encodes the first image sequence according to the text information, and the first preset service interface also obtains the encoded data of the first image sequence from the encoding device, and then returns the encoded data.

[0110] The first image sequence may include one or more images, that is, the first image sequence may be a single image or a continuous video frame. In some embodiments, the first image sequence may be an image sequence for screen sharing. After the screen sharing requester initiates a screen sharing request on the terminal device, the terminal device may obtain a screenshot of the screen to be shared or record the screen to obtain the image sequence according to the request.

[0111] For the acquired first image sequence, text information can be extracted from each image in the first image sequence, and the text information may include but is not limited to metadata of the text content in each image, which may also be referred to as attribute information of the text content. In some embodiments, the metadata of the text content may include but is not limited to position information of the area where the text content is located, position change information of the area where the text content is located in multiple consecutive images, and / or new text content information appearing in the current image, etc.

[0112] The location information of the area where the text content is located may include the relative location information of the area where the text content is located in the current image. The location change information may include the change information of the location of the area where the text content is located in consecutive multiple images in the first image sequence between the previous and the next images, etc. For example, the displacement of the area where the text content is located in the current image compared to its location in the previous image. The new text content information that appears in the current image may include, but is not limited to, whether new text content appears in the current image and the location information of the area where the new text content appears, etc. Whether new text content appears in the current image can be understood as all or part of the text content that appears in the current image did not appear in the previous image.

[0113] After extracting the text information from the first image sequence, the first image sequence can be encoded based on this text information. In some embodiments, an image encoding device or a video encoding device can be used to encode the first image sequence. When the first image sequence includes only one image, that is, the current screen sharing is a static image, an image encoder is used for encoding. When the first image sequence is a continuous multi-image sequence, that is, the current screen sharing is a dynamic video sequence, a video encoder is used for encoding. Usually, the encoding device encodes the image blocks in the image, and different encoding parameters may be selected for different image blocks, and matching encoding may also be performed through the corresponding image blocks in the reference frame. The encoding device in the embodiments of the present disclosure can encode the image blocks corresponding to the area where the text content is located based on the text information extracted from the first image sequence. This text information can be used as the prior information of the encoding device, enabling this text information to guide the encoding device to perform more effective encoding on the area where the text content is located in the first image sequence, ultimately achieving more accurate, efficient encoding and decoding of the text content in the first image sequence, and ensuring that the decoded text content is not blurred.

[0114] In the embodiments of the present disclosure, by pre-deploying a service interface, the text information is extracted from the first image sequence to be encoded by calling this server interface when needed. This text information includes the metadata of the text content included in the first image sequence, and then the area where the text content is located in the first image sequence is encoded based on this text information. In this way, a service can be provided for the demand side. When encoding an image sequence containing text content, using the text information extracted from the image sequence as prior information enables more effective encoding when encoding the area where the text content is located, improving the encoding efficiency, accuracy, and clarity of the area where the text content is located, etc.

[0115] In an alternative implementation manner of this embodiment, extracting the text information from the first image sequence includes:

[0116] Extract the text information from the first image sequence using a machine learning model.

[0117] Details of this alternative implementation can be found in Figure 1 the descriptions in the illustrated embodiments and related embodiments, which will not be elaborated here.

[0118] In an alternative implementation of this embodiment, encoding the first image sequence based on the text information includes:

[0119] Provide the first image sequence and the text information to an encoding device so that the encoding device encodes the text content included in the first image sequence based on the text information.

[0120] Details of this alternative implementation can be found in Figure 1 the descriptions in the illustrated embodiments and related embodiments, which will not be elaborated here.

[0121] In an alternative implementation of this embodiment, the method further includes:

[0122] Receive encoded data;

[0123] Call a second preset service interface so that the second preset service interface decodes the encoded data to obtain a second image sequence and performs information enhancement processing on the text content included in the second image sequence;

[0124] Return the second image sequence after the enhancement processing.

[0125] In this alternative implementation, the second preset service interface can also be pre-deployed in the cloud. The second preset service interface can also be a Saas (Software-as-a-service) interface. The requester can obtain the right to use the second preset service interface in advance and call the second preset service interface to decode the encoded data when needed to obtain a second image sequence. After receiving the encoded data of the second image sequence, the second preset service interface decodes the encoded data and performs information enhancement processing on the text content included in the decoded second image sequence.

[0126] The second image sequence may include one or more images, that is, the second image sequence may include static images or may include a dynamic video frame sequence. In addition, the second image sequence may also include text content.

[0127] Since the encoded data received from other devices is lossy compressed encoded data, texture details may be lost in the decoded second image sequence, such as the text becoming blurred. Therefore, the embodiments of the present disclosure perform information enhancement processing on the decoded second image sequence so that the text content part in the second image sequence can be made clearer by means of sharpening or the like. In some embodiments, another machine learning model can be constructed through a neural network or the like. The machine learning model can adopt, but is not limited to, the Unet neural network structure. Its input is the decoded second image sequence, and the output is the image after information enhancement processing. In other embodiments, the GAN network technology can also be adopted by the machine learning model to restore the texture information in the image, especially in the area where the text content is located, to a certain extent.

[0128] Figure 4 FIG. shows an application scenario schematic diagram of the data processing method according to an embodiment of the present disclosure in screen sharing. It should be noted that, Figure 4 The solid line box in is used to schematically illustrate the method steps executed on each hardware device, and the dashed line box is used to schematically illustrate the execution entity of each method step, that is, the hardware device. As Figure 4 shown, the screen sharing initiator initiates a screen sharing request on the terminal device 401. After the terminal device 401 receives the request, it can obtain the screen data to be shared. The screen data can be an image sequence such as a screen capture or a recorded screen video. The terminal device 401 also extracts the text information from the above image sequence. The text information can include, but is not limited to, the position information of the area where the text content is located, the position change information of the area where the text content is located in consecutive multiple images, and / or the new text content information that appears in the current image, etc. After the terminal device 401 extracts the text information in the image sequence, it sends the text information and the image sequence to the encoding device 402 for encoding.

[0129] During the encoding process, the encoding device 402 determines the region where the text content is located in the current image according to the text information, and then when encoding the region where the text content is located, it can perform encoding according to a preset encoding rule. For example, a smaller quantization parameter QP or a smaller-sized encoding block can be used to encode the image region where the text content is located, so that the image quality of the image region where the text content is located is higher, that is, the compression rate of the image region where the text content is located is reduced, and the encoding quality is improved. In addition, the position change information of the region where the text part is located extracted from the first image sequence (such as the translation of the image content caused by mouse sliding, etc.) can assist the motion estimation efficiency and accuracy of the encoding device during the inter-frame prediction process. This is because the neural network used to build the machine learning model is based on pixel motion calculations, and this method itself is more refined than the motion estimation based on image blocks in the encoding device; in addition, the motion estimation of the encoding device during the inter-frame prediction process has a certain search range limit, while the machine learning model does not have this limitation. Therefore, the machine learning model will be more accurate in extracting the motion information of the image region where the text content is located. Therefore, the text information extracted from the first image sequence can effectively guide the subsequent encoding device to perform more effective encoding.

[0130] After the encoding device 402 encodes the image sequence under the guidance of the above text information to obtain encoded data, after the encoded data is returned, it is sent by the terminal device 401 to the screen sharing recipient.

[0131] After the terminal device 403 of the screen sharing recipient receives the encoded data, it can first use the decoding device 404 to decode the encoded data. Considering that the encoded data obtained by the screen sharing initiator through lossy encoding, the image sequence after decoding the encoded data may have phenomena such as loss of texture details, such as the text becoming blurred. Therefore, the screen sharing recipient performs information enhancement processing on the decoded image sequence so that the text content part in the second image sequence can be made clearer through means such as sharpening.

[0132] The terminal device 403 of the screen sharing recipient displays the enhanced image sequence on the terminal screen.

[0133] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure.

[0134] According to a data processing device of an embodiment of the present disclosure, the device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. The data processing device includes:

[0135] A first acquisition module, configured to acquire a first image sequence; the first image sequence includes at least one image;

[0136] An extraction module, configured to extract text information from the first image sequence; the text information includes metadata of the text content included in the first image sequence;

[0137] A first encoding module, configured to encode the first image sequence based on the text information.

[0138] In an alternative implementation of this embodiment, the extraction module includes:

[0139] An extraction sub-module, configured to extract the text information from the first image sequence by using a machine learning model.

[0140] In an alternative implementation of this embodiment, the metadata includes position information of the region where the text content is located, position change information of the region where the text content is located in consecutive multiple images, and / or information of new text content appearing in the current image.

[0141] In an alternative implementation of this embodiment, the first encoding module includes:

[0142] A first encoding sub-module, configured to provide the first image sequence and the text information to an encoding device, so that the encoding device encodes the text content included in the first image sequence based on the text information.

[0143] In an alternative implementation of this embodiment, the apparatus further includes:

[0144] A first receiving module, configured to receive encoded data;

[0145] A first decoding module, configured to decode the encoded data to obtain a second image sequence;

[0146] A first processing module, configured to perform information enhancement processing on the text content included in the second image sequence.

[0147] The data processing apparatus in this embodiment corresponds to Figure 1 the data processing methods in the illustrated embodiment and related embodiments, and specific details can be referred to the descriptions in the above-mentioned Figure 1 illustrated embodiment and related embodiments, and will not be elaborated here.

[0148] According to another embodiment of the present disclosure, the data processing apparatus can be partially or fully implemented as an electronic device through software, hardware, or a combination of both. The data processing apparatus includes:

[0149] A second receiving module, configured to receive a first image sequence and text information in the first image sequence; the first image sequence includes at least one image; the text information includes metadata of text content included in the first image sequence;

[0150] A second encoding module, configured to encode the first image sequence based on the text information.

[0151] In an alternative implementation of this embodiment, the metadata includes position information of the region where the text content is located, position change information of the region where the text content is located in multiple consecutive images, and / or new text content information that appears in the current image.

[0152] In an alternative implementation of this embodiment, the metadata includes the region range where the text content is located, and the second encoding module includes:

[0153] A second encoding sub-module, configured to encode the images within the region where the text content is located according to preset encoding parameters; the preset encoding parameters include quantization step size.

[0154] In an alternative implementation of this embodiment, the metadata includes position change information of the text content in multiple consecutive images in the first image sequence; the second encoding module includes:

[0155] A first prediction sub-module, configured to perform motion estimation on the images within the region where the text content is located based on the position change information during the inter-frame prediction process.

[0156] In an alternative implementation of this embodiment, the metadata includes new text content information that appears in the current image, and the second encoding module includes:

[0157] A second prediction sub-module, configured to skip motion estimation for the images within the region where the new text content is located during the inter-frame prediction process when there is new text content in the current image.

[0158] The data processing device in this embodiment corresponds to Figure 2 the data processing methods in the illustrated embodiment and related embodiments, and specific details can be referred to the descriptions of the illustrated embodiment and related embodiments above, which will not be elaborated here. Figure 2

[0159] According to a data processing device of another embodiment of the present disclosure, the device can be implemented as part or all of an electronic device through software, hardware, or a combination of both. The data processing device includes:

[0160] ​A second acquisition module, configured to acquire a first image sequence; the first image sequence includes at least one image;

[0161] A first invocation module, configured to invoke a first preset service interface, so that the first preset service interface extracts text information from the first image sequence and encodes the first image sequence based on the text information; the text information includes metadata of the text content included in the first image sequence;

[0162] A first return module, configured to return the encoded data of the first image sequence.

[0163] In an optional implementation manner of this embodiment, the apparatus further includes:

[0164] A third reception module, configured to receive encoded data;

[0165] A second invocation module, configured to invoke a second preset service interface, so that the second preset service interface decodes the encoded data to obtain a second image sequence and performs information enhancement processing on the text content included in the second image sequence;

[0166] A second return module, configured to return the second image sequence after the enhancement processing.

[0167] The data processing apparatus in this embodiment corresponds to Figure 3 the data processing methods in the illustrated embodiment and related embodiments, and specific details can be referred to the descriptions of the Figure 3 illustrated embodiment and related embodiments above, and will not be elaborated here.

[0168] Figure 5 is a schematic structural diagram of an electronic device suitable for implementing the data processing method according to an embodiment of the present disclosure.

[0169] As Figure 5 shown, the electronic device 500 includes a processing unit 501, which can be implemented as a processing unit such as a CPU, GPU, FPGA, NPU, etc. The processing unit 501 can execute various processes in the implementation manners of any of the above methods of the present disclosure according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0170] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 510 as required so that a computer program read out therefrom is installed into the storage section 508 as required.

[0171] Specifically, according to an embodiment of the present disclosure, any of the methods described above with reference to the embodiments of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program tangibly embodied on a machine-readable medium, the computer program including program code for performing any of the methods in the embodiments of the present disclosure. In such an embodiment, the computer program may be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511.

[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the architectures, functions, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0173] The units or modules described in the embodiments of the present disclosure may be implemented in software or in hardware. The units or modules described may also be provided in a processor, and the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases.

[0174] As another aspect, the present disclosure also provides a computer-readable storage medium, which may be the computer-readable storage medium included in the device in the above-described embodiments; or it may exist separately and be a computer-readable storage medium not assembled into the device. The computer-readable storage medium stores one or more programs, and the one or more programs are used by one or more processors to execute the methods described in the present disclosure.

[0175] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present disclosure.

Claims

1. A data processing method, wherein, the data processing method is executed on a terminal device and includes: obtaining a first image sequence; the first image sequence includes at least one image; extracting text information from the first image sequence; the text information includes metadata of the text content included in the first image sequence; encoding the first image sequence based on the text information; wherein, encoding the first image sequence based on the text information includes: providing the first image sequence and the text information to an encoding device, so that the encoding device encodes the text content included in the first image sequence by using the text information as prior information; wherein, extracting the text information from the first image sequence includes: extracting the text information from the first image sequence by using a machine learning model.

2. The method according to claim 1, wherein, the metadata includes position information of the area where the text content is located, position change information of the area where the text content is located in consecutive multiple images, and / or new text content information appearing in the current image.

3. The method according to claim 1, wherein, further includes: receiving encoded data; decoding the encoded data to obtain a second image sequence; performing information enhancement processing on the text content included in the second image sequence.

4. A data processing method, wherein, the data processing method is executed on an encoding device; and includes: receiving a first image sequence and the text information in the first image sequence; the first image sequence includes at least one image; the text information includes metadata of the text content included in the first image sequence; after the text information is extracted from the first image sequence by the terminal device, it is provided to the encoding device together with the first image sequence; encoding the first image sequence by using the text information as prior information.

5. The method according to claim 4, wherein, the metadata includes position information of the area where the text content is located, position change information of the area where the text content is located in consecutive multiple images, and / or new text content information appearing in the current image.

6. The method according to claim 5, wherein, the metadata includes the area range where the text content is located, and encoding the first image sequence based on the text information includes: encoding the image within the area where the text content is located according to preset encoding parameters; the preset encoding parameters include a quantization step size.

7. The method according to any one of claims 5-6, wherein, the metadata includes position change information of the text content in consecutive multiple images in the first image sequence; encoding the first image sequence based on the text information includes: during the inter-frame prediction process, performing motion estimation on the image within the area where the text content is located based on the position change information.

8. The method according to any one of claims 5-6, wherein, The metadata includes information on new text content appearing in the current image. Encoding the first image sequence based on the text information includes: When there is new text content in the current image, during the inter-frame prediction process, skip the motion estimation for the images within the region where the new text content is located.

9. A data processing method, wherein, The data processing method is executed on a terminal device and includes: Obtain a first image sequence; the first image sequence includes at least one image; Invoke a first preset service interface to extract text information from the first image sequence by the first preset service interface and encode the first image sequence based on the text information; the text information includes metadata of the text content included in the first image sequence; wherein, encoding the first image sequence based on the text information includes: providing the first image sequence and the text information to an encoding device so that the encoding device encodes the text content included in the first image sequence using the text information as prior information; extracting the text information from the first image sequence includes: using a machine learning model to extract the text information from the first image sequence; Return the encoded data of the first image sequence.

10. The method according to claim 9, wherein, further includes: Receive encoded data; Invoke a second preset service interface to decode the encoded data by the second preset service interface to obtain a second image sequence and perform information enhancement processing on the text content included in the second image sequence; Return the second image sequence after the enhancement processing.

11. A data processing apparatus, wherein, includes: A first acquisition module configured to acquire a first image sequence; The first image sequence includes at least one image; An extraction module configured to extract text information from the first image sequence; the text information includes metadata of the text content included in the first image sequence; A first encoding module configured to encode the first image sequence based on the text information; wherein, the first encoding module includes: A first encoding sub-module configured to provide the first image sequence and the text information to an encoding device so that the encoding device encodes the text content included in the first image sequence using the text information as prior information; wherein, the extraction module includes: An extraction sub-module configured to use a machine learning model to extract the text information from the first image sequence.

12. A data processing apparatus, wherein, includes: A second receiving module configured to receive a first image sequence and the text information in the first image sequence; The first image sequence includes at least one image; the text information includes metadata of the text content included in the first image sequence; after the text information is extracted from the first image sequence by a terminal device, it is provided to an encoding device together with the first image sequence; A second encoding module configured to encode the first image sequence using the text information as prior information.

13. A data processing device, wherein, comprising: a second acquisition module configured to acquire a first image sequence; the first image sequence includes at least one image; a first call module configured to call a first preset service interface, so that the first preset service interface extracts text information from the first image sequence and encodes the first image sequence based on the text information; the text information includes metadata of the text content included in the first image sequence; a first return module configured to return the encoded data of the first image sequence; wherein, the part of the first call module that encodes the first image sequence based on the text information is configured to: provide the first image sequence and the text information to an encoding device, so that the encoding device uses the text information as prior information to encode the text content included in the first image sequence; the part of the first call module that extracts the text information from the first image sequence is configured to: extract the text information from the first image sequence by using a machine learning model.

14. An electronic device, wherein, comprising a memory and a processor; wherein, the memory is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1-10.

15. A computer-readable storage medium, on which computer instructions are stored, wherein, when the computer instructions are executed by a processor, the method according to any one of claims 1-10 is implemented.

Citation Information

Patent Citations

  • Data compression method and device and data encoding / decoding method and device

    CN109831668A