Image processing method and device, detection model training method and device and electronic equipment

By selecting the matching detection model in the image processing to determine the region of interest and encoding it with different definitions, the problems of image transmission quality and network billing cost are solved, and efficient and quality-assured image transmission is achieved.

CN120201196APending Publication Date: 2025-06-24GUANGZHOU DULING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330738.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the image data transmission, the previous technology usually mechanically encodes the entire image in low definition, making it difficult to ensure the quality of the image transmission, affecting the user's visual perception, and increasing network billing costs.

Method used

By acquiring the image to be processed, selecting the object detection model that matches it, determining the region of interest in the image, and encoding the region with high definition, while encoding other regions with low definition, generating the target image and sending it.

Benefits of technology

This method can not only appropriately reduce network billing costs, but also ensure the image transmission quality of the area of ​​interest, thereby ensuring the overall transmission quality of the images to be processed and avoiding affecting the user's perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201196A_ABST
    Figure CN120201196A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, a detection model training method and device and electronic equipment, and relates to the technical field of computers, in particular to the application fields of image processing, video processing, audio and video transmission, online games, digital humans, machine learning and the like. The specific implementation scheme is as follows: acquiring a to-be-processed image; selecting a target detection model matched with the to-be-processed image from the plurality of candidate detection models; determining a region of interest in the to-be-processed image by using the target detection model; performing high-definition coding on the region of interest in the to-be-processed image, and performing low-definition coding on the remaining region except the region of interest in the to-be-processed image to obtain a target image corresponding to the to-be-processed image; and sending the target image to the target device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to application fields such as image processing, video processing, audio-video transmission, online games, digital humans, machine learning, etc. Specifically, it relates to an image processing method, a training method of a detection model, a device, and an electronic device. Background Art

[0002] When an application program with a "client-server" architecture runs on a server, it often needs to interact with a terminal device installed with an application program client. For example, sending image data to the terminal device. Moreover, before these image data are sent to the terminal device, they usually need to be encoded by a service device. Summary of the Invention

[0003] The present disclosure provides an image processing method, a training method of a detection model, a device, and an electronic device.

[0004] According to a first aspect of the present disclosure, there is provided an image processing method, including:

[0005] Obtaining an image to be processed;

[0006] Selecting a target detection model that matches the image to be processed from multiple candidate detection models;

[0007] Using the target detection model to determine a region of interest in the image to be processed;

[0008] Performing high-definition encoding on the region of interest in the image to be processed, and performing low-definition encoding on the remaining region in the image to be processed except for the region of interest, to obtain a target image corresponding to the image to be processed;

[0009] Sending the target image to a target device.

[0010] According to a second aspect of the present disclosure, there is provided a training method of a detection model, including:

[0011] Based on different image data sources, constructing multiple training data sets; wherein, each training data set in the multiple training data sets includes a sample image and identification information for characterizing the region of interest in the sample image;

[0012] Using the multiple training data sets to train an initial detection model to obtain multiple candidate detection models corresponding one-to-one to the multiple training data sets.

[0013] According to a third aspect of the present disclosure, there is provided an image processing device, including:

[0014] An image acquisition unit for obtaining an image to be processed;

[0015] A model selection unit, configured to select a target detection model that matches the image to be processed from multiple candidate detection models;

[0016] A region determination unit, configured to use the target detection model to determine a region of interest in the image to be processed;

[0017] An image encoding unit, configured to perform high-definition encoding on the region of interest in the image to be processed, and perform low-definition encoding on the remaining region of the image to be processed except the region of interest, to obtain a target image corresponding to the image to be processed;

[0018] An image sending unit, configured to send the target image to a target device

[0019] According to a fourth aspect of the present disclosure, there is provided a training device for a detection model, including:

[0020] A dataset construction unit, configured to construct multiple training datasets based on different image data sources; wherein, each training dataset in the multiple training datasets includes a sample image and identification information for characterizing the region of interest in the sample image;

[0021] A model training unit, configured to use the multiple training datasets to train an initial detection model to obtain multiple candidate detection models corresponding one by one to the multiple training datasets.

[0022] According to a fifth aspect of the present disclosure, there is provided an electronic device, including:

[0023] At least one processor;

[0024] A memory communicatively connected to the at least one processor;

[0025] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method provided in the first aspect of the present disclosure.

[0026] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method provided in the first aspect of the present disclosure.

[0027] According to a seventh aspect of the present disclosure, there is provided a computer program product, including a computer program, which implements the method provided in the first aspect of the present disclosure when executed by a processor.

[0028] Adopting the present disclosure can not only appropriately reduce the network billing cost, but also ensure the image transmission quality of the key image region, that is, the region of interest, so as to ensure the overall transmission quality of the image to be processed and avoid affecting the user experience.

[0029] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0031] Figure 1 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure;

[0032] Figure 2 is a schematic flowchart of a training method for a detection model provided by an embodiment of the present disclosure;

[0033] Figure 3 is a schematic partial interaction flowchart of a target program provided by an embodiment of the present disclosure;

[0034] Figure 4 is a schematic application scenario diagram of an image processing method provided by an embodiment of the present disclosure;

[0035] Figure 5 is a schematic application scenario diagram of a training method for a detection model provided by an embodiment of the present disclosure;

[0036] Figure 6 is a schematic structural block diagram of an image processing device provided by an embodiment of the present disclosure;

[0037] Figure 7 is a schematic structural block diagram of a training device for a detection model provided by an embodiment of the present disclosure;

[0038] Figure 8 is a schematic structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0040] As described above, when an application with a "client-server" architecture runs on a server, it often needs to interact with a terminal device installed with the application client. For example, it sends image data to the terminal device. Moreover, before the image data is sent to the terminal device, it usually needs to be encoded by the server. However, the inventors have found that currently, when the server encodes image data (hereinafter referred to as the image to be processed), it usually mechanically encodes the entire image to be processed with low clarity. In this way, although the network billing cost can be minimized to the greatest extent, it is difficult to ensure the transmission quality of the image data, thus affecting the user experience.

[0041] In view of the above problems, embodiments of the present disclosure provide an image processing method, which can be applied to an electronic device. Among them, the electronic device can be a service device, specifically, it can be a server, a workbench, a mainframe computer or other similar computing devices. Hereinafter, Figure 1 in conjunction with the

[0042] flow chart shown, an image processing method provided by embodiments of the present disclosure will be described. It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described in the flow chart can also be executed in other orders.

[0043] Step S101, obtain the image to be processed.

[0044] Among them, the image to be processed can be the image data that needs to be sent to the target device installed with the target program client when the target program runs on the service device. Among them, the target program can be a specified game program, a specified live program, a specified monitoring program, etc.; the target device can be a terminal device, specifically, it can be a conventional computer (for example, a desktop computer, a laptop computer, a tablet computer, etc.), a smart phone, a personal digital processor or other similar computing devices.

[0045] Among them, the multiple candidate detection models can be pre-trained, and the training process can be: based on different image data sources, multiple training data sets are constructed, and the initial detection model is trained using the multiple training data sets to obtain multiple candidate detection models corresponding to the multiple training data sets one by one. Among them, when the target program is a specified game program, different image data sources can be different game programs; when the target program is a specified live program, different image data sources can be different types of live programs; when the target program is a specified monitoring program, different image data sources can be different types of monitoring programs.

[0046] In addition, it should be noted that in the embodiments of the present disclosure, the initial detection model may be a single-stage detector (YouOnly Look Once, YOLO), a single-stage multi-box detector (Single Shot MultiBox Detector, SSD), a two-stage detector (Faster Region-based Convolutional Neural Networks, Faster R-CNN), a detection transformer (Detection Transformer, DETR), etc.

[0047] After obtaining multiple candidate detection models, a target detection model that matches the image to be processed can be selected from the multiple candidate detection models. For example, based on the data source of the image to be processed (i.e., the target program), the target detection model can be selected from the multiple candidate models.

[0048] Step S103: Use the target detection model to determine the region of interest in the image to be processed.

[0049] Among them, the region of interest in the image to be processed is used to characterize the key image region in the image to be processed. For example, when the target program is a specified game program, the region of interest in the image to be processed may be game characters, monsters, treasures, etc.; for another example, when the target program is a specified live broadcast program, the region of interest in the image to be processed may be the anchor, co-anchor, sold goods, etc.; for yet another example, when the target program is a specified monitoring program, the region of interest in the image to be processed may be people, animals, items, etc.

[0050] Step S104: Perform high-definition encoding on the region of interest in the image to be processed, and perform low-definition encoding on the remaining region in the image to be processed except the region of interest, to obtain a target image corresponding to the image to be processed.

[0051] In the embodiments of the present disclosure, a preset coding algorithm (such as H.264, H.265, etc.) can be used to perform high-definition coding on the region of interest in the image to be processed, and perform low-definition coding on the remaining region in the image to be processed except the region of interest, so as to obtain a target image corresponding to the image to be processed. For example, a preset coding algorithm can be used to perform high-definition coding on the region of interest in the image to be processed according to a first compression ratio, and perform low-definition coding on the remaining region in the image to be processed except the region of interest according to a second compression ratio, so as to obtain a target image corresponding to the image to be processed. Wherein, the first compression ratio is less than the second compression ratio to ensure that the data richness of the first compressed region obtained after performing high-definition coding on the region of interest in the image to be processed according to the first compression ratio is greater than the data richness of the second compressed region obtained after performing low-definition coding on the remaining region in the image to be processed except the region of interest according to the second compression ratio.

[0052] Step S105, send the target image to the target device.

[0053] By using the image processing method provided in the embodiments of the present disclosure, after obtaining the image to be processed, a target detection model matching the image to be processed can be selected from multiple candidate detection models, and the target detection model can be used to determine the region of interest in the image to be processed, and then high-definition coding is performed on the region of interest in the image to be processed, and low-definition coding is performed on the remaining region in the image to be processed except the region of interest, so as to obtain a target image corresponding to the image to be processed, and the target image is sent to the target device. In this process, the key lies in "performing high-definition coding on the region of interest in the image to be processed, and performing low-definition coding on the remaining region in the image to be processed except the region of interest, so as to obtain a target image corresponding to the image to be processed". In this way, not only can the network charging cost be appropriately reduced, but also the image transmission quality of the key image region, that is, the region of interest, can be ensured, so as to ensure the overall transmission quality of the image to be processed and avoid affecting the user experience. Moreover, in the embodiments of the present disclosure, a target detection model matching the image to be processed is selected from multiple candidate detection models, and the target detection model is used to determine the region of interest in the image to be processed, rather than directly using a general detection model to determine the region of interest in the image to be processed. Therefore, the reliability of the region of interest can be ensured.

[0054] For the convenience of further explaining some steps in the image processing method, hereinafter, the acquisition method of multiple candidate detection models will be described in more detail. Please refer to Figure 2, embodiments of the present disclosure provide a method for training a detection model, which can be applied to an electronic device. Among them, the electronic device can be a server, a workbench, a mainframe computer, a conventional computer (such as a desktop computer, a laptop computer, a tablet computer, etc.) or other similar computing devices. Hereinafter, in combination with Figure 2 The flowchart shown will illustrate a method for training a detection model provided by an embodiment of the present disclosure. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described in the flowchart can also be executed in other orders.

[0055] Step S201, construct multiple training data sets based on different image data sources.

[0056] Among them, when the target program is a specified game program, different image data sources can be different game programs; when the target program is a specified live program, different image data sources can be different types of live programs; when the target program is a specified monitoring program, different image data sources can be different types of monitoring programs.

[0057] In addition, it should be noted that in the embodiments of the present disclosure, each training data set in the multiple training data sets may include a sample image and identification information for characterizing the region of interest in the sample image. Here, the identification information can be a location identifier, for example, coordinate information, and the embodiments of the present disclosure do not make specific limitations on this.

[0058] Step S202, use the multiple training data sets to train the initial detection model to obtain multiple candidate detection models corresponding one by one to the multiple training data sets.

[0059] Among them, the initial detection model can be YOLO, SSD, Faster R-CNN, DETR, etc.

[0060] In one example, each training data set in the multiple training data sets can be used as a target data set, obtain the current sample image from the target data set, and determine the current identification information for characterizing the region of interest in the current sample image. Then, use the initial detection model to determine the region of interest in the current sample image, which can be specifically characterized by speculative identification information, and then train the initial detection model based on the information loss between the current identification information and the speculative identification information to obtain the trained initial detection model. Among them, the speculative identification information can be speculative location identification, for example, speculative coordinate information, and the embodiments of the present disclosure do not make specific limitations on this.

[0061] In addition, it should be noted that in the embodiments of the present disclosure, when the trained initial detection model meets the preset convergence condition, the trained initial detection model can be used as a candidate detection model corresponding to the target data set. The preset convergence condition can be set according to actual application requirements, and the embodiments of the present disclosure do not limit this.

[0062] By using the training method of the detection model provided by the embodiments of the present disclosure, multiple training data sets can be constructed based on different image data sources, and the initial detection model can be trained using the multiple training data sets to obtain multiple candidate detection models corresponding to the multiple training data sets one by one. In this way, when performing the image processing method, the target detection model that matches the image to be processed can be selected from the multiple candidate detection models, and the target detection model can be used to determine the region of interest in the image to be processed, rather than directly using a general detection model to determine the region of interest in the image to be processed. Therefore, the reliability of the region of interest can be ensured.

[0063] In some alternative embodiments, step S201, that is, "constructing multiple training data sets based on different image data sources" may include:

[0064] Constructing N first training data sets based on N popular games;

[0065] Constructing a second training data set based on a game collection;

[0066] Taking the N first training data sets and the second training data set as multiple training data sets.

[0067] Where N≥2 and is an integer.

[0068] In addition, it should be noted that in the embodiments of the present disclosure, all game programs on the program management platform connected to the service device can be determined, and the top N game programs with the highest heat index can be determined from all game programs as the N popular games. The heat index can be a single index such as the download volume, usage frequency, usage duration, number of active users, etc., or a composite index obtained based on at least one of the single indices such as the download volume, usage frequency, usage duration, number of active users, etc.

[0069] After determining N popular games, N first training data sets can be constructed based on the N popular games. Specifically, each of the N popular games can be used as the current popular game, and screenshots of different game interfaces can be taken for the current popular game to obtain multiple first sample images, and first identification information of each first sample image among the multiple first sample images can be obtained. For example, the first identification information of each first sample image among the multiple first sample images can be obtained by means of manual annotation. Among them, the first identification information can be a first position identification. For example, the first coordinate information is not specifically limited in the embodiments of the present disclosure. Based on this, it can be understood that in the embodiments of the present disclosure, the first training data set corresponding to the current popular game can include multiple first sample images and the first identification information of each first sample image among the multiple first sample images.

[0070] Furthermore, in the embodiments of the present disclosure, the game collection can include multiple game programs.

[0071] In one example, "constructing a second training data set based on the game collection" can include: using the N popular games as the game collection to construct a second training data set based on the game collection. Specifically, after constructing N first training data sets based on the N popular games, the N first training data sets can be merged into a second training data set.

[0072] In the above example, finally, N first training data sets and one second training data set can be obtained as multiple training data sets. Then, when using the multiple training data sets to train the initial detection model, N first candidate models corresponding to the N first training data sets one by one and a second candidate model corresponding to the second training data set can be obtained.

[0073] In another example, "constructing a second training data set based on the game collection" can include: using multiple games of the same type under the target game type as the game collection to construct a second training data set based on the game collection. Among them, the target game type is used to represent each of the M game types, and M≥2 and is an integer. Here, the M game types can include cultivation types, casual types, real-time strategy types, sports types, clearance types, etc.

[0074] In a specific example, when constructing the second training dataset based on the game collection, each game of the same type in the game collection can be used as the current type of game, and screenshots of different game interfaces are taken for the current type of game to obtain multiple second sample images, and second identification information of each second sample image among the multiple second sample images is obtained. For example, the second identification information of each second sample image among the multiple second sample images can be obtained by means of manual annotation. The second identification information can be a second position identifier. For example, it can be second coordinate information, and the embodiments of the present disclosure do not make specific limitations on this.

[0075] In the above example, finally, N first training datasets and M second training datasets corresponding one-to-one to M game types can be obtained as multiple training datasets. Then, when using the multiple training datasets to train the initial detection model, N first candidate models corresponding one-to-one to the N first training datasets and M second candidate models corresponding one-to-one to the M second training datasets can be obtained.

[0076] In the above manner, in the embodiments of the present disclosure, N first training datasets can be constructed based on N popular games, and the second training dataset can be constructed based on the game collection. Then, the N first training datasets and the second training dataset are used as multiple training datasets to ensure the richness and diversity of the training datasets. In this way, after using the multiple training datasets to train the initial detection model to obtain multiple candidate detection models corresponding one-to-one to the multiple training datasets, when performing the image processing method, a target detection model with a high degree of matching with the image to be processed can be selected from the multiple candidate detection models, so that when using the target detection model to determine the region of interest in the image to be processed, the reliability of the region of interest can be ensured. Moreover, when constructing the second training dataset based on the game collection, the N popular games can be used as the game collection to construct the second training dataset, or multiple games of the same type under the target game type can be used as the game collection to construct the second training dataset. In this way, after using the multiple training datasets to train the initial detection model to obtain multiple candidate detection models corresponding one-to-one to the multiple training datasets, when performing the image processing method, even if the data source of the image to be processed is not one of the N popular games, a target detection model with a high degree of matching with the image to be processed can be selected from the multiple candidate detection models, so that when using the target detection model to determine the region of interest in the image to be processed, the region of interest can be accurately locked, that is, the reliability of the region of interest can be further improved.

[0077] Next, some steps in the image processing method will be further described.

[0078] First of all, it should be noted that in the embodiments of the present disclosure, the image to be processed can be independent image data or video images carried in a video data stream. For example, the image to be processed can be each video image carried in the video data stream, or can be partial video images carried in the video data stream. Based on this, when the image to be processed is a video image carried in a video data stream, step S101, that is, "obtain the image to be processed" can include:

[0079] Determine the image detection frequency based on the action interval duration and / or network latency;

[0080] Extract video images from the video data stream as the images to be processed based on the image detection frequency.

[0081] Among them, the action interval duration is used to represent the interval duration between the trigger time point of the most recent user action and the current time point. Here, the most recent user action can be interaction actions such as clicks and swipes recently triggered by the user on the target device (that is, the terminal device) installed with the target program client.

[0082] In addition, it should be noted that in the embodiments of the present disclosure, when the user triggers an interaction action on the target device, an interaction control instruction corresponding to the interaction action will be generated and sent to the service device, so that the service device generates the latest video data stream based on the interaction instruction.

[0083] Exemplarily, when the target program is a specified game program, the user triggers an interaction action on the target device, generates an interaction control instruction corresponding to the interaction action, and then sends the interaction control instruction to the service device. Please refer to Figure 3, after receiving the interaction control instruction, the service device can adjust the game interface based on the interaction instruction (for example, move the game character, change the game character's action, etc.) to obtain the latest display data of the target program, and generate the latest video data stream based on the latest display data. Thereafter, the service device can determine the image detection frequency based on the action interval duration and / or network latency, and extract video images from the video data stream as the images to be processed based on the image detection frequency. Then, the image processing method is executed, that is, after obtaining the images to be processed, select the target detection model that matches the images to be processed from multiple candidate detection models, and use the target detection model to determine the region of interest in the images to be processed. Then, perform high-definition encoding on the region of interest in the images to be processed, and perform low-definition encoding on the remaining region other than the region of interest in the images to be processed to obtain the target image corresponding to the images to be processed, and send the target image to the target device. Among them, sending the target image to the target device can be: inserting the target image into the target position in the video data stream to obtain the updated video data stream, and sending the updated video data stream to the target device. Here, the target position can be the original position where the video image corresponding to the target image is located in the video data stream.

[0084] In addition, it should be noted that in the embodiments of the present disclosure, for the remaining video images in the video data stream that are not extracted as the images to be processed, high-definition encoding of the entire image or low-definition encoding of the entire image can be performed, which can be specifically set according to application requirements, and the embodiments of the present disclosure do not make specific limitations on this.

[0085] Finally, when the target device receives the updated video data stream, it can perform a decoding operation on the updated video data stream to obtain multiple images to be displayed, and display the multiple images to be displayed on the screen of the target device.

[0086] In addition, it should be noted that in the embodiments of the present disclosure, the network latency can be the network delay duration between the service device and the target device.

[0087] After obtaining the action interval duration and / or network latency, an image detection frequency that is inversely correlated with the action interval duration can be set (that is, the shorter the action interval duration, the higher the set image detection frequency; the longer the action interval duration, the lower the set image detection frequency), or an image detection frequency that is positively correlated with the network latency can be set (that is, the longer the network latency, the higher the set image detection frequency; the shorter the network latency, the lower the set image detection frequency).

[0088] In the above manner, in the embodiments of the present disclosure, a reasonable image detection frequency can be determined based on the action interval duration and / or network delay, and video images can be extracted from the video data stream based on the image detection frequency as the images to be processed. In this way, excessive computing resources of the service device can be avoided.

[0089] In one example, "determining the image detection frequency based on the action interval duration and / or network delay" may include one of the following three:

[0090] (1) When the action interval duration is within a preset short-duration interval, the first detection frequency is used as the image detection frequency; or, when the action interval duration is not within the preset short-duration interval, the second detection frequency is used as the image detection frequency.

[0091] Among them, the preset short-duration interval can be (0, X1), and X1 can be set according to actual application requirements, and the embodiments of the present disclosure do not limit this; the second detection frequency is lower than the first detection frequency.

[0092] In a specific example, when the action interval duration is within the preset short-duration interval (0, X1), the first detection frequency can be used as the image detection frequency. Correspondingly, when the action interval duration is not within the preset short-duration interval (0, X1), for example, within the duration interval [X1, +∞), the second detection frequency can be used as the image detection frequency. Among them, the second detection frequency is lower than the first detection frequency. For example, the second detection frequency can be 1 / 5, that is, one video image is extracted from every five video images in the video data stream as the image to be processed; the first detection frequency can be 1 / 2, that is, one video image is extracted from every two video images in the video data stream as the image to be processed.

[0093] (2) When the network delay is within a preset high-delay interval, the third detection frequency is used as the image detection frequency; or, when the network delay is not within the preset high-delay interval, the fourth detection frequency is used as the image detection frequency.

[0094] Among them, the preset high-delay interval can be (X2, +∞), and X2 can be set according to actual application requirements, and the embodiments of the present disclosure do not limit this; the fourth detection frequency is lower than the third detection frequency.

[0095] In a specific example, when the network latency is in the preset high-latency interval (X2, +∞), the third detection frequency can be used as the image detection frequency. Correspondingly, when the network latency is not in the preset high-latency interval (X2, +∞), for example, when it is in the duration interval (0, X2], the fourth detection frequency can be used as the image detection frequency. Among them, the fourth detection frequency is lower than the third detection frequency. For example, the fourth detection frequency can be 1 / 5, that is, one video image is extracted from every five video images in the video data stream as the image to be processed; the third detection frequency can be 1 / 2, that is, one video image is extracted from every two video images in the video data stream as the image to be processed.

[0096] (3) When the action interval duration is in the preset short-duration interval and the network latency is in the preset high-latency interval, the fifth detection frequency is used as the image detection frequency; or, when only one of the action interval duration being in the preset short-duration interval and the network latency being in the preset high-latency interval is true, the sixth detection frequency is used as the image detection frequency; or, when the action interval duration is not in the preset short-duration interval and the network latency is not in the preset high-latency interval, the seventh detection frequency is used as the image detection frequency.

[0097] Among them, the preset short-duration interval can be (0, X3), the preset high-latency interval can be (X4, +∞), and X3 and X4 can be set according to actual application requirements, and the embodiments of the present disclosure do not limit this; the seventh detection frequency is less than the sixth detection frequency, and the sixth detection frequency is less than the fifth detection frequency.

[0098] In a specific example, the fifth detection frequency can be used as the image detection frequency when the action interval duration is within a preset short-duration range (0, X3) and the network latency is within a preset high-latency range (X4, +∞); or, the sixth detection frequency can be used as the image detection frequency when only one of the action interval duration within the preset short-duration range (0, X3) and the network latency within the preset high-latency range (X4, +∞) is true; or, the seventh detection frequency can be used as the image detection frequency when the action interval duration is not within the preset short-duration range (0, X3), for example, within the duration range [X3, +∞), and the network latency is not within the preset high-latency range (X4, +∞), for example, within the duration range (0, X4]. Among them, the seventh detection frequency is less than the sixth detection frequency, and the sixth detection frequency is less than the fifth detection frequency. For example, the seventh detection frequency can be 1 / 5, that is, one video image is extracted from every five video images in the video data stream as the image to be processed; the sixth detection frequency can be 1 / 3, that is, one video image is extracted from every three video images in the video data stream as the image to be processed; the first detection frequency can be 1, that is, each video image in the video data stream is used as the image to be processed.

[0099] In the above examples, a simple logic can be used to determine a reasonable image detection frequency based on the action interval duration and / or network latency, thereby improving the acquisition efficiency of the image detection frequency and the execution efficiency of the image processing method.

[0100] In an alternative embodiment, step S102, that is, "select a target detection model that matches the image to be processed from multiple candidate detection models" can include:

[0101] Obtain N first candidate models corresponding to N popular games one by one;

[0102] Obtain a second candidate model corresponding to the game collection;

[0103] Select a target detection model from multiple candidate detection models including N first candidate models and the second candidate model.

[0104] Where N ≥ 2 and is an integer; the first target model is obtained by training an initial detection model using a first training data set constructed based on the target popular game; the first target model is used to represent each of the N first candidate models; the target popular game is the popular game corresponding to the first target model among the N popular games.

[0105] In addition, it should be noted that in the embodiments of the present disclosure, the N popular games are determined from all game programs on the program management platform connected to the service device, and are the top N game programs with the highest heat index. Among them, the heat index can be a single index such as the download volume, usage frequency, usage duration, number of active users, etc., or a composite index obtained based on at least one of the single indexes such as the download volume, usage frequency, usage duration, number of active users, etc.

[0106] Through the above method, in the embodiments of the present disclosure, N first candidate models corresponding to the N popular games can be obtained, and a second candidate model corresponding to the game collection can be obtained. Then, a target detection model is selected from the multiple candidate detection models including the N first candidate models and the second candidate model to ensure that the target detection model has a high matching degree with the image to be processed, so that when using the target detection model to determine the region of interest in the image to be processed, the region of interest can be accurately locked, that is, the reliability of the region of interest can be improved.

[0107] Furthermore, in the embodiments of the present disclosure, the game collection may include multiple game programs.

[0108] In one example, "obtaining the second candidate model corresponding to the game collection" may include: obtaining the second candidate model obtained by using the second training data set constructed based on the game collection to train the initial detection model with the N popular games as the game collection. Among them, the second training data set may be a training data set obtained by merging the N first training data sets corresponding to the N popular games one by one.

[0109] Continuing with the above example, when selecting the target detection model from the multiple candidate detection models including the N first candidate models and the second candidate model, the data source of the image to be processed (that is, the target program) can be determined. When the data source is one of the N popular games, the first candidate model corresponding to the data source is selected from the N first candidate models as the target detection model; or when the data source is not one of the N popular games, the second candidate model is used as the target detection model. During this process, even if the data source of the image to be processed is not one of the N popular games, a target detection model with a high matching degree with the image to be processed can be obtained, so that when using the target detection model to determine the region of interest in the image to be processed, the region of interest can be accurately locked, that is, the reliability of the region of interest can be further improved.

[0110] In another example, "obtaining a second candidate model corresponding to a game collection" may include: obtaining M second candidate models that correspond one-to-one to M game types. Here, M≥2 and is an integer; the second target model is obtained by using multiple games of the same type under the target game type as a game collection to train an initial detection model with a second training data set constructed based on the game collection; the second target model is used to represent each of the M second candidate models; the target game type is the game type corresponding to the second target model among the M game types. Here, the M game types may include cultivation type, casual type, real-time strategy type, sports type, clearance type, etc.

[0111] Continuing with the above example, when selecting a target detection model from multiple candidate detection models including N first candidate models and second candidate models, the data source (i.e., the target program) of the image to be processed may be determined, and when the data source is one of the N popular games, the first candidate model corresponding to the data source is selected from the N first candidate models as the target detection model; or, when the data source is not one of the N popular games, the second candidate model corresponding to the data source is selected from the M second candidate models as the target detection model. Specifically, the game type to which the data source belongs may be determined from the M game types as the type to be queried, and the second candidate model corresponding to the type to be queried is determined from the M second candidate models as the target detection model. In this process, even if the data source of the image to be processed is not one of the N popular games, a target detection model with a high degree of matching with the image to be processed can be selected from multiple candidate detection models (specifically, from the M second candidate models), so that when using the target detection model to determine the region of interest in the image to be processed, the region of interest can be accurately locked, that is, the reliability of the region of interest can be further improved.

[0112] Please refer to Figure 4 , which is a schematic diagram of an application scenario of an image processing method provided by an embodiment of the present disclosure.

[0113] The image processing method provided by an embodiment of the present disclosure is applied to an electronic device. Here, the electronic device may be a service device, specifically, a server, a workbench, a mainframe computer, or other similar computing devices.

[0114] Here, the electronic device is used to:

[0115] Obtain an image to be processed;

[0116] Select a target detection model that matches the image to be processed from multiple candidate detection models;

[0117] Use the target detection model to determine the region of interest in the image to be processed;

[0118] Perform high-definition encoding on the region of interest in the image to be processed, and perform low-definition encoding on the remaining region in the image to be processed except for the region of interest, to obtain a target image corresponding to the image to be processed;

[0119] Send the target image to the target device.

[0120] It should be noted that in the embodiments of the present disclosure, Figure 4 The schematic diagram of the application scenario shown is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 4 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0121] Please refer to Figure 5 , which is a schematic diagram of the application scenario of a training method for a detection model provided by the embodiments of the present disclosure.

[0122] The training method for the detection model provided by the embodiments of the present disclosure is applied to an electronic device. Among them, the electronic device can be a server, a workbench, a mainframe computer, a conventional computer (such as a desktop computer, a laptop computer, a tablet computer, etc.) or other similar computing devices.

[0123] Here, the electronic device is used for:

[0124] Based on different image data sources, construct multiple training data sets; among them, each training data set in the multiple training data sets includes a sample image and identification information for characterizing the region of interest in the sample image;

[0125] Use the multiple training data sets to train an initial detection model to obtain multiple candidate detection models corresponding to the multiple training data sets one by one.

[0126] It should be noted that in the embodiments of the present disclosure, Figure 5 The schematic diagram of the application scenario shown is only illustrative and not restrictive. Those skilled in the art can make various obvious changes and / or substitutions based on Figure 5 the examples, and the obtained technical solutions still fall within the scope of the disclosure of the embodiments of the present disclosure.

[0127] To better implement the foregoing method, the embodiments of the present disclosure further provide an image processing apparatus, which can be integrated into an electronic device. Among them, the electronic device can be a service device, specifically a server, a workbench, a mainframe computer or other similar computing devices. Hereinafter, with reference to Figure 6 the schematic structural block diagram shown, an image processing apparatus 600 provided by the embodiments of the disclosure will be described.

[0128] The image processing device 600 includes:

[0129] An image acquisition unit 601 for acquiring an image to be processed;

[0130] A model selection unit 602 for selecting a target detection model that matches the image to be processed from multiple candidate detection models;

[0131] A region determination unit 603 for determining a region of interest in the image to be processed by using the target detection model;

[0132] An image encoding unit 604 for performing high-definition encoding on the region of interest in the image to be processed and performing low-definition encoding on the remaining region other than the region of interest in the image to be processed, to obtain a target image corresponding to the image to be processed;

[0133] An image sending unit 605 for sending the target image to a target device.

[0134] In some alternative embodiments, the model selection unit 602 is configured to:

[0135] Obtain N first candidate models respectively corresponding to N popular games; where N≥2 and is an integer; the first target model is obtained by training an initial detection model by using a first training data set constructed based on the target popular game; the first target model is used to represent each of the N first candidate models; the target popular game is the popular game corresponding to the first target model among the N popular games;

[0136] Obtain a second candidate model corresponding to a game collection;

[0137] Select a target detection model from multiple candidate detection models including the N first candidate models and the second candidate model.

[0138] In some alternative embodiments, the model selection unit 602 is configured to:

[0139] Obtain a second candidate model obtained by using the N popular games as a game collection to train an initial detection model by using a second training data set constructed based on the game collection.

[0140] In some alternative embodiments, the model selection unit 602 is configured to:

[0141] Determine the data source of the image to be processed;

[0142] In the case where the data source is one of the N popular games, select the first candidate model corresponding to the data source from the N first candidate models as the target detection model;

[0143] Alternatively, in the case where the data source is not one of the N popular games, the second candidate model is used as the target detection model.

[0144] In some alternative embodiments, the model selection unit 602 is configured to:

[0145] Obtain M second candidate models corresponding one-to-one to M game types; where M≥2 and is an integer; the second target model is obtained by using multiple games of the same type under the target game type as a game collection and training an initial detection model with a second training data set constructed based on the game collection; the second target model is used to represent each of the M second candidate models; the target game type is the game type corresponding to the second target model among the M game types.

[0146] In some alternative embodiments, the model selection unit 602 is configured to:

[0147] Determine the data source of the image to be processed;

[0148] In the case where the data source is one of the N popular games, select the first candidate model corresponding to the data source from the N first candidate models as the target detection model;

[0149] Alternatively, in the case where the data source is not one of the N popular games, select the second candidate model corresponding to the data source from the M second candidate models as the target detection model.

[0150] In some alternative embodiments, the image acquisition unit 601 is configured to:

[0151] Determine the image detection frequency based on the action interval duration and / or network latency; where the action interval duration is used to represent the interval duration between the trigger time point of the most recent user action and the current time point;

[0152] Extract video images from the video data stream as the images to be processed based on the image detection frequency.

[0153] In some alternative embodiments, the image acquisition unit 601 is configured to:

[0154] In the case where the action interval duration is within a preset short duration interval, use the first detection frequency as the image detection frequency;

[0155] Alternatively, in the case where the action interval duration is not within the preset short duration interval, use the second detection frequency as the image detection frequency; where the second detection frequency is lower than the first detection frequency.

[0156] In some alternative embodiments, the image acquisition unit 601 is configured to:

[0157] When the network delay is within a preset high-delay interval, the third detection frequency is used as the image detection frequency;

[0158] Alternatively, when the network delay is not within the preset high-delay interval, the fourth detection frequency is used as the image detection frequency; wherein, the fourth detection frequency is lower than the third detection frequency.

[0159] In the embodiments of the present disclosure, for the specific functions and examples of each unit in the image processing device 600, reference may be made to the relevant descriptions of the corresponding steps in the foregoing embodiments of the image processing method, which will not be elaborated herein.

[0160] To better implement the foregoing method for training a detection model, the embodiments of the present disclosure further provide a training device for a detection model, which may be integrated into an electronic device. Among them, the electronic device may be a server, a workbench, a mainframe computer, a conventional computer (for example, a desktop computer, a laptop computer, a tablet computer, etc.) or other similar computing devices. Hereinafter, with reference to Figure 7 the following schematic structural block diagram, a training device 700 for a detection model provided by the embodiments of the present disclosure will be described.

[0161] The training device 700 for a detection model includes:

[0162] A data set construction unit 701, configured to construct a plurality of training data sets based on different image data sources; wherein, each training data set in the plurality of training data sets includes a sample image and identification information for characterizing the region of interest in the sample image;

[0163] A model training unit 702, configured to use the plurality of training data sets to train an initial detection model to obtain a plurality of candidate detection models corresponding one by one to the plurality of training data sets.

[0164] In some optional embodiments, the data set construction unit 701 is configured to:

[0165] Construct N first training data sets based on N popular games; wherein, N≥2 and is an integer;

[0166] Construct a second training data set based on a game collection;

[0167] Use the N first training data sets and the second training data set as the plurality of training data sets.

[0168] In some optional embodiments, the data set construction unit 701 is configured to:

[0169] Use the N popular games as a game collection to construct a second training data set based on the game collection;

[0170] Alternatively, multiple games of the same type under the target game type are used as a game collection to construct a second training data set based on the game collection; wherein the target game type is used to represent each of the M game types; M≥2 and is an integer.

[0171] In the embodiments of the present disclosure, for the specific functions and examples of each unit in the training device 700 of the detection model, reference may be made to the relevant descriptions of the corresponding steps in the foregoing embodiments of the detection model training method, which will not be elaborated herein.

[0172] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0173] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0174] Figure 8 FIG. shows a schematic structural block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure. The electronic device 800 is intended to represent various forms of digital computers, such as in-vehicle computing devices, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 800 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0175] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0176] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of renderers, speakers, etc.; a storage unit 808, such as a disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0177] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the training method of the detection model and / or the image processing method. For example, in some embodiments, the training method of the detection model and / or the image processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training method of the detection model and / or the image processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform the training method of the detection model and / or the image processing method in any other suitable manner (e.g., by means of firmware).

[0178] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0179] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0180] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0181] For providing interaction with a user, the systems and techniques described herein can be implemented on a computer having: a rendering device for rendering information to the user (e.g., a cathode ray tube (CRT) renderer or a liquid crystal display (LCD) renderer); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0182] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0183] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0184] Embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute a training method of a detection model and / or an image processing method.

[0185] Embodiments of the present disclosure also provide a computer program product, including a computer program which, when executed by a processor, implements a training method of a detection model and / or an image processing method.

[0186] It should be understood that various forms of the processes shown above may be used, steps may be reordered, added or deleted. For example, the steps recited in the present disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and this is not limited herein. In addition, in the present disclosure, relational terms such as "first", "second", "third", etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In addition, "a plurality" in the present disclosure can be understood as at least two.

[0187] The foregoing specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. An image processing method, comprising: Get the image to be processed; Selecting a target detection model that matches the image to be processed from multiple candidate detection models; Determine a region of interest in the image to be processed using the target detection model; Performing high-definition encoding on the region of interest in the image to be processed, and performing low-definition encoding on the remaining region of the image to be processed except the region of interest, to obtain a target image corresponding to the image to be processed; The target image is sent to a target device.

2. The method according to claim 1, wherein: The step of selecting a target detection model that matches the image to be processed from a plurality of candidate detection models includes: Obtain N first candidate models corresponding to N popular games one by one; wherein N ≥ 2 and is an integer; the first target model is obtained by training the initial detection model using a first training data set constructed based on the target popular game; the first target model is used to characterize each first candidate model in the N first candidate models; the target popular game is a popular game in the N popular games corresponding to the first target model; Obtain a second candidate model corresponding to the game collection; The target detection model is selected from the multiple candidate detection models including the N first candidate models and the second candidate model.

3. The method according to claim 2, wherein: The obtaining of the second candidate model corresponding to the game collection includes: The second candidate model is obtained by taking the N popular games as the game collection and training the initial detection model using a second training data set constructed based on the game collection.

4. The method according to claim 3, wherein: The selecting the target detection model from the plurality of candidate detection models including the N first candidate models and the second candidate model comprises: Determining the data source of the image to be processed; In a case where the data source is one of the N popular games, selecting a first candidate model corresponding to the data source from the N first candidate models as the object detection model; Alternatively, when the data source is not one of the N popular games, the second candidate model is used as the target detection model.

5. The method according to claim 2, wherein: The obtaining of the second candidate model corresponding to the game collection includes: Obtain M second candidate models corresponding one to one with M game types; wherein M ≥ 2 and is an integer; a second target model is obtained by training the initial detection model by taking a plurality of games of the same type under the target game type as the game collection and utilizing a second training data set constructed based on the game collection; the second target model is used to characterize each second candidate model in the M second candidate models; the target game type is a game type in the M game types corresponding to the second target model.

6. The method according to claim 5, wherein: The selecting the target detection model from the plurality of candidate detection models including the N first candidate models and the second candidate model comprises: Determining the data source of the image to be processed; In a case where the data source is one of the N popular games, selecting a first candidate model corresponding to the data source from the N first candidate models as the object detection model; Alternatively, when the data source is not one of the N popular games, a second candidate model corresponding to the data source is selected from the M second candidate models as the target detection model.

7. The method according to any one of claims 1 to 6, wherein: The step of obtaining the image to be processed comprises: Determine the image detection frequency based on the action interval duration and / or network delay; wherein the action interval duration is used to characterize the interval duration between the trigger time point of the most recent user action and the current time point; Based on the image detection frequency, a video image is extracted from a video data stream as the image to be processed.

8. The method according to claim 7, wherein: The determining of the image detection frequency based on the action interval duration and / or network delay includes: When the action interval duration is within a preset short duration interval, using the first detection frequency as the image detection frequency; Alternatively, when the action interval duration is not within the preset short duration interval, a second detection frequency is used as the image detection frequency; wherein the second detection frequency is lower than the first detection frequency.

9. The method according to claim 7, wherein: The determining of the image detection frequency based on the action interval duration and / or network delay includes: When the network delay is within a preset high delay interval, using the third detection frequency as the image detection frequency; Alternatively, when the network delay is not within the preset high delay interval, the fourth detection frequency is used as the image detection frequency; wherein the fourth detection frequency is lower than the third detection frequency.

10. A method for training a detection model, comprising: Based on different image data sources, construct multiple training data sets; wherein each of the multiple training data sets includes a sample image and identification information for characterizing a region of interest in the sample image; The initial detection model is trained using the multiple training data sets to obtain multiple candidate detection models that correspond one-to-one to the multiple training data sets.

11. The method according to claim 10, wherein: The method constructs multiple training data sets based on different image data sources, including: Based on N popular games, construct N first training data sets; where N ≥ 2 and is an integer; Based on the game collection, construct a second training data set; The N first training data sets and the second training data set are used as the multiple training data sets.

12. The method according to claim 11, wherein: The step of constructing a second training data set based on the game collection includes: Using the N popular games as the game collection, so as to construct the second training data set based on the game collection; Alternatively, a plurality of games of the same type under the target game type are taken as the game collection to construct the second training data set based on the game collection; wherein the target game type is used to characterize each game type in the M game types; M ≥ 2 and is an integer.

13. An image processing device, comprising: An image acquisition unit, used for acquiring an image to be processed; A model selection unit, used to select a target detection model that matches the image to be processed from multiple candidate detection models; A region determination unit, configured to determine a region of interest in the image to be processed using the target detection model; An image encoding unit, configured to perform high-definition encoding on a region of interest in the image to be processed, and to perform low-definition encoding on a remaining region of the image to be processed except the region of interest, so as to obtain a target image corresponding to the image to be processed; An image sending unit is used to send the target image to a target device.

14. A training device for a detection model, comprising: A data set construction unit, configured to construct a plurality of training data sets based on different image data sources; wherein each of the plurality of training data sets includes a sample image and identification information for characterizing a region of interest in the sample image; The model training unit is used to train the initial detection model using the multiple training data sets to obtain multiple candidate detection models that correspond one-to-one to the multiple training data sets.

15. An electronic device, comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 12.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 12.

17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.