Game props classification and neural network training method and device
By jointly training the game prop classification network and the environment classification network, and sharing the initial feature extraction network, the classification and training problems under the influence of complex game environments are solved, and more robust neural network training and reduced training costs are achieved.
Patent Information
- Application Number
- CN202180002742.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-27
- Filing Date
- 2021-09-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-09-28
AI Technical Summary
In game scenarios, the classification of game props and the training of neural networks are affected by the complex game environment, resulting in high complexity in sample images collection and labeling, which increases the training cost of neural networks.
By conducting joint training on the first initial classification network and the second initial classification network, sharing the same initial feature extraction network, the game environment category information extracted from the input image is used to assist in the training of the first initial classification network.
The robustness of the trained neural network is improved, allowing it to adapt to a variety of different gaming environments, reducing the complexity of sample images acquisition and labeling, and thus reducing the training cost of neural networks.
Smart Images

Figure CN113853243B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] The present disclosure claims priority to Singapore patent application number 10202110639U filed on September 26, 2021, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the field of computer vision technology, and in particular to a method and device for classifying game props and training a neural network. Background Art
[0004] In the game scene, game props need to be classified. The above process is generally implemented through neural networks. However, game props often have many categories, and the actual game environment, such as lighting, background, shadows, etc. in the game area, is often complex and changeable. In order to train a more robust neural network, it is necessary to consider the above two factors at the same time. In this way, the sample images used to train the neural network need to include multiple game props in different game environments and multiple different categories, and the acquisition and annotation of sample images are more complex. Summary of the invention
[0005] The present invention provides a method and device for classifying game props and training a neural network.
[0006] According to a first aspect of an embodiment of the present disclosure, a method for classifying game props is provided, the method comprising: inputting an image to be processed including a target game prop into a pre-trained first target classification network; obtaining the category of the target game prop output by the first target classification network; the first target classification network and the second target classification network are obtained by jointly training a first initial classification network and a second initial classification network, and the first initial classification network and the second initial classification network share the same initial feature extraction network; the first initial classification network is used to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network; the second initial classification network is used to classify the game environment in which the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network.
[0007] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; the features of the first sample image and the features of the second sample image are both extracted by the initial feature extraction network.
[0008] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the method further includes: based on the first sample image and the prop classification loss, performing a first training on the initial feature extraction network and the initial prop classification network to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; based on the first sample image, the second sample image and the environment classification loss, performing a second training on the intermediate feature extraction network and the initial environment classification network to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0009] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0010] In some embodiments, the second training of the intermediate feature extraction network and the initial environment classification network is performed based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, including: fixing the intermediate feature extraction network, and training the initial environment classification network based on the first loss, the first sample image and the second sample image to obtain the target environment classification network; fixing the target environment classification network, and training the intermediate feature extraction network based on the second loss, the first sample image and the second sample image to obtain the target feature extraction network.
[0011] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0012] In some embodiments, the method also includes: acquiring an image set, the image set including a first image subset and a second image subset, the images in the first image subset including a first label, the images in the second image subset not including the first label, the first label being used to characterize the category of the game props; marking the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, the second label being used to characterize the category of the game environment in which the game props are located; determining the images in the first image subset as the first sample images, and determining the images in the second image subset as the second sample images.
[0013] According to a second aspect of an embodiment of the present disclosure, a neural network training method is provided, which is used to train a first initial classification network to obtain a first target classification network, wherein the first target classification network is used to classify game props; the method comprises: obtaining a first sample image and a second sample image; jointly training the first initial classification network and the second initial classification network based on the first sample image and the second sample image to obtain the first target classification network, wherein the first initial classification network and the second initial classification network share the same initial feature extraction network; the first initial classification network is used to classify game props in the first sample image based on features extracted from the first sample image by the initial feature extraction network; the second initial classification network is used to classify the game environment in which the game props in the first sample image and the second sample image are located based on features extracted from the first sample image and the second sample image by the initial feature extraction network.
[0014] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image.
[0015] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the joint training of the first initial classification network and the second initial classification network includes: based on the first sample image and the prop classification loss, first training the initial feature extraction network and the initial prop classification network to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; based on the first sample image, the second sample image and the environment classification loss, second training the intermediate feature extraction network and the initial environment classification network to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0016] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0017] In some embodiments, the second training of the intermediate feature extraction network and the initial environment classification network is performed based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, including: fixing the intermediate feature extraction network, and training the initial environment classification network based on the first loss, the first sample image and the second sample image to obtain the target environment classification network; fixing the target environment classification network, and training the intermediate feature extraction network based on the second loss, the first sample image and the second sample image to obtain the target feature extraction network.
[0018] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0019] In some embodiments, the method also includes: acquiring an image set, the image set including a first image subset and a second image subset, the images in the first image subset including a first label, the images in the second image subset not including the first label, the first label being used to characterize the category of the game props; marking the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, the second label being used to characterize the category of the game environment in which the game props are located; determining the images in the first image subset as the first sample images, and determining the images in the second image subset as the second sample images.
[0020] According to a third aspect of an embodiment of the present disclosure, a game prop classification device is provided, the device comprising: an input module, used to input a to-be-processed image including a target game prop into a pre-trained first target classification network; a classification module, used to obtain the category of the target game prop output by the first target classification network; the first target classification network and the second target classification network are obtained by jointly training a first initial classification network and a second initial classification network, and the first initial classification network and the second initial classification network share the same initial feature extraction network; the first initial classification network is used to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network; the second initial classification network is used to classify the game environment in which the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network.
[0021] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; the features of the first sample image and the features of the second sample image are both extracted by the initial feature extraction network.
[0022] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the device also includes: a first training module for performing a first training on the initial feature extraction network and the initial prop classification network based on the first sample image and the prop classification loss to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; a second training module for performing a second training on the intermediate feature extraction network and the initial environment classification network based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0023] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0024] In some embodiments, the second training module is used to: fix the intermediate feature extraction network, and train the initial environment classification network based on the first loss, the first sample image, and the second sample image to obtain the target environment classification network; fix the target environment classification network, and train the intermediate feature extraction network based on the second loss, the first sample image, and the second sample image to obtain the target feature extraction network.
[0025] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0026] In some embodiments, the device also includes: an image set acquisition module, used to acquire an image set, the image set includes a first image subset and a second image subset, the images in the first image subset include a first label, and the images in the second image subset do not include the first label, and the first label is used to characterize the category of the game props; a labeling module, used to label the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, and the second label is used to characterize the category of the game environment in which the game props are located; a sample image determination module, used to determine the image in the first image subset as the first sample image, and determine the image in the second image subset as the second sample image.
[0027] According to a fourth aspect of an embodiment of the present disclosure, a neural network training device is provided, which is used to train a first initial classification network to obtain a first target classification network, wherein the first target classification network is used to classify game props; the device includes: a sample image acquisition module, which is used to acquire a first sample image and a second sample image; a training module, which is used to jointly train the first initial classification network and the second initial classification network based on the first sample image and the second sample image to obtain the first target classification network, wherein the first initial classification network and the second initial classification network share the same initial feature extraction network; the first initial classification network is used to classify game props in the first sample image based on features extracted from the first sample image by the initial feature extraction network; the second initial classification network is used to classify the game environment in which the game props in the first sample image and the second sample image are located based on features extracted from the first sample image and the second sample image by the initial feature extraction network.
[0028] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image.
[0029] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the training module includes: a first training unit for performing a first training on the initial feature extraction network and the initial prop classification network based on the first sample image and the prop classification loss to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; a second training unit for performing a second training on the intermediate feature extraction network and the initial environment classification network based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0030] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0031] In some embodiments, the second training unit is used to: fix the intermediate feature extraction network, and train the initial environment classification network based on the first loss, the first sample image, and the second sample image to obtain the target environment classification network; fix the target environment classification network, and train the intermediate feature extraction network based on the second loss, the first sample image, and the second sample image to obtain the target feature extraction network.
[0032] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0033] In some embodiments, the device also includes: an image set acquisition module, used to acquire an image set, the image set includes a first image subset and a second image subset, the images in the first image subset include a first label, and the images in the second image subset do not include the first label, and the first label is used to characterize the category of the game props; a labeling module, used to label the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, and the second label is used to characterize the category of the game environment in which the game props are located; a sample image determination module, used to determine the image in the first image subset as the first sample image, and determine the image in the second image subset as the second sample image.
[0034] According to a fifth aspect of an embodiment of the present disclosure, there is provided a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in any one of the embodiments is implemented.
[0035] According to a sixth aspect of an embodiment of the present disclosure, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of the embodiments is implemented.
[0036] According to a seventh aspect of an embodiment of the present disclosure, a computer program is provided, comprising a computer-readable code, wherein when the computer-readable code is executed in an electronic device, the processor in the electronic device executes the method described in any embodiment of the present disclosure.
[0037] The disclosed embodiment obtains the first target classification network and the second target classification network by jointly training the first initial classification network and the second initial classification network. Since the first initial classification network and the second initial classification network share the same feature extraction network, and the second initial classification network can classify the game environment in which the game props in the input image are located based on the features extracted from the input image by the feature extraction network, the category information of the game environment determined from the input image by the second initial classification network can be used to assist the training of the first initial classification network. In this way, the robustness of the trained first target classification network can be improved, so that the first target classification network can adapt to a variety of different game environments, thereby eliminating the need to collect a large number of sample images of game props of different categories in different game environments, reducing the high complexity of sample image collection and annotation, and thus reducing the training cost of the neural network.
[0038] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used to illustrate the technical solutions of the present disclosure together with the specification.
[0040] Figure 1A , Figure 1B , Figure 1C and Figure 1D These are schematic diagrams of different game props.
[0041] Figure 2A and Figure 2B They are schematic diagrams of game props in different game environments.
[0042] Figure 3 It is a flowchart of the method for classifying game props according to an embodiment of the present disclosure.
[0043] Figure 4A It is a general schematic diagram of the training process of the neural network of the embodiment of the present disclosure.
[0044] Figure 4B It is a schematic diagram of the network structure during the training process of the neural network in the embodiment of the present disclosure.
[0045] Figure 5 It is a flowchart of a neural network training method according to an embodiment of the present disclosure.
[0046] Figure 6 It is a block diagram of a game prop classification device according to an embodiment of the present disclosure.
[0047] Figure 7It is a block diagram of a neural network training device according to an embodiment of the present disclosure.
[0048] Figure 8 It is a schematic diagram of the structure of a computer device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0049] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0050] The terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. The singular forms of "a", "said" and "the" used in this disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items. In addition, the term "at least one" herein means any combination of at least two of any one or more of a plurality of.
[0051] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same category from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0052] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and to make the above-mentioned purposes, features and advantages of the embodiments of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure are further described in detail below in conjunction with the accompanying drawings.
[0053] Game scenes often include multiple categories of game props, such as game coins, cards, game markers, and dice. Different game props may further include multiple subcategories. For example, cards include multiple cards with different points or different patterns, and game coins include multiple game coins with different denominations. Figure 1A and Figure 1B As shown in the figure, they are schematic diagrams of two different types of cards; Figure 1C and Figure 1DAs shown in the figure, they are schematic diagrams of two different types of game coins. It can be seen that different types of cards use different patterns, and different types of game coins have different colors and patterns. Of course, in actual applications, in addition to color and pattern, other attributes (for example, size, material, etc.) of different types of game props may also be different.
[0054] In addition, different areas of the game scene, or the game scene at different time periods may have different categories of game environments. Factors affecting the category of the game environment may include, but are not limited to, the category of the game area, the color of the light, and the shadow area. Different game environment categories correspond to one or more different factors. For example, game environment A and game environment B have the same game area category, the same light color, and different shadow areas, while game environment B and game environment C have different game area categories, the same light color, and different shadow areas.
[0055] In order to train a neural network with relatively robust performance to identify game props in game scenes, it is usually necessary to include as many sample images as possible of game props of different categories in different game environments. For example, if the total number of game prop categories is M and the total number of game environment categories is N, then the number of sample images required is M×N, so that the neural network can learn the characteristics of different game props in different game environments. Figure 2A and Figure 2B The following is a schematic diagram of game props in different game environments. The figure uses different shadow categories to illustrate the game environment. It can be seen that Figure 2A In the game environment shown, the light is almost directly above the game coin, so the shadow area of the game coin is small, and the shadow of game coin A blocks less of game coin B. Figure 2B In the game environment shown, the light is on the upper left of the game coin, so the shadow area of the game coin is larger, and the shadow of game coin A covers more of game coin B. It can be seen that the game environment may have a certain impact on the recognition of the neural network. Therefore, in order to improve the robustness of the neural network, the sample images used to train the neural network need to include the game coin in the upper left corner. Figure 2A Sample images of the game environment shown, and game coins in Figure 2B Sample image of the game environment shown.
[0056] However, the above method requires collecting sample images in a variety of different game scenarios, and the collected sample images all need to be labeled, which leads to high complexity in collecting and labeling sample images, and thus high training costs for the neural network.
[0057] Based on this, the present disclosure provides a method for classifying game props. Figure 3 As shown, the method includes:
[0058] Step 301: inputting the image to be processed including the target game prop into a pre-trained first target classification network;
[0059] Step 302: Obtain the category of the target game prop output by the first target classification network;
[0060] The first target classification network and the second target classification network are obtained by jointly training the first initial classification network and the second initial classification network, and the first initial classification network and the second initial classification network share the same initial feature extraction network;
[0061] The first initial classification network is used to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network;
[0062] The second initial classification network is used to classify the game environment in which the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network.
[0063] In step 301, the target game props may be game props in the game area, such as game coins, cards, etc. The game area may be imaged to obtain an image to be processed including the game props. In some embodiments, an image acquisition device may be arranged around the game area to collect the video of the game process in real time during the game, and the video frame including the target game props in the collected video may be input into the first target classification network as the image to be processed. Alternatively, all video frames in the collected video may be input into the first target classification network, and the image to be processed including the target game props may be screened out by the first target classification network, and then subsequent processing may be performed.
[0064] In step 302, the first target classification network can output the category of the target game props, for example, whether the target game props are game coins or cards. The subcategories of the target game props can also be output. For example, when the target game props are game coins, each subcategory of game coins corresponds to a denomination, and the first target classification network can output the denomination category of the game coins; for another example, when the target game props are cards, each subcategory of cards corresponds to a point and a suit, and the first target classification network can output the point and suit of the card. Further, the first target classification network can also output the number of target game props, for example, a pile of game coins stacked in the vertical direction, the first target classification network can output how many pieces of game coins are stacked in this pile of game coins.
[0065] In this embodiment, the first target classification network can be obtained by multi-task joint training. That is, the first target classification network and the second target classification network are obtained by jointly training the first initial classification network and the second initial classification network. The first target classification network is a neural network obtained by jointly training the first initial classification network, and the second target classification network is a neural network obtained by jointly training the second initial classification network.
[0066] The first initial classification network is used to perform the prop classification task, that is, to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network. The second initial classification network is used to perform the environment classification task, that is, to classify the game environment in which the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network. The input image generally refers to an image input to the feature extraction network. In different training stages, the input image may refer to the first sample image or the second sample image.
[0067] Through multi-task joint training, the category information of the game environment determined by the second initial classification network from the input image can be used to assist the training of the first initial classification network. In this way, the robustness of the trained first target classification network can be improved, so that the first target classification network can adapt to a variety of different game environments, thereby eliminating the need to collect a large number of sample images of game props of different categories in different game environments, reducing the complexity of sample image collection and annotation, and thus reducing the training cost of the neural network.
[0068] In some embodiments, the first initial classification network is trained based on the prop classification loss, and the prop classification loss is the classification loss of the first initial classification network for classifying the game props in the first sample image based on the features of the first sample image. The second initial classification network is trained based on the environment classification loss, and the environment classification loss is the classification loss of the second initial classification network for classifying the game environment in which the game props in the first sample image and the second sample image are located based on the features of the first sample image and the second sample image. Wherein, the features of the first sample image and the features of the second sample image are both extracted by the initial feature extraction network.
[0069] By using the item classification loss to train the first initial classification network, the trained first target classification network can learn the features used to distinguish different categories of game items, so that the first target classification network can obtain sufficient classification accuracy. By using the environment classification loss to train the second initial classification network, the classification results of the second initial classification network can be used during the training process to constrain the common part (i.e., the feature extraction network) between the second initial classification network and the first initial classification network, thereby reducing the impact of different game environments on the first target classification network, thereby improving the robustness of the first target classification network in different game environments.
[0070] The first sample image may carry a first label and a second label, wherein the first label is used to characterize the category of the game props in the first sample image. The second sample image may only carry the second label, wherein the second label is used to characterize the category of the game environment in which the game props in the second sample image are located. In some embodiments, some images may be collected, and these images may be marked with the first label as the first sample image, and other images may be collected, and these images may be marked with the second label as the second sample image. That is, the first sample image and the second sample image are different images.
[0071] In some other embodiments, an image set may also be obtained, the image set including a first image subset and a second image subset, the images in the first image subset may include a first label and a second label, and the images in the second image subset do not include the first label. The second labels of the images in the first image subset and the second labels of the images in the second image subset may be marked as different second labels, and the images in the first image subset may be determined as the first sample images, and the images in the second image subset may be determined as the second sample images.
[0072] For example, the images in the first image subset may be pre-collected and annotated with a first label, wherein the images in the first image subset may include images collected under multiple different game environments. The second labels of all images in the first image subset may be annotated with the same label, for example, all are annotated with "1". The images in the second image subset may also include images collected under multiple different game environments. The images in the second image subset may not include the first label, and the second labels of all images in the second image subset may be annotated with the same label, and the second labels of the images in the second image subset are different from the second labels of the images in the first image subset, for example, the second labels of the images in the second image subset are all annotated with "0". In this way, only two different second labels need to be directly specified, and there is no need to determine the category of the specific game environment corresponding to each image in the image set separately, and only the first labels of the images in the first image subset need to be annotated, and the images in the second image subset do not need to be annotated with the first labels, thereby reducing the annotation complexity of the sample images.
[0073] The images in the first image subset and the images in the second image subset may be video frames reflowed from videos captured in real time on-site in the game area, or may be images collected by simulating a real game scene and collectively capturing the images in the simulated game scene.
[0074] The disclosed embodiment does not need to collect and annotate sample images corresponding to the combinations of various game props and various game environments. It only needs to collect and annotate some first sample images with a first label, and specify the second labels of the first sample images annotated with the first label and the second sample images not annotated with the first label as different labels. In the related art, assuming that the total number of categories of game props is M, the total number of categories of game environments is N, and assuming that the number of each sample image is 1, the number of sample images required to be collected and annotated is M×N. In the present embodiment, only M first sample images carrying the first label and the second label and n (1<n≤N) second sample images that do not carry the first label and only carry the second label can be collected and annotated, which effectively reduces the number of sample images required, thereby reducing the complexity of collecting and annotating sample images.
[0075] In order to obtain sample images that are more valuable to the training process and to improve the classification accuracy of the first target classification network, the first sample image and the second sample image can be screened out from the game image based on preset conditions. The preset conditions include: detecting a preset event related to the game props in the game area from the game image. The preset event can be an event in which an error occurs in the operation of the game props. In an application scenario, the detection and recognition results of the game image obtained by imaging the game area can be output to the business layer for business processing, so that the business layer determines whether an error occurs in the operation of the game props. For example, the placement position of a specific game prop can be identified from a game image, and the placement position is sent to the business layer. The business layer can determine whether the placement position of the game prop is within the preset placement area. If it is not within the placement area, the business layer reports an error. For another example, the placement order of specific game props can be identified from multiple game images, and the placement order is sent to the business layer. The business layer can determine whether the placement order of the game props meets the preset order. If it does not meet the preset order, the business layer reports an error. Part or all of the game images in which errors are reported at the business layer may be annotated, thereby obtaining a first sample image and / or a second sample image.
[0076] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an image input to the first initial classification network; and an initial prop classification network for classifying game props in the image input to the first initial classification network based on the features extracted by the initial feature extraction network. The second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props are located in the image input to the second initial classification network based on the features extracted by the initial feature extraction network. Optionally, the initial feature extraction network is a convolutional neural network, and the network structure of the initial feature extraction network can be a resnet network body. The initial prop classification network includes a fully connected layer and a softmax layer for outputting the category of a single game prop, or the initial prop classification network includes a fully connected layer and a CTC (Connectionist Temporal Classification) network for identifying game props such as stacked game coins and outputting sequence recognition results.
[0077] Reference below Figure 4A and Figure 4B The training process of the first target classification network is explained.
[0078] First, based on the first sample image and the prop classification loss, the initial feature extraction network and the initial prop classification network are first trained to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network. The prop classification loss can be determined based on the difference between the classification result of the target prop classification network and the true category of the game props in the first sample image. In some embodiments, the prop classification loss is a cross entropy loss.
[0079] Then, based on the first sample image, the second sample image and the environment classification loss, the intermediate feature extraction network and the initial environment classification network are subjected to a second training to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network. In the second training process, adversarial training is performed between the initial environment classification network and the intermediate feature extraction network, so that the trained target environment classification network is difficult to distinguish the environment category corresponding to the features extracted by the target feature extraction network, that is, the features extracted by the target feature extraction network have the same distribution characteristics in different game environments. In this way, the influence of different categories of game environments on the classification results of game props can be reduced, thereby improving the robustness of the first target classification network.
[0080] In some embodiments, the second classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props in the first sample image and the second sample image are located, and the second loss is used to characterize the ability of the initial environment classification network to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong. By adopting the first loss, a target environment classification network with good classification performance can be trained. By adopting the second loss, a target feature extraction network with good feature extraction performance can be trained. By making the second labels of the first sample image and the second sample image different, even if the two involve the same game environment, the features extracted by the target feature extraction network obtained by training the intermediate feature extraction network pre-trained only based on the first sample image based on the first sample image and the second sample image can confuse the classification performance of the target environment classification network trained based on the first sample image and the second sample image, so that the target environment classification network cannot distinguish the environment category corresponding to the features extracted by the target feature extraction network.
[0081] The second training includes two stages. In the first stage, the intermediate feature extraction network is fixed, and the initial environment classification network is trained based on the first loss, the first sample image, and the second sample image, so as to train a target environment classification network with better classification performance. In the second stage, the target environment classification network is fixed, and the intermediate feature extraction network is trained based on the second loss, the first sample image, and the second sample image, so as to train a target feature extraction network with better feature extraction performance. Through the two-stage training method, the convergence speed in the training process can be improved, and the trained first target classification network can be made more stable.
[0082] The training process of the entire first target classification network is to iterate each training step alternately until convergence (the target environment classification network cannot distinguish the environment category, and the target prop classification network plays a role). The specific process is as follows:
[0083] A) Using the first sample image and the prop classification loss (loss3) to train the initial prop classification network, update the network parameters of the initial feature extraction network and the initial prop classification network, and obtain an intermediate neural network including an intermediate feature extraction network and a target prop classification network. The initial environment classification network is not trained in this step.
[0084] B) Fix the intermediate feature extraction network, use the first sample image, the second sample image and the first loss (loss1) to train the initial environment classification network, and obtain the intermediate feature extraction network and the target environment classification network. The target prop classification network is not trained in this step. The first loss can be a cross entropy loss. When the environment category is 2, the first loss is recorded as:
[0085]
[0086] Among them, p(i) is the true environment category probability vector, [1,0] means that the second label is 1, [0,1] means that the second label is 0, and q(i) is the category of the game environment predicted by the initial environment classification network.
[0087] C) Fix the target environment classification network, use the first sample image, the second sample image and the second loss (loss2) to train the intermediate feature extraction network to obtain the target feature extraction network. The target prop classification network is not trained in this step. The purpose of this step is to make the target feature extraction network unable to distinguish the category of the game environment, that is, to optimize the target p(i) to be uniformly distributed [0.5, 0.5], and the second loss is recorded as:
[0088]
[0089] Through the above training process, a first target classification network including a target feature extraction network and a target prop classification network, and a second target classification network including a target feature extraction network and a target environment classification network are obtained. In the inference stage, only the first target classification network can be used to classify the game props in the processed image without using the second target classification network.
[0090] Furthermore, the performance of the first target classification network can be tested by collecting and annotating test images from real game scenes. If the classification accuracy of the first target classification network is higher than a preset accuracy threshold, the first target classification network is determined as the neural network ultimately used to classify game props; otherwise, the first target classification network is retrained.
[0091] In practical applications, different game scenes often have different game environments. When the first target classification network in one game scene is used in another game scene, it is often necessary to re-collect and annotate sample images to train the first target classification network. The method of the disclosed embodiment can make the first target classification network adapt to various different game scenes. The disclosed embodiment can achieve rapid generalization performance improvement for the first target classification network that has achieved high accuracy in a limited scene but is prone to errors in a new test environment, adapt to a variety of new and different test environments, and improve the robustness of the first target classification network.
[0092] like Figure 5 As shown, the embodiment of the present disclosure also provides a neural network training method for training a first initial classification network to obtain a first target classification network, wherein the first target classification network is used to classify game props; the method comprises:
[0093] Step 501: Acquire a first sample image and a second sample image;
[0094] Step 502: jointly training the first initial classification network and the second initial classification network based on the first sample image and the second sample image to obtain the first target classification network, wherein the first initial classification network and the second initial classification network share the same initial feature extraction network;
[0095] The first initial classification network is used to classify the game props in the first sample image based on the features extracted from the first sample image by the initial feature extraction network;
[0096] The second initial classification network is used to classify the game environment in which the game props in the first sample image and the second sample image are located based on the features extracted from the first sample image and the second sample image by the initial feature extraction network.
[0097] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image.
[0098] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the joint training of the first initial classification network and the second initial classification network includes: based on the first sample image and the prop classification loss, first training the initial feature extraction network and the initial prop classification network to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; based on the first sample image, the second sample image and the environment classification loss, second training the intermediate feature extraction network and the initial environment classification network to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0099] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0100] In some embodiments, the second training of the intermediate feature extraction network and the initial environment classification network is performed based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, including: fixing the intermediate feature extraction network, and training the initial environment classification network based on the first loss, the first sample image and the second sample image to obtain the target environment classification network; fixing the target environment classification network, and training the intermediate feature extraction network based on the second loss, the first sample image and the second sample image to obtain the target feature extraction network.
[0101] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0102] In some embodiments, the method also includes: acquiring an image set, the image set including a first image subset and a second image subset, the images in the first image subset including a first label, the images in the second image subset not including the first label, the first label being used to characterize the category of the game props; marking the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, the second label being used to characterize the category of the game environment in which the game props are located; determining the images in the first image subset as the first sample images, and determining the images in the second image subset as the second sample images.
[0103] The details of the training method of the above neural network are detailed in the embodiment of the aforementioned game prop classification method, which will not be repeated here.
[0104] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0105] like Figure 6 As shown, the embodiment of the present disclosure also provides a game prop classification device, the device comprising:
[0106] An input module 601 is used to input the to-be-processed image including the target game prop into a pre-trained first target classification network;
[0107] A classification module 602, used to obtain the category of the target game prop output by the first target classification network;
[0108] The first target classification network and the second target classification network are obtained by jointly training the first initial classification network and the second initial classification network, and the first initial classification network and the second initial classification network share the same initial feature extraction network;
[0109] The first initial classification network is used to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network;
[0110] The second initial classification network is used to classify the game environment in which the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network.
[0111] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; the features of the first sample image and the features of the second sample image are both extracted by the initial feature extraction network.
[0112] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the device also includes: a first training module for performing a first training on the initial feature extraction network and the initial prop classification network based on the first sample image and the prop classification loss to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; a second training module for performing a second training on the intermediate feature extraction network and the initial environment classification network based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0113] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0114] In some embodiments, the second training module is used to: fix the intermediate feature extraction network, and train the initial environment classification network based on the first loss, the first sample image, and the second sample image to obtain the target environment classification network; fix the target environment classification network, and train the intermediate feature extraction network based on the second loss, the first sample image, and the second sample image to obtain the target feature extraction network.
[0115] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0116] In some embodiments, the device also includes: an image set acquisition module, used to acquire an image set, the image set includes a first image subset and a second image subset, the images in the first image subset include a first label, and the images in the second image subset do not include the first label, and the first label is used to characterize the category of the game props; a labeling module, used to label the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, and the second label is used to characterize the category of the game environment in which the game props are located; a sample image determination module, used to determine the image in the first image subset as the first sample image, and determine the image in the second image subset as the second sample image.
[0117] like Figure 7 As shown, the embodiment of the present disclosure also provides a neural network training device, which is used to train a first initial classification network to obtain a first target classification network, and the first target classification network is used to classify game props; the device includes:
[0118] The sample image acquisition module 701 is used to acquire a first sample image and a second sample image;
[0119] A training module 702 is used to jointly train the first initial classification network and the second initial classification network based on the first sample image and the second sample image to obtain the first target classification network, wherein the first initial classification network and the second initial classification network share the same initial feature extraction network;
[0120] The first initial classification network is used to classify the game props in the first sample image based on the features extracted from the first sample image by the initial feature extraction network; the second initial classification network is used to classify the game environment in which the game props in the first sample image and the second sample image are located based on the features extracted from the first sample image and the second sample image by the initial feature extraction network.
[0121] In some embodiments, the first initial classification network is trained based on a prop classification loss, where the prop classification loss is the classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; the second initial classification network is trained based on an environment classification loss, where the environment classification loss is the classification loss of the second initial classification network for classifying the game environments in which game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image.
[0122] In some embodiments, the first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on the features extracted by the initial feature extraction network; the second initial classification network includes the initial feature extraction network and an initial environment classification network for classifying the game environment in which the game props in the input image are located based on the features extracted by the initial feature extraction network; the training module includes: a first training unit for performing a first training on the initial feature extraction network and the initial prop classification network based on the first sample image and the prop classification loss to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; a second training unit for performing a second training on the intermediate feature extraction network and the initial environment classification network based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
[0123] In some embodiments, the environment classification loss includes a first loss and a second loss, the first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the actual category of the game environment where the game props are located in the first sample image and the second sample image, and the second loss is used to characterize the initial environment classification network's ability to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
[0124] In some embodiments, the second training unit is used to: fix the intermediate feature extraction network, and train the initial environment classification network based on the first loss, the first sample image, and the second sample image to obtain the target environment classification network; fix the target environment classification network, and train the intermediate feature extraction network based on the second loss, the first sample image, and the second sample image to obtain the target feature extraction network.
[0125] In some embodiments, the first sample image and the second sample image are screened from a game image based on preset conditions, and the game image is obtained by imaging a game area. The preset conditions include: detecting from the game image that a preset event related to game props occurs in the game area.
[0126] In some embodiments, the device also includes: an image set acquisition module, used to acquire an image set, the image set includes a first image subset and a second image subset, the images in the first image subset include a first label, and the images in the second image subset do not include the first label, and the first label is used to characterize the category of the game props; a labeling module, used to label the second label of the images in the first image subset and the second label of the images in the second image subset as different second labels, and the second label is used to characterize the category of the game environment in which the game props are located; a sample image determination module, used to determine the image in the first image subset as the first sample image, and determine the image in the second image subset as the second sample image.
[0127] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0128] An embodiment of the present specification also provides a computer device, which at least includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any of the above embodiments is implemented.
[0129] Figure 8A more specific hardware structure diagram of a computing device provided in an embodiment of this specification is shown, and the device may include: a processor 801, a memory 802, an input / output interface 803, a communication interface 804, and a bus 805. The processor 801, the memory 802, the input / output interface 803, and the communication interface 804 are connected to each other in communication within the device through the bus 805.
[0130] The processor 801 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., for executing relevant programs to implement the technical solutions provided in the embodiments of this specification. The processor 801 can also include a graphics card, which can be a Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0131] The memory 802 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 802 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 802 and are called and executed by the processor 801.
[0132] The input / output interface 803 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0133] The communication interface 804 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0134] The bus 805 includes a path to transmit information between various components of the device (eg, the processor 801 , the memory 802 , the input / output interface 803 , and the communication interface 804 ).
[0135] It should be noted that, although the above device only shows the processor 801, the memory 802, the input / output interface 803, the communication interface 804 and the bus 805, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0136] The present disclosure also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in any of the above embodiments is implemented.
[0137] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0138] The present disclosure also provides a computer program, including a computer-readable code. When the computer-readable code is executed in an electronic device, a processor in the electronic device can execute the method described in any of the above embodiments.
[0139] It can be known from the above description of the implementation mode that the technicians in this field can clearly understand that the embodiments of this specification can be implemented by means of software plus the necessary general hardware platform. Based on such an understanding, the technical solution of the embodiments of this specification can be essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.
[0140] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver, a game console, a tablet computer, a wearable device or a combination of any of these devices.
[0141] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely schematic, wherein the modules described as separate components may or may not be physically separated, and the functions of each module can be implemented in the same one or more software and / or hardware when implementing the embodiment scheme of this specification. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment scheme. A person of ordinary skill in the art can understand and implement it without paying creative labor.
[0142] The above is only a specific implementation of the embodiments of this specification. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiments of this specification. These improvements and modifications should also be regarded as the protection scope of the embodiments of this specification.
Claims
1. A method for classifying game props, the method include: Inputting the image to be processed including the target game prop into a pre-trained first target classification network; Obtaining the category of the target game prop output by the first target classification network; The first target classification network and the second target classification network are obtained by jointly training the first initial classification network and the second initial classification network, and the first initial classification network and the second initial classification network share the same initial feature extraction network; The first initial classification network is used to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network; The second initial classification network is used to classify the game environment where the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network; in, The first initial classification network is trained based on a prop classification loss, where the prop classification loss is a classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; The second initial classification network is trained based on an environment classification loss, where the environment classification loss is a classification loss for the second initial classification network to classify the game environments in which the game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; The features of the first sample image and the features of the second sample image are both extracted by the initial feature extraction network; The method further comprises: Acquire an image set, the image set including a first image subset and a second image subset, the images in the first image subset all include a first label, the images in the second image subset do not include the first label, and the first label is used to represent a category of a game prop; marking the second labels of the images in the first image subset and the second labels of the images in the second image subset as different second labels, wherein the second labels are used to characterize the category of the game environment where the game props are located; determining an image in the first image subset as the first sample image, An image in the second image subset is determined as the second sample image.
2. The method according to claim 1, The first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on features extracted by the initial feature extraction network; The second initial classification network includes the initial feature extraction network and an initial environment classification network, and is used to classify the game environment where the game props in the input image are located based on the features extracted by the initial feature extraction network; The method also include: Based on the first sample image and the prop classification loss, performing a first training on the initial feature extraction network and the initial prop classification network to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; Based on the first sample image, the second sample image and the environment classification loss, the intermediate feature extraction network and the initial environment classification network are subjected to a second training to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
3. The method according to claim 2, wherein the environmental classification loss include: A first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the real category of the game environment where the game props are located in the first sample image and the second sample image, The second loss is used to characterize the ability of the initial environment classification network to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
4. The method according to claim 3, wherein the intermediate feature extraction network and the initial environment classification network are second trained based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, include: Fixing the intermediate feature extraction network, and training the initial environment classification network based on the first loss, the first sample image, and the second sample image to obtain the target environment classification network; The target environment classification network is fixed, and the intermediate feature extraction network is trained based on the second loss, the first sample image, and the second sample image to obtain the target feature extraction network.
5. The method according to any one of claims 1 to 4, wherein the first sample image and the second sample image are obtained by screening a game image based on a preset condition, the game image is obtained by imaging a game area, and the preset condition include: A preset event related to game props occurring in the game area is detected from the game image.
6. A neural network training method, used to train a first initial classification network to obtain a first target classification network, wherein the first target classification network is used to classify game props; the method include: Acquire a first sample image and a second sample image; Jointly training the first initial classification network and the second initial classification network based on the first sample image and the second sample image to obtain the first target classification network, wherein the first initial classification network and the second initial classification network share the same initial feature extraction network; The first initial classification network is used to classify the game props in the first sample image based on the features extracted from the first sample image by the initial feature extraction network; The second initial classification network is used to classify the game environments where the game props in the first sample image and the second sample image are located based on the features extracted from the first sample image and the second sample image by the initial feature extraction network; The first initial classification network is trained based on a prop classification loss, where the prop classification loss is a classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; The second initial classification network is trained based on an environment classification loss, where the environment classification loss is a classification loss for the second initial classification network to classify the game environment in which the game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; Acquire an image set, the image set including a first image subset and a second image subset, the images in the first image subset all include a first label, the images in the second image subset do not include the first label, and the first label is used to represent a category of a game prop; marking the second labels of the images in the first image subset and the second labels of the images in the second image subset as different second labels, wherein the second labels are used to characterize the category of the game environment where the game props are located; determining an image in the first image subset as the first sample image, An image in the second image subset is determined as the second sample image.
7. The method according to claim 6, The first initial classification network includes an initial feature extraction network for extracting features from an input image; and an initial prop classification network for classifying game props in the input image based on features extracted by the initial feature extraction network; The second initial classification network includes the initial feature extraction network and an initial environment classification network, and is used to classify the game environment where the game props in the input image are located based on the features extracted by the initial feature extraction network; The jointly training the first initial classification network and the second initial classification network comprises: Based on the first sample image and the prop classification loss, performing a first training on the initial feature extraction network and the initial prop classification network to obtain an intermediate classification network including an intermediate feature extraction network and a target prop classification network; Based on the first sample image, the second sample image and the environment classification loss, the intermediate feature extraction network and the initial environment classification network are subjected to a second training to obtain a target feature extraction network and a target environment classification network, wherein the first target classification network includes the target feature extraction network and the target prop classification network.
8. The method according to claim 7, wherein the environmental classification loss include: A first loss is used to characterize the difference between the category of the game environment predicted by the initial environment classification network and the real category of the game environment where the game props are located in the first sample image and the second sample image, The second loss is used to characterize the ability of the initial environment classification network to distinguish the category of the game environment to which the features extracted by the intermediate feature extraction network belong.
9. The method according to claim 8, wherein the intermediate feature extraction network and the initial environment classification network are subjected to a second training based on the first sample image, the second sample image and the environment classification loss to obtain a target feature extraction network and a target environment classification network, include: Fixing the intermediate feature extraction network, and training the initial environment classification network based on the first loss, the first sample image, and the second sample image to obtain the target environment classification network; The target environment classification network is fixed, and the intermediate feature extraction network is trained based on the second loss, the first sample image, and the second sample image to obtain the target feature extraction network.
10. The method according to any one of claims 6 to 9, wherein the first sample image and the second sample image are obtained by screening a game image based on a preset condition, the game image is obtained by imaging a game area, and the preset condition include: A preset event related to game props occurring in the game area is detected from the game image.
11. A device for classifying game props, the device include: An input module, used for inputting the to-be-processed image including the target game prop into a pre-trained first target classification network; A classification module, used to obtain the category of the target game prop output by the first target classification network; The first target classification network and the second target classification network are obtained by jointly training the first initial classification network and the second initial classification network, and the first initial classification network and the second initial classification network share the same initial feature extraction network; The first initial classification network is used to classify the game props in the input image based on the features extracted from the input image by the initial feature extraction network; The second initial classification network is used to classify the game environment where the game props in the input image are located based on the features extracted from the input image by the initial feature extraction network; The first initial classification network is trained based on a prop classification loss, where the prop classification loss is a classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; The second initial classification network is trained based on an environment classification loss, where the environment classification loss is a classification loss for the second initial classification network to classify the game environment in which the game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; The features of the first sample image and the features of the second sample image are both extracted by the initial feature extraction network; The device further includes: an image set acquisition module, configured to acquire an image set, wherein the image set includes a first image subset and a second image subset, wherein images in the first image subset all include a first label, and images in the second image subset do not include the first label, and the first label is used to represent a category of a game prop; a labeling module, used to label the second labels of the images in the first image subset and the second labels of the images in the second image subset as different second labels, wherein the second labels are used to characterize the category of the game environment where the game props are located; The sample image determination module is used to determine an image in the first image subset as the first sample image, and determine an image in the second image subset as the second sample image.
12. A neural network training device, used to train a first initial classification network to obtain a first target classification network, wherein the first target classification network is used to classify game props; the device include: A sample image acquisition module, used to acquire a first sample image and a second sample image; A training module, configured to jointly train the first initial classification network and the second initial classification network based on the first sample image and the second sample image to obtain the first target classification network, wherein the first initial classification network and the second initial classification network share the same initial feature extraction network; The first initial classification network is used to classify the game props in the first sample image based on the features extracted from the first sample image by the initial feature extraction network; The second initial classification network is used to classify the game environments where the game props in the first sample image and the second sample image are located based on the features extracted from the first sample image and the second sample image by the initial feature extraction network; The first initial classification network is trained based on a prop classification loss, where the prop classification loss is a classification loss of the first initial classification network for classifying game props in the first sample image based on features of the first sample image; The second initial classification network is trained based on an environment classification loss, where the environment classification loss is a classification loss for the second initial classification network to classify the game environment in which the game props in the first sample image and the second sample image are located based on features of the first sample image and the second sample image; The device further includes: an image set acquisition module, configured to acquire an image set, wherein the image set includes a first image subset and a second image subset, wherein images in the first image subset all include a first label, and images in the second image subset do not include the first label, and the first label is used to represent a category of a game prop; a labeling module, used to label the second labels of the images in the first image subset and the second labels of the images in the second image subset as different second labels, wherein the second labels are used to characterize the category of the game environment where the game props are located; The sample image determination module is used to determine an image in the first image subset as the first sample image, and determine an image in the second image subset as the second sample image.
13. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 10 when executing the program.
Citation Information
Patent Citations
Training device, training method and detection method
CN101655914A
System and method for synthetic image training of a neural network associated with a casino table game monitoring system
US20200402342A1