Content identification method, device, storage medium, and electronic device

By extracting edge features from the target video to generate target labels, the problem of low accuracy in game content recognition is solved, and accurate recognition of game information is achieved.

CN113569616BActive Publication Date: 2025-09-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110215192.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-24
Publication Date
2025-09-12
Estimated Expiration
2041-02-24

AI Technical Summary

Technical Problem

The accuracy of game content recognition in the existing technology is low, especially because the content of games of the same type is similar, which makes it difficult to accurately identify game information.

Method used

By acquiring N frames of images from the target video, the edge features of each frame are extracted. The reference information where the distance between the edge area and the center point is greater than or equal to a preset threshold is used to indicate the information associated with the controlled virtual object in the virtual scene, and a target label is generated to identify the game content.

Benefits of technology

The accuracy of game content recognition is improved, and accurate recognition of game information corresponding to the target video screen is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113569616B_ABST
    Figure CN113569616B_ABST
Patent Text Reader

Abstract

The present invention discloses a content recognition method, device, storage medium and electronic device in the field of artificial intelligence, and also relates to image processing, image recognition and other technologies in the field of computer vision. The method comprises: obtaining N frames of images from a target video; extracting edge features of each frame of the N frames, wherein the edge features are used to represent the features of reference information displayed in the edge area of ​​the image frame, the distance between the edge area and the center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by the picture played by the target video; when a target label corresponding to the picture played by the target video is obtained based on the edge features, the target label is displayed. The present invention solves the technical problem of low accuracy in game content recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to a content identification method, device, storage medium, and electronic device. Background Art

[0002] The live streaming industry has experienced rapid growth in recent years, particularly in the live streaming of games. However, due to the diverse nature of live streaming platforms, many games need to be identified and blocked. However, due to the similarity of content within the same genre, it's difficult to accurately identify the corresponding game information based on the image being played by the target video. This results in low accuracy in game content recognition in the related art.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] Embodiments of the present invention provide a content recognition method, device, storage medium, and electronic device to at least solve the technical problem of low accuracy in game content recognition.

[0005] According to one aspect of an embodiment of the present invention, a method for dynamic region adjustment is provided, comprising: acquiring N frames of images from a target video, wherein N is an integer greater than or equal to 1; extracting edge features of each of the N frames of images, wherein the edge features are used to represent features of reference information displayed in an edge region of the image frame, and a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a picture played by the target video; and displaying the target tag when a target tag corresponding to the picture played by the target video is acquired according to the edge features, wherein the target tag is used to indicate that the application to which the target video belongs is a target application.

[0006] According to another aspect of an embodiment of the present invention, a region dynamic adjustment device is also provided, including: a first acquisition unit, used to acquire N frames of images from a target video, wherein N is an integer greater than or equal to 1; a first extraction unit, used to extract edge features of each frame of the N frames of images, wherein the edge features are used to represent features of reference information displayed in an edge area of ​​the image frame, the distance between the edge area and the center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by the picture played by the target video; a first display unit, used to display the target tag when a target tag corresponding to the picture played by the target video is acquired according to the edge features, wherein the target tag is used to indicate that the application to which the target video belongs is a target application.

[0007] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned method for dynamic region adjustment when running.

[0008] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-mentioned method for dynamic region adjustment through the computer program.

[0009] In an embodiment of the present invention, N frames of images are obtained from a target video, where N is an integer greater than or equal to 1; edge features of each of the N frames of images are extracted, where the edge features are used to represent features of reference information displayed in an edge region of the image frame, a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by the screen played by the target video; when a target tag corresponding to the screen played by the target video is obtained based on the edge features, the target tag is displayed, where the target tag is used to indicate that the application to which the target video belongs is a target application, and the target tag corresponding to the screen played by the target video played in the target video is identified using information associated with the controlled virtual object distributed in the edge region of a frame of image, thereby achieving the purpose of accurately identifying game information corresponding to the screen played by the target video, thereby achieving the technical effect of improving the accuracy of game content recognition, and thus solving the technical problem of low accuracy in game content recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0011] Figure 1 is a schematic diagram of an application environment of an optional content identification method according to an embodiment of the present invention;

[0012] Figure 2 is a schematic diagram of a flow chart of an optional content identification method according to an embodiment of the present invention;

[0013] Figure 3 is a schematic diagram of an optional content identification method according to an embodiment of the present invention;

[0014] Figure 4is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0015] Figure 5 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0016] Figure 6 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0017] Figure 7 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0018] Figure 8 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0019] Figure 9 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0020] Figure 10 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0021] Figure 11 is a schematic diagram of another optional content identification method according to an embodiment of the present invention;

[0022] Figure 12 is a schematic diagram of an optional game content recognition device according to an embodiment of the present invention;

[0023] Figure 13 FIG. 4 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0027] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0028] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision techniques such as using cameras and computers to replace the human eye in identifying, tracking, and measuring objects. Further image processing is performed to transform the computer-generated images into images more suitable for human observation or transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0029] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0030] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, robots, smart medical care, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0031] The solutions provided in the embodiments of this application involve artificial intelligence computer vision, machine learning and other technologies, which are specifically described through the following embodiments:

[0032] The solution of the present application can be applied to identify various application interfaces, such as game interfaces, monitoring interfaces, remote control operation interfaces, etc.

[0033] According to one aspect of an embodiment of the present invention, a content identification method is provided. Optionally, as an optional implementation, the content identification method can be applied to, but is not limited to, Figure 1 In the environment shown in FIG. , it may include but is not limited to a user device 102, a network 110 and a server 112, wherein the user device 102 may include but is not limited to a display 108, a processor 106 and a memory 104. Optionally, in Figure 1 In the example, the display 108 is playing a target video (the target video is a frame of a game screen as an example).

[0034] The specific process can be as follows:

[0035] Step S102 : The user device 102 obtains N frames of images from a target video played on the display 108 , wherein the target video may be, but is not limited to, images played in a target game;

[0036] Steps S104-S106, the user device 102 sends N frames of image to the server 112 via the network 110;

[0037] Step S108: The server 112 extracts edge features of each of the N frames of images through the processing engine 116, thereby generating a target label (or recognition result) corresponding to the edge feature;

[0038] In steps S110-S112, the server 112 sends the target tag (or recognition result) to the user device 102 (or other user device, such as a device held by an auditor, etc.) through the network 110, and the processor 106 in the user device 102 displays the target tag (or recognition result) on the display 108 and stores the target tag (or recognition result) in the memory 104.

[0039] remove Figure 1 In addition to the examples shown, the above steps can be independently completed by the user device 102. That is, the user device 102 performs steps such as edge feature extraction and target label generation, thereby reducing the processing pressure on the server. For example, the server 112 generates a recognition result, and the user device 102 generates a target label based on the received recognition result. The user device 102 includes but is not limited to a handheld device (such as a mobile phone), a laptop computer, a desktop computer, an in-vehicle device, etc. The present invention does not limit the specific implementation of the user device 102.

[0040] Alternatively, as an optional implementation, Figure 2 As shown, the content identification method includes:

[0041] S202, acquiring N frames of images from the target video, where N is an integer greater than or equal to 1;

[0042] S204, extracting edge features of each of the N frames of image, where the edge features are used to represent features of reference information displayed in an edge region of the image frame, a distance between the edge region and a center point of the image frame being greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a frame played by the target video;

[0043] S206 , when a target tag corresponding to the screen played by the target video is acquired according to the edge feature, the target tag is displayed, wherein the target tag is used to indicate that the application to which the target video belongs is the target application.

[0044] Optionally, in this embodiment, the content recognition method can be applied to, but is not limited to, the recognition scenario of shooting game videos played on a video platform, and is used to identify shooting game information corresponding to the shooting game video, and further the reviewer processes the shooting game video according to the corresponding shooting game information, for example, obtaining N frames of images from a video that plays a screen played by a shooting target video, and determining the game tag corresponding to the screen played by the target video played by the video based on the features of the reference information displayed in the edge area of ​​each frame of the N frames, wherein the game tag is used to represent relevant information of the corresponding game, such as the game name, the manufacturer to which the game belongs, whether the game is allowed to be played, etc.

[0045] Optionally, in this embodiment, the content recognition method can be applied to, but is not limited to, recognition scenarios of multiple types of game videos played on a video platform, to identify the game type corresponding to the game video, and further the staff classifies the game video according to the corresponding game type and plays it on the corresponding sub-platform under the video platform. For example, N frames of images are obtained from a video of a screen played by a target video of an unknown type, and a game tag corresponding to the screen played by the target video of the video is determined based on the features of the reference information displayed in the edge area of ​​each frame of the N frames, wherein the game tag is used to represent relevant information of the corresponding game, such as the game type, game name, etc.

[0046] Optionally, in this embodiment, the content recognition method may be applied to, but is not limited to, video content recognition scenarios displayed in a control interface, where the control interface may be, but is not limited to, a control interface of a smart device, such as a drone, an unmanned vehicle, or VR. This is merely an example and is not intended to be limiting.

[0047] Optionally, in this embodiment, a group of continuous multi-frame images can be obtained from the target video, but is not limited to, or the target video can be extracted at a certain time interval to form an image sequence arranged in chronological order, and the group of continuous multi-frame images or image sequence can be used as N frame images, or the group of continuous multi-frame images or image sequence can be further processed, such as filtering and deleting blank or low-information frames, and then deduplicating the image frames, so that the images to be recognized are as inconsistent and diverse as possible to improve recognition efficiency.

[0048] Optionally, in this embodiment, the reference information is used to indicate information associated with the controlled virtual object in the virtual scene indicated by the screen played by the target video, such as the direction, score, health, map, location, etc. Furthermore, the information associated with the controlled virtual object in the virtual scene indicated by the screen played by the target video may be, but is not limited to, first reference information, and may also include, but is not limited to, second reference information unrelated to the controlled virtual object, such as a game icon, game logo, or game name. The reference information includes the first reference information and the second reference information.

[0049] To further illustrate, considering the high similarity of game content, especially the higher similarity of game content with similar game themes, it is necessary to capture the obvious differences in appearance of different games in order to improve the accuracy of game content recognition. For example, Figure 3 As shown, most of the game elements in the game screen 302 may be only slightly different from those of other games, but the game elements corresponding to the reference information 304 in the shadow can well distinguish the difference between games. The reason is that the reference information 304 mostly corresponds to basic information with game-specific attributes (such as the direction associated with the controlled virtual object, score, health, map, position) arranged at the edge of the game screen 302, rather than the main information (such as the controlled virtual object, virtual props, game background, etc.) displayed in the center of the game screen 302 but lacking game-specific attributes.

[0050] Optionally, in this embodiment, the target tag can be displayed on the associated user device, but is not limited to, so that the reviewer can perform an audit operation on the target video according to the target tag. In addition, the target operation can be performed automatically on the target video according to the target tag, but is not limited to, wherein the target operation can include, but is not limited to, a first operation for classifying the video and a second operation for reviewing the video. The classified video is used to indicate that the target video is classified into the class platform corresponding to the target tag for display and playback according to the target tag. The review video is used to indicate whether the target video meets the playback conditions according to the target tag. If it does not meet the conditions, the playback of the target video is prohibited, and a prompt message is sent to prompt the player of the target video to make corrections. If it meets the conditions, the playback of the target video is allowed, and the next video with game screen is retrieved.

[0051] Optionally, in this embodiment, the edge region may be, but is not limited to, a rectangular, circular, elliptical, or irregular shape. Furthermore, the distance between the edge region and the center point of the image frame may be, but is not limited to, a maximum distance, a minimum distance, an average distance, a selected distance, a variance distance, and the like, which are not limited herein.

[0052] To illustrate further, the optional Figure 4As shown, image 402 includes a center point 404 and an edge area 408 (shaded portion) excluding a circle having the center point 404 as the center and a distance 406 as the radius. It can be seen that the distance 406 is the minimum distance between the edge area 408 and the center point 404.

[0053] In addition, in order to indicate that the acquisition of edge areas is not limited, it is optional to use Figure 4 The scenario shown, continuing with e.g. Figure 5 As shown, in the image 402 , a center point 404 and an edge area 506 are included. It can be seen that the distance 502 is the maximum distance between the edge area 506 and the center point 404 , and the distance 504 is the minimum distance between the edge area 506 and the center point 404 .

[0054] It should be noted that N frames of images are obtained from the target video, where N is an integer greater than or equal to 1; edge features of each frame of the N frames are extracted, where the edge features are used to represent features of reference information displayed in an edge area of ​​the image frame, and the distance between the edge area and the center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by the picture played by the target video; when a target tag corresponding to the picture played by the target video is obtained according to the edge features, the target tag is displayed, where the target tag is used to indicate that the application to which the target video belongs is the target application.

[0055] For further example, it is optional to assume that one of the N frames of images obtained in the target video corresponds to the screen played by the target video as follows: Figure 6 As shown in the game screen 602, the edge area 604 can be but is not limited to the partial area above the game screen 602 near the edge, and then the features of the reference information displayed in the edge area 604 are extracted, and the target tag corresponding to the game shown in the game screen 602 is determined based on the features.

[0056] To further illustrate, the optional hypothetical content identification method execution process is as follows Figure 7 As shown, first obtain the image to be identified 702 (obtain N frames of images from the target video), then extract the edge features 704 of the image to be identified 702 (extract the edge features of each frame of the N frames), and finally determine the target label 706 corresponding to the image to be identified 702 based on the edge features 704 (obtain the target label corresponding to the picture played by the target video based on the edge features).

[0057] Through the embodiments provided by the present application, N frames of images are obtained from a target video, where N is an integer greater than or equal to 1; edge features of each frame of the N frames are extracted, where the edge features are used to represent features of reference information displayed in an edge area of ​​the image frame, the distance between the edge area and the center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a picture played by the target video; when a target tag corresponding to the picture played by the target video is obtained according to the edge features, the target tag is displayed, where the target tag is used to indicate that the application to which the target video belongs is a target application, and the target tag corresponding to the picture played by the target video played in the target video is identified using the information associated with the controlled virtual object distributed in the edge area of ​​a frame of the image, thereby achieving the purpose of accurately identifying the game information corresponding to the picture played by the target video, thereby realizing the technical effect of improving the accuracy of game content recognition.

[0058] As an optional solution, after obtaining N frames of images from the target video, the following steps are included:

[0059] S1, extracting global features of each frame of image, wherein the global features are used to represent the global image information of the picture corresponding to the frame of image;

[0060] S2, fuses global features and edge features to obtain target features;

[0061] S3, displays the target label corresponding to the target feature.

[0062] Optionally, in this embodiment, on the one hand, taking into account that the pictures played by some target videos may be missing or for other reasons, the reference value of the reference information is insufficient to determine the game information of the pictures played by the target video; on the other hand, taking into account the comprehensiveness of the basis for identifying game content, the features of the global image information of each frame image used to represent the picture played by the target video corresponding to the frame image are extracted. Optionally, the whole-family image information may include but is not limited to the main content of the picture played by the target video and the background of the picture played by the target video. For example, the main content is a first-person or third-person player holding a gun or a knife, and the background of the picture played by the target video is more complex, such as a room, the wild, or inside a car, etc.

[0063] Optionally, in this embodiment, the fusion of global features and edge features may be, but is not limited to, directly splicing global features and edge features of different vector lengths; it may also be, but is not limited to, first convolving global features and edge features of different lengths into vectors of the same length and then splicing the vectors of the same length; it may also be, but is not limited to, intercepting global features and edge feature convolutions of the same length according to a preset length limit and then splicing them.

[0064] It should be noted that the global features of each frame of image are extracted, wherein the global features are used to represent the global image information of the picture corresponding to a frame of image; when the edge features are extracted, the global features and edge features are fused to obtain the target features; and the target label corresponding to the target features is displayed.

[0065] To further illustrate, the optional Figure 7 The scenario shown, continuing with e.g. Figure 8 As shown, first obtain the image to be identified 702 (obtain N frames of images from the target video), then extract the edge features 704 and global features 802 of the image to be identified 702 (extract the edge features and global features of each frame of the N frames), and then determine the target label 806 corresponding to the image to be identified 702 based on the target features 804 fused from the edge features 704 and the global features 802 (obtain the target label corresponding to the picture played by the target video based on the target features fused from the edge features and the global features).

[0066] Through the embodiments provided in the present application, the global features of each frame of image are extracted, wherein the global features are used to represent the global image information of the picture corresponding to the image frame; the global features and edge features are fused to obtain the target features; the target labels corresponding to the target features are displayed, thereby achieving the purpose of improving the comprehensiveness of the features based on which the target labels are obtained, and realizing the effect of improving the accuracy of obtaining the target labels.

[0067] As an optional solution, extract the global features of a frame of image, including:

[0068] S1, extracting the underlying features of each frame of image, wherein the underlying features are used to represent the image information corresponding to the image frame;

[0069] S2, performs the first convolution operation on the underlying features to obtain global features.

[0070] Optionally, in this embodiment, a feature extraction structure (e.g., a ResNet, Densent, etc. network structure) may be used, but is not limited to, to perform an extraction operation on each frame of image to extract its underlying features (underlying features). A convolutional structure (e.g., a LeNet, AlexNet, ZF Net, etc. network structure) may be used, but is not limited to, to perform a first convolution operation on the underlying features.

[0071] It should be noted that the underlying features of each frame of image are extracted, wherein the underlying features are used to represent image information corresponding to the image frame; and a first convolution operation is performed on the underlying features to obtain global features.

[0072] To further illustrate, an optional classification method, such as one based on deep learning, uses a pre-constructed convolutional network to process (extract the underlying features of each frame of image and perform a first convolution operation on the underlying features) each frame of image.

[0073] Through the embodiments provided in the present application, the underlying features of each frame of image are extracted, wherein the underlying features are used to represent the image information corresponding to the image frame; a first convolution operation is performed on the underlying features to obtain global features, thereby achieving the purpose of obtaining more accurate global features through a series of processing, and realizing the effect of improving the representation accuracy of global features.

[0074] As an optional solution, edge features of each of the N frames of images are extracted, including:

[0075] S1, when the underlying features are extracted, the underlying features are split to obtain local features, wherein the local features are used to represent the image information corresponding to the edge area;

[0076] S2, performing a second convolution operation on the local features to obtain high-level local features, wherein the high-level local features are used to represent information associated with the controlled virtual object in the image information corresponding to the edge region;

[0077] S3, determining the high-level local features as edge features.

[0078] Optionally, in this embodiment, the splitting of the underlying features may be, but is not limited to, cutting out the underlying feature blobs, for example, splitting the top, bottom, left, and right 1 / 8 features (local features); and then, but not limited to, using a convolutional structure (such as LeNet, AlexNet, ZF Net, etc.) to perform a second convolution operation on the top, bottom, left, and right 1 / 8 features to obtain local high-level features (high-level local features), so that the local high-level features can represent relevant game information distributed at the upper, lower, left, and right boundaries of the game screen.

[0079] It should be noted that when the underlying features are extracted, the underlying features are split to obtain local features, wherein the local features are used to represent the image information corresponding to the edge area; a second convolution operation is performed on the local features to obtain high-level local features, wherein the high-level local features are used to represent the information associated with the controlled virtual object in the image information corresponding to the edge area; and the high-level local features are determined as edge features.

[0080] To further illustrate, an optional classification method, such as one based on deep learning, uses a pre-constructed convolutional network to process (extract the underlying features of each frame of image, split the underlying features, and perform a second convolution operation on the local features) each frame of image.

[0081] Through the embodiments provided by the present application, when the underlying features are extracted, the underlying features are split to obtain local features, wherein the local features are used to represent the image information corresponding to the edge area; a second convolution operation is performed on the local features to obtain high-level local features, wherein the high-level local features are used to represent the information associated with the controlled virtual object in the image information corresponding to the edge area; the high-level local features are determined as edge features, thereby achieving the purpose of obtaining more accurate edge features through a series of processing, and realizing the effect of improving the accuracy of the representation of edge features.

[0082] As an optional solution, after obtaining N frames of images from the target video, the following steps are included:

[0083] S1, when a global feature and an edge feature that meets a first recognition condition are extracted, displaying a target label corresponding to the edge feature, wherein the first recognition condition is that the amount of information of the reference information reaches a first threshold;

[0084] S2: When the global feature and the edge feature that does not meet the first recognition condition are extracted, the target label corresponding to the target feature is displayed.

[0085] Optionally, in this embodiment, taking into account that the reference value of edge features varies, if the game content is identified solely based on edge features, on the one hand, it is not conducive to improving the comprehensiveness of the game content identification, and on the other hand, when the reference value of the edge feature is low, it is also impossible to improve the basic recognition accuracy of the game content. On this basis, a first recognition condition is set to determine the reference value of the edge feature, and when the reference value of the edge feature is high (the global feature is extracted and the edge feature meets the first recognition condition), the target label corresponding to the edge feature is displayed, and when the reference value of the edge feature is low (the global feature is extracted and the edge feature does not meet the first recognition condition), the target label corresponding to the target feature is displayed.

[0086] It should be noted that when global features and edge features that meet the first recognition condition are extracted, the target label corresponding to the edge feature is displayed, wherein the first recognition condition is that the amount of information of the reference information reaches the first threshold; when global features and edge features that do not meet the first recognition condition are extracted, the target label corresponding to the target feature is displayed.

[0087] To further illustrate, the optional Figure 7The scenario shown, continuing with e.g. Figure 8 As shown, the image to be identified 702 is first obtained (N frames of images are obtained from the target video), and then the edge features 704 and global features 802 of the image to be identified 702 are extracted (edge ​​features and global features of each frame of the N frames are extracted). Then, assuming that the edge feature 704 does not meet the recognition conditions, the target label 806 corresponding to the image to be identified 702 is determined based on the target feature 804 fused from the edge feature 704 and the global feature 802 (the target label corresponding to the picture played by the target video is obtained based on the target feature fused from the edge feature and the global feature).

[0088] Through the embodiments provided by the present application, when global features and edge features that meet the first recognition condition are extracted, the target label corresponding to the edge feature is displayed, wherein the first recognition condition is that the information amount of the reference information reaches a first threshold; when global features and edge features that do not meet the first recognition condition are extracted, the target label corresponding to the target feature is displayed, thereby achieving the purpose of fully utilizing global features to complete game content recognition even when the edge features do not meet the first recognition condition, and realizing the effect of ensuring the recognition accuracy of game content.

[0089] As an optional solution, after obtaining N frames of images from the target video, the following steps are included:

[0090] S1, when a global feature is extracted and an edge feature that does not meet a second recognition condition is extracted, displaying a target label corresponding to the global feature, wherein the second recognition condition is that the amount of information of the reference information reaches a second threshold;

[0091] S2: When the global feature and the edge feature that meet the second recognition condition are extracted, the target label corresponding to the target feature is displayed.

[0092] Optionally, in this embodiment, taking into account that the reference value of edge features varies, if the edge features are used blindly to identify game content, on the one hand, it is not conducive to improving the comprehensiveness of game content recognition, and on the other hand, when the reference value of the edge features is low, the basic recognition accuracy of the game content cannot be improved. On this basis, a second recognition condition is set to determine the reference value of the edge feature, and when the reference value of the edge feature is low (the global feature is extracted and the edge feature does not meet the second recognition condition), the target label corresponding to the global feature is displayed, and when the reference value of the edge feature is high (the global feature is extracted and the edge feature meets the second recognition condition), the target label corresponding to the target feature is displayed.

[0093] It should be noted that when global features are extracted and edge features that do not meet the second recognition condition are extracted, the target label corresponding to the global feature is displayed, wherein the second recognition condition is that the amount of information of the reference information reaches the second threshold; when global features and edge features that meet the second recognition condition are extracted, the target label corresponding to the target feature is displayed.

[0094] To further illustrate, the optional Figure 8 The scenario shown, continuing with e.g. Figure 9 As shown, the image to be identified 702 is first obtained (N frames of images are obtained from the target video), and then the edge features 704 and global features 802 of the image to be identified 702 are extracted (the edge features and global features of each frame of the N frames are extracted). Then, assuming that the edge features 704 do not meet the recognition conditions, only the global features 802 are used to determine the target label 902 corresponding to the image to be identified 702 (the target label corresponding to the picture played by the target video is obtained based on the target features fused from the edge features and the global features).

[0095] To further illustrate, the optional Figure 7 The scenario shown, continuing with e.g. Figure 8 As shown, the image to be identified 702 is first obtained (N frames of images are obtained from the target video), and then the edge features 704 and global features 802 of the image to be identified 702 are extracted (the edge features and global features of each frame of the N frames are extracted). It is then assumed that the edge features 704 have met the recognition conditions, but for the sake of comprehensive recognition, the target label 806 corresponding to the image to be identified 702 is still determined based on the target features 804 fused from the edge features 704 and the global features 802 (the target label corresponding to the picture played by the target video is obtained based on the target features fused from the edge features and the global features).

[0096] Through the embodiments provided by the present application, when global features are extracted and edge features that do not meet the second recognition condition are extracted, the target label corresponding to the global feature is displayed, wherein the second recognition condition is that the amount of information of the reference information reaches the second threshold; when global features and edge features that meet the second recognition condition are extracted, the target label corresponding to the target feature is displayed, thereby achieving the purpose of improving the comprehensiveness of features used to identify game content and realizing the effect of improving the comprehensiveness of recognition of game content.

[0097] As an optional solution, edge features of each of the N frames of images are extracted, including:

[0098] Inputting N frames of images sequentially into a first network structure of an image recognition model to obtain edge features output by the first network structure, wherein the image recognition model is a model obtained by training an initial neural network model using multiple sample images extracted from the video. The first network structure is used to split and convolve the underlying features corresponding to each frame of the image, and the underlying features are used to represent image information corresponding to the image frame;

[0099] As an optional solution, extract the global features of each frame of the image, including:

[0100] N frames of images are input into the second network structure of the image recognition model to obtain the global features output by the second network structure. The second network structure is used to convolve the underlying features corresponding to each frame of image.

[0101] It should be noted that N frames of images are sequentially input into the first network structure of the image recognition model to obtain edge features output by the first network structure, wherein the image recognition model is a model obtained by training the initial neural network model using multiple sample images extracted from the video. The first network structure is used to split and convolve the underlying features corresponding to each frame of the image, and the underlying features are used to represent the image information corresponding to the image frame; N frames of images are input into the second network structure of the image recognition model to obtain global features output by the second network structure, and the second network structure is used to convolve the underlying features corresponding to each frame of the image.

[0102] To illustrate further, the optional Figure 10 The specific steps are as follows:

[0103] S1002, inputting the sample image into the image recognition model;

[0104] S1004, extracting underlying features of the sample image in the input layer of the image recognition model;

[0105] S1006-1, splitting local features of underlying features in the first network structure;

[0106] S1006-2, extracting high-level features from the underlying features in the second network structure;

[0107] S1008-1, obtaining high-level local features output by the first network structure;

[0108] S1008-2, obtaining global features output by the second network structure;

[0109] S1010, integrating high-level local features and global features;

[0110] S1012, generating a game tag (ie, a target tag).

[0111] Through the embodiment provided by the present application, N frames of images are sequentially input into the first network structure of the image recognition model to obtain edge features output by the first network structure, wherein the image recognition model is a model obtained by training the initial neural network model using multiple sample images extracted from the video, and the first network structure is used to split and convolve the underlying features corresponding to each frame of the image, and the underlying features are used to represent the image information corresponding to the image frame; N frames of images are input into the second network structure of the image recognition model to obtain global features output by the second network structure, and the second network structure is used to convolve the underlying features corresponding to each frame of the image, thereby achieving the purpose of using an efficient network structure to complete the recognition process of game content and realizing the effect of improving the recognition efficiency of game content.

[0112] As an optional solution, before obtaining N frames of images from the target video, the following steps are included:

[0113] S1, obtain multiple sample images;

[0114] S2, performing a first mark on the information associated with each sample image to obtain a plurality of sample images after the first mark;

[0115] S3: Input the plurality of sample images after the first labeling into the initial image recognition model to train and obtain the image recognition model.

[0116] Optionally, in this embodiment, the first mark is used to mark the game information corresponding to the sample image, for example, marking the first sample image as corresponding to game A, the second sample image as corresponding to game B, and so on.

[0117] It should be noted that a plurality of sample images are obtained; information associated with each sample image is first marked to obtain a plurality of sample images after the first marking; and the plurality of sample images after the first marking are input into the initial image recognition model to train the image recognition model.

[0118] Through the embodiments provided in the present application, multiple sample images are obtained; the information associated with each sample image is first marked to obtain multiple sample images after the first marking; the multiple sample images after the first marking are input into the initial image recognition model to train the image recognition model, thereby achieving the purpose of using the trained complete image recognition model to complete the recognition of game content, and realizing the effect of improving the recognition efficiency of game content.

[0119] As an optional solution, a plurality of sample images after the first labeling are input into an initial image recognition model to train an image recognition model, including:

[0120] S1, repeat the following steps until the image recognition model is obtained:

[0121] S2, determining a current sample image from the labeled multiple sample images and determining a current image recognition model;

[0122] S3, splitting and convolving the current sample image through the first network structure of the current image recognition model to obtain the current edge features corresponding to the current sample image;

[0123] S4, performing convolution on the current sample image through the second network structure of the current image recognition model to obtain a current global feature corresponding to the current sample image;

[0124] S5, when a current target feature corresponding to the current sample image obtained by fusing the current edge feature and the current global feature is obtained, a current output result corresponding to the current target feature is obtained through the output structure of the current image recognition model, wherein the current output result is used to indicate game information matching the current sample image;

[0125] S6, when the current output result does not meet the recognition convergence condition, obtaining the next sample image as the current sample image;

[0126] S7: When the current output result reaches the recognition convergence condition, determine that the current image recognition model is the image recognition model.

[0127] It should be noted that the following steps are repeated until an image recognition model is obtained: a current sample image is determined from a plurality of labeled sample images, and a current image recognition model is determined; the current sample image is split and convolved through the first network structure of the current image recognition model to obtain the current edge features corresponding to the current sample image; the current sample image is convolved through the second network structure of the current image recognition model to obtain the current global features corresponding to the current sample image; when the current target features corresponding to the current sample image obtained by fusing the current edge features and the current global features are obtained, the current output result corresponding to the current target features is obtained through the output structure of the current image recognition model, wherein the current output result is used to indicate the game information matched by the current sample image; when the current output result does not meet the recognition convergence conditions, the next sample image is obtained as the current sample image; when the current output result meets the recognition convergence conditions, the current image recognition model is determined to be the image recognition model.

[0128] For further explanation, we can optionally use the identification of shooting game content as an example, as follows:

[0129] First, for subsequent training, the sample images can be further cropped, but are not limited to being cropped. Specifically, four sample image sets are randomly selected from multiple sample images, and the top, bottom, leftmost, and rightmost 1 / 8 of the image are cut off respectively, denoted as Cut1, Cut2, Cut3, and Cut4. The image set without cropping is denoted as Ori, and the left and right images are scaled to the same size.

[0130] Furthermore, for any input sample image, the underlying features are first extracted, such as using a resnet, dense network, etc. After that, it is divided into two operations:

[0131] 1. Continue to extract high-level features; 2. Perform feature splitting on the underlying features of the image.

[0132] The details are as follows:

[0133] 1. If we continue to perform convolution operations on the underlying features, we can obtain high-level features of the image. This high-level feature can represent the game features of the entire image, including information such as color and texture;

[0134] 2. Split the underlying features. This involves cutting out the top, bottom, left, and right 1 / 8 of the underlying feature blob. These features serve as local features of the underlying features. After further convolution operations, local high-level features are obtained. These features represent the distribution of relevant content at the top, bottom, left, and right boundaries of the target video (including high-level features such as directional information, health points, and maps in shooting games).

[0135] Furthermore, a classifier can be provided, but is not limited to, to determine which shooting game the entire image belongs to. Its input is a concatenation of high-level local features and global features. For the images in Cut 1, 2, 3, and 4, the classifier can use softmax or cross entropy as a loss function, or can also use, but is not limited to, more discriminative classification loss functions such as arcface, cosface, center-loss, etc.

[0136] Through the embodiment provided by the present application, the following steps are repeatedly performed until an image recognition model is obtained: a current sample image is determined from a plurality of labeled sample images, and a current image recognition model is determined; the current sample image is split and convolved through the first network structure of the current image recognition model to obtain the current edge features corresponding to the current sample image; the current sample image is convolved through the second network structure of the current image recognition model to obtain the current global features corresponding to the current sample image; when the current target features corresponding to the current sample image obtained by fusing the current edge features and the current global features are obtained, the current output result corresponding to the current target features is obtained through the output structure of the current image recognition model, wherein the current output result is used to indicate the game information matched by the current sample image; when the current output result does not meet the recognition convergence conditions, the next sample image is obtained as the current sample image; when the current output result meets the recognition convergence conditions, the current image recognition model is determined to be the image recognition model, thereby achieving the purpose of the training scheme of the complete image recognition model and realizing the effect of improving the training completeness of the image recognition model.

[0137] As an optional solution, after splitting and convolving the current sample image through the first network structure of the current image recognition model to obtain the current edge features corresponding to the current sample image, the following steps are included:

[0138] S1, when the amount of reference information corresponding to the current edge feature reaches a third threshold, performing a second marking on the current sample image;

[0139] S2. When the amount of reference information corresponding to the current edge feature does not reach a third threshold, the current sample image is marked for the third time, wherein the training weight of the sample image after the second marking in the process of training the image recognition model is greater than that of the sample image after the third marking.

[0140] Optionally, in this embodiment, for edge features, it is also possible but not limited to judge whether the amount of information of the corresponding reference information reaches a third threshold, and give corresponding labels based on the judgment, for example, if it reaches it, label 1 (second label) is given, and if it does not reach it, label 0 (third label) is given. Therefore, in the process of iterative training, it is possible but not limited to updating learning parameters and training weights according to the above-mentioned label 1 or label 0 to improve the training effect of the image recognition model.

[0141] For further explanation, the recognition scenario of the shooting game content mentioned above can be optionally used as an example, as follows:

[0142] For high-level local features, whether they contain game edge information can be judged by, but not limited to, a preset classifier. For example, for the image in Cut1 above, the top game information is cut off, so its corresponding high-level features do not contain local information of the upper edge of the game. Therefore, the classifier should judge that it does not contain the upper information of the shooting game screen and give a label of 0; for the images in Ori, they contain the upper edge information of the game, and the classifier judges that they contain the upper information and gives a label of 1. Similarly, for the images in Cut2, 3, and 4, the classifier should also give a label of 0. The classifier here can also use softmax or cross entropy as the loss function, and can also use, but not limited to, a more discriminative classification loss function, such as arcface, cosface, center-loss and other loss functions.

[0143] It should be noted that when the amount of information of the reference information corresponding to the current edge feature reaches the third threshold, the current sample image is marked for the second time; when the amount of information of the reference information corresponding to the current edge feature does not reach the third threshold, the current sample image is marked for the third time, wherein the training weight of the sample image after the second marking in the process of training the image recognition model is greater than that of the sample image after the third marking.

[0144] To further illustrate, the optional Figure 10 The scenario shown, continuing with e.g. Figure 11 The specific steps are as follows:

[0145] S1002, inputting the sample image into the image recognition model;

[0146] S1004, extracting underlying features of the sample image in the input layer of the image recognition model;

[0147] S1006-1, splitting local features of underlying features in the first network structure;

[0148] S1006-2, extracting high-level features from the underlying features in the second network structure;

[0149] S1008-1, obtaining high-level local features output by the first network structure;

[0150] S1008-2, obtaining global features output by the second network structure;

[0151] S1010, integrating high-level local features and global features;

[0152] S1012, generating a game tag (i.e., a target tag);

[0153] S1102: Determine whether the high-level local feature contains image peripheral information.

[0154] Through the embodiments provided by the present application, when the amount of reference information corresponding to the current edge feature reaches a third threshold, the current sample image is marked for the second time; when the amount of reference information corresponding to the current edge feature does not reach the third threshold, the current sample image is marked for the third time, wherein the training weight of the sample image after the second marking in the process of training the image recognition model is greater than that of the sample image after the third marking, thereby achieving the purpose of improving the training effect of the image recognition model and realizing the effect of improving the recognition accuracy of the game content.

[0155] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0156] According to another aspect of the embodiment of the present invention, a game content recognition device for implementing the above content recognition method is also provided. Figure 12 As shown, the device includes:

[0157] A first acquiring unit 1202 is configured to acquire N frames of images from a target video, where N is an integer greater than or equal to 1;

[0158] A first extraction unit 1204 is configured to extract edge features of each of the N frames of imagery, wherein the edge features represent features of reference information displayed within an edge region of the image frame, wherein a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information indicates information associated with a controlled virtual object in a virtual scene indicated by a frame played by a target video;

[0159] The first display unit 1206 is configured to display a target tag when a target tag corresponding to the screen played by the target video is acquired according to the edge feature, wherein the target tag is used to indicate that the application to which the target video belongs is the target application.

[0160] Optionally, in this embodiment, the game content recognition device can be applied to, but is not limited to, the recognition scenario of shooting game videos played on the video platform, and is used to identify shooting game information corresponding to the shooting game video, and further the reviewer processes the shooting game video according to the corresponding shooting game information, for example, obtaining N frames of images from a video that plays a screen played by a shooting target video, and determining the game tag corresponding to the screen played by the target video played by the video based on the characteristics of the reference information displayed in the edge area of ​​each frame of the N frames, wherein the game tag is used to represent relevant information of the corresponding game, such as the game name, the manufacturer to which the game belongs, whether the game is allowed to be played, etc.

[0161] Optionally, in this embodiment, the game content recognition device can be applied to, but is not limited to, recognition scenarios of multiple types of game videos played on a video platform, to identify the game type corresponding to the game video, and further classify the game video according to the corresponding game type and play it on the corresponding sub-platform under the video platform by the staff. For example, N frames of images are obtained from a video of a screen played by a target video of an unknown type, and a game tag corresponding to the screen played by the target video of the video is determined based on the features of the reference information displayed in the edge area of ​​each frame of the N frames, wherein the game tag is used to represent relevant information of the corresponding game, such as the game type, game name, etc.

[0162] Optionally, in this embodiment, the content recognition device may be used, but is not limited to, in a video content recognition scenario displayed in a control interface, wherein the control interface may be, but is not limited to, a control interface of a smart device, such as a drone, an unmanned vehicle, or VR. This is merely an example and is not intended to be limiting.

[0163] Optionally, in this embodiment, a group of continuous multi-frame images can be obtained from the target video, but is not limited to, or the target video can be extracted at a certain time interval to form an image sequence arranged in chronological order, and the group of continuous multi-frame images or image sequence can be used as N frame images, or the group of continuous multi-frame images or image sequence can be further processed, such as filtering and deleting blank or low-information frames, and then deduplicating the image frames, so that the images to be recognized are as inconsistent and diverse as possible to improve recognition efficiency.

[0164] Optionally, in this embodiment, the reference information is used to indicate information associated with the controlled virtual object in the virtual scene indicated by the screen played by the target video, such as the direction, score, health, map, location, etc. Furthermore, the information associated with the controlled virtual object in the virtual scene indicated by the screen played by the target video may be, but is not limited to, first reference information, and may also include, but is not limited to, second reference information unrelated to the controlled virtual object, such as a game icon, game logo, or game name. The reference information includes the first reference information and the second reference information.

[0165] Optionally, in this embodiment, the target tag can be displayed on the associated user device, but is not limited to, so that the reviewer can perform an audit operation on the target video according to the target tag. In addition, the target operation can be performed automatically on the target video according to the target tag, but is not limited to, wherein the target operation can include, but is not limited to, a first operation for classifying the video and a second operation for reviewing the video. The classified video is used to indicate that the target video is classified into the class platform corresponding to the target tag for display and playback according to the target tag. The review video is used to indicate whether the target video meets the playback conditions according to the target tag. If it does not meet the conditions, the playback of the target video is prohibited, and a prompt message is sent to prompt the player of the target video to make corrections. If it meets the conditions, the playback of the target video is allowed, and the next video with game screen is retrieved.

[0166] Optionally, in this embodiment, the edge region may be, but is not limited to, a rectangular, circular, elliptical, or irregular shape. Furthermore, the distance between the edge region and the center point of the image frame may be, but is not limited to, a maximum distance, a minimum distance, an average distance, a selected distance, a variance distance, and the like, which are not limited herein.

[0167] It should be noted that N frames of images are obtained from the target video, where N is an integer greater than or equal to 1; edge features of each frame of the N frames are extracted, where the edge features are used to represent features of reference information displayed in an edge area of ​​the image frame, and the distance between the edge area and the center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by the picture played by the target video; when a target tag corresponding to the picture played by the target video is obtained according to the edge features, the target tag is displayed, where the target tag is used to indicate that the application to which the target video belongs is the target application.

[0168] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0169] Through the embodiments provided by the present application, N frames of images are obtained from a target video, where N is an integer greater than or equal to 1; edge features of each frame of the N frames are extracted, where the edge features are used to represent features of reference information displayed in an edge area of ​​the image frame, the distance between the edge area and the center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a picture played by the target video; when a target tag corresponding to the picture played by the target video is obtained according to the edge features, the target tag is displayed, where the target tag is used to indicate that the application to which the target video belongs is a target application, and the target tag corresponding to the picture played by the target video played in the target video is identified using the information associated with the controlled virtual object distributed in the edge area of ​​a frame of the image, thereby achieving the purpose of accurately identifying the game information corresponding to the picture played by the target video, thereby realizing the technical effect of improving the accuracy of game content recognition.

[0170] As an optional solution, it includes:

[0171] A second extraction unit is configured to extract global features of each frame of the image after acquiring N frames of the image from the target video, wherein the global features are used to represent global image information of a screen played by the target video corresponding to the frame of the image;

[0172] A fusion unit is used to fuse global features and edge features to obtain target features after acquiring N frames of images from the target video;

[0173] The second display unit is used to display the target label corresponding to the target feature after acquiring N frames of images from the target video.

[0174] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0175] As an optional solution, the second extraction unit includes:

[0176] An extraction module is used to extract the underlying features of each frame of image, wherein the underlying features are used to represent the image information corresponding to the image frame;

[0177] The first convolution module is used to perform a first convolution operation on the underlying features to obtain global features.

[0178] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0179] As an optional solution, the first extraction unit 1204 includes:

[0180] A splitting module is used to split the underlying features to obtain local features when the underlying features are extracted, wherein the local features are used to represent the image information corresponding to the edge area;

[0181] a second convolution module, configured to perform a second convolution operation on the local features to obtain high-level local features, wherein the high-level local features are used to represent information associated with the controlled virtual object in the image information corresponding to the edge region;

[0182] The first determining module is used to determine the high-level local features as edge features.

[0183] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0184] As an optional solution, it includes:

[0185] a third display unit, configured to, after acquiring N frames of images from the target video, display a target label corresponding to the edge feature when a global feature and an edge feature meeting a first recognition condition are extracted, wherein the first recognition condition is that an amount of information of the reference information reaches a first threshold;

[0186] The fourth display unit is used to display the target label corresponding to the target feature after acquiring N frames of images from the target video and extracting the global feature and the edge feature that does not meet the first recognition condition.

[0187] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0188] As an optional solution, it includes:

[0189] a fifth display unit, configured to, after acquiring N frames of images from the target video, display a target label corresponding to the global feature when the global feature is extracted and the edge feature does not meet the second recognition condition, wherein the second recognition condition is that the amount of information of the reference information reaches a second threshold;

[0190] The sixth display unit is used to display the target label corresponding to the target feature after acquiring N frames of images from the target video and extracting the global feature and the edge feature that meets the second recognition condition.

[0191] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0192] As an optional solution,

[0193] A first extraction unit includes: a first input module for sequentially inputting N frames of images into a first network structure of an image recognition model to obtain edge features output by the first network structure, wherein the image recognition model is a model obtained by training an initial neural network model using multiple sample images extracted from a video, and the first network structure is used to split and convolve underlying features corresponding to each frame of the image, and the underlying features are used to represent image information corresponding to the image frame;

[0194] The second extraction unit includes: a second input module, which is used to input N frames of images into the second network structure of the image recognition model to obtain the global features output by the second network structure, and the second network structure is used to convolve the underlying features corresponding to each frame of image.

[0195] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0196] As an optional solution, it includes:

[0197] A second acquisition unit is used to acquire a plurality of sample images before acquiring N frames of images from the target video;

[0198] a marking unit, configured to perform a first marking on information associated with each sample image before acquiring N frames of images from a target video, to obtain a plurality of sample images after the first marking;

[0199] The training unit is used to input a plurality of sample images after the first marking into the initial image recognition model before obtaining N frames of images from the target video, so as to train the image recognition model.

[0200] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0201] As an optional option, training units include:

[0202] The repeat module is used to repeatedly execute the following steps until an image recognition model is obtained:

[0203] a second determining module, configured to determine a current sample image from the labeled plurality of sample images and determine a current image recognition model;

[0204] A first acquisition module is used to split and convolve the current sample image through the first network structure of the current image recognition model to obtain the current edge features corresponding to the current sample image;

[0205] A second acquisition module is used to convolve the current sample image through the second network structure of the current image recognition model to obtain the current global feature corresponding to the current sample image;

[0206] A third acquisition module is configured to obtain, after obtaining a current target feature corresponding to the current sample image obtained by fusing the current edge feature and the current global feature, a current output result corresponding to the current target feature through the output structure of the current image recognition model, wherein the current output result is used to indicate game information matching the current sample image;

[0207] A fourth acquisition module, configured to acquire a next sample image as a current sample image if the current output result does not meet the recognition convergence condition;

[0208] The third determination module is used to determine that the current image recognition model is an image recognition model when the current output result meets the recognition convergence condition.

[0209] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0210] As an optional solution, it includes:

[0211] a first labeling module, configured to, after splitting and convolving the current sample image using the first network structure of the current image recognition model to obtain a current edge feature corresponding to the current sample image, perform a second labeling on the current sample image when the amount of reference information corresponding to the current edge feature reaches a third threshold;

[0212] The second labeling module is used to split and convolve the current sample image through the first network structure of the current image recognition model to obtain the current edge feature corresponding to the current sample image, and then, when the amount of information of the reference information corresponding to the current edge feature does not reach a third threshold, perform a third labeling on the current sample image, wherein the training weight of the sample image after the second labeling in the process of training the image recognition model is greater than that of the sample image after the third labeling.

[0213] For specific embodiments, reference may be made to the examples shown in the above content identification method, which will not be described in detail in this example.

[0214] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above content recognition method is also provided. Figure 13 As shown, the electronic device includes a memory 1302 and a processor 1304. The memory 1302 stores a computer program, and the processor 1304 is configured to execute the steps in any of the above method embodiments through the computer program.

[0215] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0216] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0217] S1, obtaining N frames of images from the target video, where N is an integer greater than or equal to 1;

[0218] S2, extracting edge features of each of the N frames of image, wherein the edge features are used to represent features of reference information displayed in an edge region of the image frame, a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a frame played by a target video;

[0219] S3. When a target tag corresponding to the screen played by the target video is obtained according to the edge feature, the target tag is displayed, wherein the target tag is used to indicate that the application to which the target video belongs is the target application.

[0220] Alternatively, those skilled in the art will appreciate that Figure 13 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 13 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 13 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 13 Different configurations shown.

[0221] Among them, the memory 1302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the content recognition method and device in the embodiment of the present invention. The processor 1304 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302, that is, realizing the above-mentioned content recognition method. The memory 1302 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1302 may further include a memory remotely located relative to the processor 1304, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1302 can be used specifically but not limited to store information such as edge features, N frames of images, and target labels. As an example, if Figure 13As shown, the memory 1302 may include, but is not limited to, the first acquisition unit 1202, the first extraction unit 1204, and the first display unit 1206 in the game content recognition device. In addition, it may also include, but is not limited to, other module units in the game content recognition device, which will not be repeated in this example.

[0222] Optionally, the transmission device 1306 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1306 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1306 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0223] In addition, the electronic device further includes: a display 1308 for displaying the edge features, N frames of images, target labels and other information; and a connection bus 1310 for connecting various module components in the electronic device.

[0224] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0225] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned content identification method, wherein the computer program is configured to perform the steps of any of the aforementioned method embodiments when executed.

[0226] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0227] S1, obtaining N frames of images from the target video, where N is an integer greater than or equal to 1;

[0228] S2, extracting edge features of each of the N frames of image, wherein the edge features are used to represent features of reference information displayed in an edge region of the image frame, a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a frame played by a target video;

[0229] S3. When a target tag corresponding to the screen played by the target video is obtained according to the edge feature, the target tag is displayed, wherein the target tag is used to indicate that the application to which the target video belongs is the target application.

[0230] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0231] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0232] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0233] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0234] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0235] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0236] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0237] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A content identification method, characterized in that: include: Obtain N frames of image from the target video, where N is an integer greater than or equal to 1; Extracting underlying features of each of the N frames of images, wherein the underlying features are used to represent image information corresponding to the image frame; Obtaining a global feature of each frame of the image through the underlying features, wherein the global feature is used to represent global image information of the picture corresponding to the image frame; Splitting the underlying features to obtain local features, wherein the local features are used to represent image information corresponding to edge areas; Obtaining edge features of each frame of the image based on the local features, wherein the edge features are used to represent features of reference information displayed in an edge region of the image frame, a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by a frame played by the target video; fusing the global features and the edge features to obtain target features; When the edge feature meets the first recognition condition, displaying a game tag corresponding to the edge feature, wherein the game tag is used to represent relevant information of the target game application and to indicate that the application to which the target video belongs is the target game application; If the edge feature does not meet the first recognition condition, displaying a game tag corresponding to the target feature; In the case where the global feature and the edge feature that does not meet the second recognition condition are extracted, displaying the target label corresponding to the global feature, wherein the second recognition condition is that the information amount of the reference information reaches a second threshold; When the global feature and the edge feature that meets the second recognition condition are extracted, the target label corresponding to the target feature is displayed.

2. The method according to claim 1, characterized in that Obtaining the global features of each frame of image through the underlying features includes: A first convolution operation is performed on the underlying features to obtain the global features.

3. The method according to claim 2, characterized in that Obtaining the edge features of each frame of image through the local features includes: performing a second convolution operation on the local features to obtain high-level local features, wherein the high-level local features are used to represent information associated with the controlled virtual object in the image information corresponding to the edge region; The high-level local features are determined as the edge features.

4. The method according to claim 1, wherein After acquiring N frames of images from the target video, the method further includes: In the case where the global feature and the edge feature that meets the first recognition condition are extracted, displaying the target label corresponding to the edge feature, wherein the first recognition condition is that the information amount of the reference information reaches a first threshold; When the global feature and the edge feature that does not meet the first recognition condition are extracted, a target label corresponding to the target feature is displayed.

5. The method according to claim 1, wherein After extracting the underlying features of each of the N frames of images, the method further includes: sequentially inputting the N frames of images into a first network structure of an image recognition model to obtain the edge features output by the first network structure, wherein the image recognition model is a model obtained by training an initial neural network model using multiple sample images extracted from a video, and the first network structure is used to split and convolve the underlying features corresponding to each frame of the image, and the underlying features are used to represent image information corresponding to the image frame; The global features of each frame of image are obtained through the underlying features, including: inputting the N frames of image into the second network structure of the image recognition model to obtain the global features output by the second network structure, and the second network structure is used to convolve the underlying features corresponding to each frame of image.

6. The method according to claim 5, characterized in that Before acquiring N frames of images from the target video, the method includes: acquiring the plurality of sample images; Performing a first mark on the information associated with each of the sample images to obtain the plurality of sample images after the first mark; The plurality of sample images after the first marking are input into an initial image recognition model to train the image recognition model.

7. The method according to claim 6, characterized in that Inputting the plurality of sample images after the first marking into an initial image recognition model to train the image recognition model includes: Repeat the following steps until the image recognition model is obtained: Determining a current sample image from the plurality of marked sample images and determining a current image recognition model; Splitting and convolving the current sample image through the first network structure of the current image recognition model to obtain current edge features corresponding to the current sample image; Performing convolution on the current sample image through the second network structure of the current image recognition model to obtain a current global feature corresponding to the current sample image; Upon obtaining a current target feature corresponding to the current sample image obtained by fusing the current edge feature and the current global feature, obtaining a current output result corresponding to the current target feature through the output structure of the current image recognition model, wherein the current output result is used to indicate game information matching the current sample image; If the current output result does not meet the recognition convergence condition, obtaining the next sample image as the current sample image; When the current output result meets the recognition convergence condition, the current image recognition model is determined to be the image recognition model.

8. The method according to claim 7, characterized in that After splitting and convolving the current sample image by the first network structure of the current image recognition model to obtain a current edge feature corresponding to the current sample image, the method includes: When the amount of reference information corresponding to the current edge feature reaches a third threshold, performing a second marking on the current sample image; When the amount of information of the reference information corresponding to the current edge feature does not reach the third threshold, the current sample image is marked for the third time, wherein the training weight of the sample image after the second marking in the process of training the image recognition model is greater than that of the sample image after the third marking.

9. A game content recognition device, characterized in that: include: A first acquisition unit is configured to acquire N frames of images from a target video, where N is an integer greater than or equal to 1; A first extraction unit is configured to extract underlying features of each image frame in the N image frames, wherein the underlying features are used to represent image information corresponding to the image frame; obtain global features of each image frame through the underlying features, wherein the global features are used to represent global image information of the picture corresponding to the image frame; split the underlying features to obtain local features, wherein the local features are used to represent image information corresponding to an edge region; obtain edge features of each image frame through the local features, wherein the edge features are used to represent features of reference information displayed in an edge region of the image frame, wherein a distance between the edge region and a center point of the image frame is greater than or equal to a preset threshold, and the reference information is used to indicate information associated with a controlled virtual object in a virtual scene indicated by the picture played by the target video; a first display unit, configured to display a target tag corresponding to the screen played by the target video when a target tag corresponding to the screen played by the target video is acquired according to the edge feature, wherein the target tag is used to indicate that the application to which the target video belongs is a target application; a second extraction unit, configured to extract global features of each frame of image after acquiring N frames of image from the target video, wherein the global features are used to represent global image information of a screen played by the target video corresponding to one frame of image; A fusion unit, configured to fuse the global feature and the edge feature to obtain a target feature; a second display unit, configured to display a game tag corresponding to the target feature if the edge feature does not meet the first recognition condition, wherein the game tag is used to represent relevant information of the target game application and to indicate that the application to which the target video belongs is the target game application; The device is further configured to display a game tag corresponding to the edge feature when the edge feature meets the first recognition condition, wherein the game tag is used to represent relevant information of the target game application and to indicate that the application to which the target video belongs is the target game application; The device further comprises: a fifth display unit, configured to, after acquiring N frames of images from the target video, display a target label corresponding to the global feature when the global feature and the edge feature that do not meet the second recognition condition are extracted, wherein the second recognition condition is that the amount of information of the reference information reaches a second threshold; The sixth display unit is used to display the target label corresponding to the target feature after acquiring N frames of images from the target video and extracting the global feature and the edge feature that meets the second recognition condition.

10. The device according to claim 9, characterized in that The second extraction unit includes: An extraction module, configured to extract underlying features of each frame of image, wherein the underlying features are used to represent image information corresponding to the image frame; The first convolution module is used to perform a first convolution operation on the underlying features to obtain the global features.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 8 when executed.

12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 8 through the computer program.

Citation Information

Patent Citations

  • Object property recognizing method, device and system

    CN107784282A

  • Game monitoring method and device

    CN109189648A

  • Live broadcast data classification method and device

    CN110267057A

  • Method and device for image processing, computer readable storage medium, and electronic device

    US20190377944A1