Data generation method and electronic equipment
By obtaining the associated data of the target object in the sports event video stream and using artificial intelligence processing models to generate commentary data, the problems of subjective bias and poor interactivity in the existing sports event commentary methods are solved, and a more objective and interactive commentary effect is achieved.
Patent Information
- Application Number
- CN202510241381.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-24
AI Technical Summary
The existing sports event commentary methods rely on professional commentators, which have problems of subjective bias and poor interactivity.
By obtaining the associated data of the target object in the target video stream, an artificial intelligence processing model is used to generate interpreted data for the target object, including describing and evaluating the performance of the target object in the video stream.
It realizes objective explanation of sports events, reduces subjective bias, and improves the accuracy and interactivity of the content of the explanation.
Smart Images

Figure CN120201255A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a data generation method and an electronic device. Background Art
[0002] Currently, the commentary of sports events mainly relies on the on-site commentary of professional commentators. However, this commentary method has problems such as affecting the objective understanding of the game by the audience due to the subjective bias of the commentator during the commentary and poor interactivity. Summary of the Invention
[0003] The present disclosure provides a data generation method and an electronic device.
[0004] According to a first aspect of the present disclosure, there is provided a data generation method, including:
[0005] Obtaining target association data of a target object in a target video stream;
[0006] Generating commentary data for the target object based on the target association data, where the commentary data at least includes data describing and / or evaluating the performance data of the target object in the target video stream.
[0007] In an embodiment of the present application, the obtaining of the target association data of the target object in the target video stream includes at least one of the following:
[0008] Determining a target object to be recognized based on target reference information, and obtaining target video frame data associated with the target object from the target video stream, where the target object includes at least one of the objects in the target video stream;
[0009] Invoking an artificial intelligence processing model to obtain the target association data of the target object in the target video stream, where the artificial intelligence model can respond to the invocation of a target application to perform an operation of obtaining the association data, and the target application is the application that outputs the target video stream.
[0010] In an embodiment of the present application, where determining an object to be recognized based on target reference information includes at least one of the following:
[0011] Obtaining first attribute information of the target video stream, and determining a target object to be recognized based on the first attribute information, so as to obtain target video frame data associated with the target object from the target video stream;
[0012] Obtaining user information of a target user, and determining a target object to be recognized based on the user information, so as to obtain target video frame data associated with the target object from the target video stream;
[0013] Obtain the first attribute information and audience information of the target video stream, determine the target object to be recognized based on the first attribute information and the audience information, so as to obtain the target video frame data associated with the target object from the target video stream;
[0014] Obtain the news hot spot information for the target video stream, determine the target object to be recognized based on the news hot spot information, so as to obtain the target video frame data associated with the target object from the target video stream.
[0015] In an embodiment of the present application, determining the target object to be recognized includes at least one of the following:
[0016] In the case where the target video stream is a sports event video stream, determine all the participants in the corresponding sports event as the target object;
[0017] In the case where the target video stream is a film / TV variety show video stream, determine at least one person in the film / TV variety show video as the target object;
[0018] In the case where the target video stream is a live video stream, at least determine the goods in the live video stream as the target object;
[0019] Determine at least one object in the target video stream that matches the user portrait information of the target user as the target object;
[0020] And / or,
[0021] Obtaining the target video frame data associated with the target object from the target video stream includes:
[0022] Segment the target video stream into independent video frame images;
[0023] Use an artificial intelligence processing model to identify the target object in the video frame image and analyze the behavior actions of the target object;
[0024] Obtain the video frame image with the target behavior action as the target video frame data.
[0025] In an embodiment of the present application, generating the commentary data for the target object based on the target association data includes:
[0026] Extract the structured data related to the target object based on the target association data;
[0027] Analyze the structured data based on the video information of the target video stream and the attribute configuration information of the target object;
[0028] Invoke the target processing model to perform generation processing on the analysis result of the structured data, and obtain the explanatory data for the target object.
[0029] In one embodiment of the present application, generating the explanatory data for the target object based on the target associated data includes at least one of the following:
[0030] Obtain the user profile data of the target user, and input the user profile data and the analysis result obtained based on the target associated data into the target processing model to generate the explanatory data for the target object;
[0031] Obtain the target knowledge base data, and control the target processing model to generate the explanatory data for the target object based on the target knowledge base data and the target associated data, where the target knowledge base data is located at the electronic device end or in the cloud;
[0032] Obtain the feedback data from the target user, and generate the explanatory data for the target object based on the feedback data and the target associated data, where the feedback data includes the feedback data of the target user for the target video stream and / or the feedback data for the explanatory data.
[0033] In one embodiment of the present application, it further includes at least one of the following:
[0034] Obtain the question data from the target user, and invoke the target processing model to generate the explanatory data for the question data and its answer result;
[0035] Optimize the model parameters and / or the knowledge base data of the target processing model based on the feedback data of the target user;
[0036] Generate the target recommendation data for the target user based on the interaction data with the target user.
[0037] In one embodiment of the present application, it further includes at least one of the following:
[0038] Based on the output device configuration information of the electronic device, display and output the explanatory data to the target display area of the electronic device and / or play it to the target space environment;
[0039] Display and output the target associated data and / or the analysis result corresponding to the explanatory data to the target display area of the electronic device;
[0040] Synchronously output the explanatory data to at least one terminal device connected to the electronic device.
[0041] According to the second aspect of the present disclosure, there is provided a data generation method, including:
[0042] Obtain the event data of event participants in the event video stream;
[0043] Generate commentary data for a target participant among the event participants based on the event data, where the target participant includes at least one of the event participants, and the commentary data at least includes data describing and / or evaluating the event performance of the target participant.
[0044] According to a third aspect of the present disclosure, there is provided an electronic device, including at least one processor and at least one processing model that can run on the processor, and the processing model can be called by a target application to perform at least one of the following:
[0045] Identify a target object and its associated data in a target video stream;
[0046] Generate commentary data for the target object based on the associated data, where the commentary data at least includes data describing and / or evaluating the performance data of the target object in the target video stream.
[0047] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] By referring to the accompanying drawings and reading the following detailed description, the above and other purposes, features, and advantages of the exemplary embodiments of the present disclosure will become easily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, where:
[0049] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.
[0050] Figure 1 Shows a schematic implementation flow diagram of the data generation method provided by the embodiment of the present disclosure;
[0051] Figure 2 Shows a schematic commentary data generation flow diagram provided by the embodiment of the present disclosure;
[0052] Figure 3 Shows another schematic implementation flow diagram of the data generation method provided by the embodiment of the present disclosure;
[0053] Figure 4 Shows a schematic composition structure diagram of an electronic device provided by the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] To make the objectives, features, and advantages of the present disclosure more apparent and understandable, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present disclosure.
[0055] Currently, there are problems in the sports event commentary method that relies on on-site commentary by professional commentators, such as the problem of subjective commentary bias affecting the objective understanding of the game by the audience and the problem of poor interactivity. Therefore, to solve these problems, the present application provides a data generation method and an electronic device. The electronic device provided by the present application can be devices such as mobile phones, computers, and tablet computers.
[0056] The technical solutions in the embodiments of the present disclosure will be described below with reference to the accompanying drawings in the embodiments of the present disclosure.
[0057] Figure 1 FIG. shows a schematic implementation flowchart of the data generation method provided by the embodiments of the present disclosure. As Figure 1 shown, the data generation method is applied to an electronic device, and the method includes:
[0058] S101, obtaining target association data of a target object in a target video stream.
[0059] The target video stream may include a live video stream, a cloud video, and a local video. The target video stream may be a video stream on topics such as sports events, game events, new product releases, live shopping, academic courses, or entertainment evenings.
[0060] The target object may include one or more objects in the video stream. The types and quantities of target objects in video streams of different topics may be different. For example, the target object in a sports event video stream may be a game participant, the target object in an academic course video stream may be the course lecturer, and the target object in an entertainment evening video stream may be the evening host. The target object in an animal recording video stream may be the animals appearing in the video stream.
[0061] The target association data of the target object in the target video stream may include video frame data that can reflect the personal performance of the target object, the relationship between the target object and other objects, and the environmental scene where the target object is located. The target association data may also include personal performance data of the target object captured from the target video stream, relationship data between the target object and other objects, and environmental scene data where the target object is located. Among them, the personal performance data of the target object may include the behavior data and / or achievement data of the target object, etc. For example, if the target video stream is a movie or TV drama video stream and the target object is the male lead actor of the movie or TV drama, the personal performance data of the male lead actor may include the actor's line level data and the matching level data between the actor's physical performance and the line content, etc. If the target video stream is a sports event video stream and the target object is each athlete in the event, the personal performance data of each athlete may include the athlete's achievement data, sports data, and technical and tactical data, etc. The athlete's achievement data may include the athlete's basic event performance data, the athlete's historical record statistics data, the athlete's technical statistics data in the event, and the athlete's opponent and event-related data, etc. For example, the basic event performance data of tennis player A may include the score data of tennis player A in the event, and the historical record statistics data of tennis player A may include data such as the ranking data in the tennis world, the important event data and award-winning data that the player has participated in, the first-serve winning rate in the season, and the second-serve winning rate in the season. The technical statistics data of tennis player A in the event may include data such as the number of aces, the first-serve winning rate, the second-serve winning rate, and double faults. The opponent and event-related data of tennis player A may include the historical record statistics data of tennis player A's opponent, the important event data and award-winning data that the opponent has participated in, and the level of the current event, etc. The basic event performance data of football player B may include data such as the playing time and the number of goals scored by football player B in the event. The historical record statistics data of football player B may include data such as the historical number of goals, the important event data and award-winning data that the player has participated in, and the number of shots in the season. The technical statistics data of football player B in the event may include data such as the number of shots, the number of assists, the number of mistakes, and the successful pass rate. The opponent and event-related data of football player B may include the historical record statistics data of football player B's opponent, the important event data and award-winning data that the opponent has participated in, and the level of the current event, etc. The technical and tactical data of the athlete may include the athlete's offensive data, defensive data, team cooperation offensive / defensive data, offensive / defensive tactical change data, offensive / defensive tactical success rate, etc. For example, the offensive data of basketball player C may include data such as the shooting percentage and the scoring method, and the defensive data of basketball player C may include data such as the number of blocks and the defensive winning rate. The team cooperation offensive / defensive data of basketball player C may include data such as the number of assists and the number of combined attacks.
[0062] The relationship data between the target object and other objects may include the change data of the positional relationship between the target object and other objects and / or the emotional relationship data, etc. For example, if the target object is basketball player C, the relationship data between basketball player C and other objects may include the standing position relationship data between basketball player C and his teammates and the standing position relationship data between basketball player C and his opponents. If the target object is the male lead in a movie or TV drama, the relationship data between the male lead and other objects may include the standing position relationship between the male lead and other cast members and the emotional communication data between the male lead and other cast members.
[0063] The environmental scene data where the target object is located may include the location information where the target object is located, the environmental change information of the scene where the target object is located, and the scene change information of the scene where the target object is located, etc. For example, if the target object is basketball player C, the environmental scene data where the target object is located may include the standing position information of basketball player C and the scene change information of the basketball court. The basketball court scene may include the halftime break scene, the game scene, and the game suspension scene, etc. If the target object is tennis player A, the environmental scene data where the target object is located may include the standing position information of tennis player A, the scene change information of the tennis court, and the weather change information of the tennis court.
[0064] S102, generating explanation data for the target object based on the target association data, where the explanation data at least includes data for describing and / or evaluating the performance data of the target object in the target video stream.
[0065] The target object may include a specific object or all objects in the target video stream. For example, in a basketball game video stream, basketball player C can be specified as the target object, or all objects in the basketball game video stream can also be used as the target object. In the embodiments of the present disclosure, different explanation data can be generated according to the quantity and type of the target object. For example, if the target object includes all objects in the target video stream, explanation data for explaining the target video stream can be generated for the target video stream. The explanation data for explaining the target video stream may include the performance data of each target object. Explanation data for describing and / or evaluating the performance of each target object in the target video stream can also be generated, which is convenient for users to obtain the explanation data of specific target objects according to their needs. For example, if all basketball players in the basketball game video stream are target objects, explanation data for describing and / or evaluating the performance of each basketball player can be generated. Users can choose to receive the explanation data of all or part of the basketball players according to their own personalized needs.
[0066] In the embodiments of the present disclosure, if the target object is an athlete, the commentary data may include evaluations and descriptions of the athlete's performance in the event video. The event performance may include the athlete's emotional control, mental quality, contribution to the game, physical fitness, and execution of technical movements in the event video. If the target object is an actor, the commentary data may include evaluations and descriptions of the actor's performance level in the film and television video. The performance level may include the clarity of the actor's lines, the coordination of expressions and body movements, and the ability of position scheduling, etc.
[0067] In the embodiments of the present disclosure, the commentary data may further include response data to the Q&A data. The electronic device may pre-collect the user's Q&A data, or may also collect the user's Q&A data in real time during the playback of the target video stream. The user's Q&A data may include Q&A data of the user regarding the performance and basic information of the target object, etc. For example, if the target object is an athlete, the Q&A data may include Q&A data regarding the athlete's historical records, event performance, name, and events participated in, etc. If the target object is a host, the Q&A data may include Q&A data regarding the programs hosted by the host, hosting style, award-winning situations, and hosting level, etc.
[0068] In the embodiments of the present disclosure, the commentary data that meets the user's personalized needs may also be generated using the user portrait information. For example, if the user portrait information reflects that the user likes basketball and likes centers, for a basketball game video, commentary data that focuses on describing and / or evaluating the technical performance of centers in the basketball game video may be generated according to the user portrait information. If the user portrait information reflects that the user likes actor M, for the film and television video in which actor M participates, commentary data that describes and / or evaluates the acting performance of actor M in the film and television video may be generated according to the user portrait information.
[0069] In the embodiments of the present disclosure, the commentary data for the target object may also be generated according to the interaction data with the user. The electronic device may collect the interaction data of users regarding the performance of the target video stream and / or the target object, analyze the user's preferences, and generate commentary data that meets the user's preferences. The interaction data may include voice interaction data and / or text interaction data, etc.
[0070] The usage scenarios of the commentary data can include single-person usage scenarios, multi-person usage scenarios, official platform usage scenarios, non-official platform usage scenarios, etc. In the embodiments of the present disclosure, corresponding commentary data can also be generated according to the usage scenarios of the commentary data, or the commentary emotional tendency, commentary words, etc. of the generated commentary data can be adjusted according to the usage scenarios of the commentary data to make the commentary data more suitable for use. For example, for a single-person usage scenario, the commentary emotional tendency and commentary words of the generated commentary data can be adjusted according to the degree of preference of the user using it for the target object, so as to avoid hitting the user's mood and improve the user experience. For an official platform usage scenario, the commentary emotional tendency and commentary words of the generated commentary data can be adjusted to make the commentary data more professional.
[0071] The knowledge base can store commentary templates for various types of videos such as sports event videos, movie and TV drama videos, evening party videos, animal documentary videos, etc. The knowledge base can also store background information, statistical data, etc. of objects in various industries such as athletes, actors, and hosts. In the embodiments of the present disclosure, the background information, statistical data, etc. of the target object and the commentary template corresponding to the target video stream can also be obtained based on the knowledge base, and the commentary data can be generated by combining the background information and statistical data of the target object, the commentary template, and the associated data of the target object in the target video stream. For example, the knowledge base stores the background information and statistical data of each player of basketball team T, and stores multiple basketball game commentary data and commentary templates of basketball team T. The commentary template can include an opening introduction part, a game process part, a critical moment analysis part, and a player performance summary part. The commentary data for the performance of each player can be generated by combining the background information and statistical data of the target object and the commentary template, and the associated data of each player of basketball team T in the target video stream: "Opening introduction part: Welcome everyone to today's XX game showdown, team A will play against team B; Game process part: After the game starts, team A quickly gets into state, and basketball player 1 assists basketball player 2 to complete a dunk at the beginning; Critical moment analysis part: Now the game enters the last ten seconds, and the point difference is only 2 points. Can team B tie it? Let's wait and see; Player performance summary part: Today, basketball player 2 performed outstandingly, getting 20 points and 8 rebounds throughout the game, while basketball player 1 got a high score of 30 points. Thank you for watching, and see you next time!"
[0072] In the present disclosure, the data generation method can be integrated into the video application client in the form of an AI commentary function, so that the video application client can enable the AI commentary function when playing target video streams such as local videos, cloud videos, or live videos. The data generation method provided by the embodiments of the present disclosure generates commentary data. Alternatively, the functional program corresponding to the data generation method can be integrated into the electronic device in the form of a plugin. When it is detected that the electronic device plays a target video stream through a browser or a video application client, the data generation method provided by the embodiments of the present disclosure can be executed to generate commentary data.
[0073] By using the data generation method provided by the embodiments of the present disclosure, commentary data for describing / evaluating the performance of an object is generated through the associated data of the object in the video stream, which can objectively reflect the performance of the object and help the audience accurately understand the performance of the object in the video.
[0074] In a possible implementation manner, obtaining the target associated data of the target object in the target video stream may include at least one of the following implementation manners A1 and A2:
[0075] Implementation manner A1: Determine the target object to be recognized based on the target reference information, and obtain the target video frame data associated with the target object from the target video stream, where the target object includes at least one of the objects in the target video stream.
[0076] In the embodiments of the present disclosure, the target reference information may include information such as the type of the target video stream, the source information of the target video stream, the audience information of the target video stream, and news hotspots. The type of the target video stream may include at least one of film and television dramas, sports events, art performances, animation videos, or current affairs news videos, etc. The source information may include at least one of the title of the video, release information, copyright information, video format, geographical location information of the shooting location, language information, or which platform or production party the video comes from, etc. The audience information includes the age range of the audience, the regional location information of the audience, and the preference information of the audience, etc. The news hotspot information includes the social media news information associated with the target video stream. For example, the athlete news and event news associated with the sports event video, etc.
[0077] In the embodiments of the present disclosure, a target object to be recognized can be determined using target reference information. For example, the target reference information includes: the type of the target video stream is a tennis match video, the source of the target video stream includes the title "2024 Tennis Final", the shooting location is City P in Country F, and the language is English. The audience information of the target video stream includes sports enthusiasts of all ages in China, and the news hotspot includes Chinese tennis player Z breaking into the final. The target reference information reflects that Chinese tennis player Z has a high degree of attention, and the video audience is Chinese sports enthusiasts. Based on this, Chinese tennis player Z in the tennis match video can be determined as the target object. Then, multiple video frames associated with Chinese tennis player Z in the tennis match video can be recognized as target video frame data, or only the video frames associated with Chinese tennis player Z can be obtained as target video frame data.
[0078] Implementation method A2: Call an artificial intelligence processing model to obtain the target association data of the target object in the target video stream. The artificial intelligence model can respond to the call of the target application to perform the operation of obtaining the association data. The target application is the application that outputs the target video stream.
[0079] In the embodiments of the present disclosure, the target application can include a video application client, a browser application, etc. The artificial intelligence model can include an AI large model or a multi-modal large model that can process videos, etc. The artificial intelligence model can be integrated into the operating system of the electronic device, and the operating system drives the artificial intelligence model to obtain the target association data of the target object in the target video stream. Alternatively, the electronic device can call the artificial intelligence model through an API (Application Programming Interface) and an SDK (Software Development Kit) to obtain the target association data of the target object in the target video stream using the artificial intelligence model.
[0080] In the embodiments of the present disclosure, the artificial intelligence model can analyze the action information, position information, and interaction information with other objects of the target object using a deep learning model based on CNN (Spatio-Temporal Convolutional Network) or LSTM (Long Short-Term Memory Network) to identify the action type of the target object (for example, the shooting action or passing action of a football player, the dunk action or passing action of a basketball player, etc.), and evaluate the action type and action effect of the target object to obtain the target association data.
[0081] In the embodiments of the present disclosure, the target video stream can also be subjected to feature extraction and multi-modal processing operations through a multi-modal large model to obtain the target association data. The multi-modal processing operations include: extracting the action features of the target object in the target video stream through a deep learning model, extracting the temporal features of the video frames in the target video stream through a temporal model, and fusing the features of different modalities to obtain the target association data.
[0082] In the embodiments of the present disclosure, by integrating a multimodal large model in an electronic device or in the cloud, and leveraging the multimodal large model to fuse video analysis, natural language processing, and knowledge base query technologies, all-round intelligent video commentary is achieved. By using the multimodal large model, video stream data and language data can be processed simultaneously to extract information such as the actions and techniques and tactics of target objects from the target video stream, and generate coherent voice commentary content, analysis charts, etc. At the same time, by using the multimodal large model, the knowledge base can be queried in real time or the Internet can be searched to obtain detailed background information of the target object, making the commentary content more comprehensive and the information more abundant. The multimodal large model improves the depth and breadth of video commentary and can provide more personalized, more detailed, and less subjectively biased commentary content for users.
[0083] In a possible implementation manner, determining the object to be recognized based on the target reference information may include at least one of the following implementation manners B1 - B4:
[0084] Implementation manner B1: Obtain the first attribute information of the target video stream, and determine the target object to be recognized based on the first attribute information, so as to obtain the target video frame data associated with the target object from the target video stream.
[0085] The first attribute information of the target video stream may include: the video type of the target video stream and / or the video stream source. The video type may include at least one of sports event videos, game event videos, movie videos, variety show videos, gala videos, shopping live videos, etc. The video stream source may include at least one of a browser, webcast, local video player, or cloud push, etc.
[0086] In the embodiments of the present disclosure, the type of the target video stream may be obtained by recognizing the content of the target video stream. The video stream source of the target video stream may be obtained by parsing at least one of the creation time, device model, encoding information, streaming media protocol, and URL information of the target video stream.
[0087] In the present disclosure, the target object to be recognized may be determined by the video type and / or the video stream source of the target video stream. In different video stream scenarios, the type and / or quantity of the determined target object are different. For example, if the target video stream is a sports event video or a game event video, at least one participant in the target video stream may be determined as the target object. If the target video stream is a movie and television video, variety show video, or gala video, the main performers or the performers that the user is interested in may be determined as the target object. If the target video stream is a shopping live video, the item promoted by the anchor may be determined as the target object.
[0088] Depending on the source of the video stream, the type and / or quantity of the determined target objects may vary. For example, a local video may contain family members, friends, and items, etc., and family members and friends can be determined as target objects. Videos from browser sources usually include feature films, variety shows, etc., and the main actors in the videos can be determined as target objects. Videos from live streaming sources usually include shopping live videos, and items can be determined as target objects.
[0089] In the embodiments of the present disclosure, after determining the target object, video frames in the target video stream that include the target object and video frames where other objects having emotional communication interaction and / or action communication with the target object are located can be identified as target video frame data associated with the target object. Video frames related to the emotional, facial expression, and value attribute changes of the target object and video frames related to the material introduction of the target object in the target video stream can also be identified as target video frame data.
[0090] Implementation method B2: Obtain the user information of the target user, and determine the target object to be identified based on the user information, so as to obtain the target video frame data associated with the target object from the target video stream.
[0091] In the embodiments of the present disclosure, the target user may include the current user of the electronic device or the owner of the electronic device. The user information of the target user may include user identification information and / or user portrait information. The user identification information can characterize the identity, age, gender, etc. of the target user. The user portrait information may include user location information, historical behavior information, and interest preference information. The user portrait information can characterize the preferences and habits of the user. For example, the user portrait of user A indicates that user A frequently watches feature films and variety shows of actor Y on the electronic device through the browser. Then, through the user portrait of user A, it can be determined that user A is accustomed to watching videos through the browser, and the preferences of user A include actor Y, feature film videos, and variety show videos.
[0092] In the embodiments of the present disclosure, the electronic device can obtain the user portrait by mining the behavior data of the user on the browser and various client applications, and analyzing the user behavior. The electronic device can also obtain the user identification information of the user from the user registration information. It can also collect and analyze the questionnaires filled in by the user by sending user information questionnaires to the user to obtain the user information.
[0093] In the embodiments of the present disclosure, if the user information of the target user includes user portrait information, user preference or habit information can be obtained based on the user portrait information, the target object that the user is interested in can be determined according to the user preference or habit information, and the video frames associated with the target object in the target video stream are used as target video frame data. For example, the user portrait of the user reflects that the user likes tennis players and actors. If the target video stream is a tennis match video, the tennis players in the tennis match video can be used as the target object. The video frames associated with the tennis players in the tennis match video are obtained as target video frame data.
[0094] If the user information of the target user includes user identification information, athletes, actors, hosts, items, etc. that the user prefers can be predicted based on the user identification information and the large model. The predicted athletes, actors, hosts, or items that the user prefers are determined as the target object, and then the video frames associated with the target object in the target video stream are used as target video frame data. For example, the user identification of the user includes that the user's age belongs to teenagers. Athletes, actors, hosts, or items that teenagers like can be obtained through the large model. If the target video stream is a basketball game video, the basketball players that teenagers like in the basketball game video can be used as the target object. The video frames associated with the target object in the basketball game video are obtained as target video frame data.
[0095] Implementation method B3: Obtain the first attribute information and the audience information of the target video stream, and determine the target object to be recognized based on the first attribute information and the audience information, so as to obtain the target video frame data associated with the target object from the target video stream.
[0096] In the embodiments of the present disclosure, the first attribute information of the target video stream includes the video stream source or video type, etc. The audience information may include the number of the audience or the age range of the audience. The target object that meets the personalized needs of the audience can be determined according to the video stream source or video type, and the number of the audience or the audience range. For example, if the audience is the elderly, middle-aged and elderly actors or middle-aged and elderly hosts, etc. can be determined as the target object. If the audience is teenagers, young actors or young idols, etc. can be determined as the target object. If the audience is an individual, athletes, actors, hosts, or items can be determined as the target object according to the individual's preferences. If the audience is multiple people, athletes, actors, hosts, or items that all the audience can accept can be determined as the target object. Then, the target video frame data associated with the target object is obtained from the target video stream.
[0097] Implementation method B4: Obtain the news hot spot information for the target video stream, and determine the target object to be recognized based on the news hot spot information, so as to obtain the target video frame data associated with the target object from the target video stream.
[0098] In the embodiments of the present disclosure, the news hot - spot information may include publicity data in external news or news media. For example, if the news hot - spot information includes that the target video stream is about the retirement or comeback game of sports athlete A, then sports athlete A can be used as the target object. If the news hot - spot information includes the promotion information of item A, then item A can be used as the target object. If the news hot - spot information includes the movie release publicity information of actor B, then actor B can be used as the target object. Then, target video frame data associated with the target object is obtained from the target video stream.
[0099] In a possible implementation manner, the determination of the target object to be recognized may include at least one of the following implementation manners C1 - C4:
[0100] Implementation manner C1: When the target video stream is a sports event - type video stream, all participants in the corresponding event are determined as the target object.
[0101] For example, if the target video stream is a basketball game video stream, all basketball players and referees in the basketball game video stream can be determined as the target object. If the target video stream is a tennis game video stream, the tennis players and referees on both sides in the tennis game video stream can be determined as the target object.
[0102] Implementation manner C2: When the target video stream is a film / variety - show - type video stream, at least one person in the film / variety - show video is determined as the target object.
[0103] For example, if the target video stream is a movie video, the leading actors of the movie can be determined as the target object. If the target video stream is a travel variety - show video, the performing artists participating in the travel in the travel variety - show video can be determined as the target object.
[0104] Implementation manner C3: When the target video stream is a live - broadcast - type video stream, at least the goods in the live - broadcast video stream are determined as the target object.
[0105] For example, if the target video stream is a shopping live - broadcast video stream, the goods in the shopping live - broadcast video stream can be determined as the target object. If the target video stream is an evening - party live - broadcast video stream, the host and the performers in the live - broadcast video stream can be determined as the target object.
[0106] Implementation manner C4: At least one object in the target video stream that matches the user portrait information of the target user is determined as the target object.
[0107] For example, the user profile of a user reflects that the user likes tennis players, football players, and actors. If the target video stream is a football game video, all football players in the football game video can be used as the target objects. If the target video stream is a movie video, the actors in the movie video that match the user profile can be determined as the target objects.
[0108] In a possible implementation manner, obtaining the target video frame data associated with the target object from the target video stream may include steps D1 - D3:
[0109] Step D1: Split the target video stream into independent video frame images.
[0110] In the embodiments of the present disclosure, the electronic device can call an image processing tool or program to split the target video stream into multiple independent video frame images by using the image processing tool or program. For example, the OpenCV tool can be called to read the target video stream, and each frame image of the target video stream can be identified and saved as an independent video frame image.
[0111] Step D2: Use an artificial intelligence processing model to identify the target object in the video frame image and analyze the behavioral actions of the target object.
[0112] In the embodiments of the present disclosure, the actions of the objects in each video frame image can be identified through an artificial intelligence processing model, and the actions of the target object can be analyzed based on the actions of the objects in each video.
[0113] Step D3: Obtain the video frame images with target behavioral actions as the target video frame data.
[0114] In the embodiments of the present disclosure, the target behavioral actions of different types of target objects are different. For example, the target behavioral actions of a basketball player may include running, dunking, and passing the ball, etc.; the target behavioral actions of a football player may include running, passing the ball, shooting, and goalkeeping, etc.; the target behavioral actions of an e-sports player may include attacking and releasing ultimate moves, etc.; the target behavioral actions of a shopping live broadcast object may include voice output actions such as putting up a link.
[0115] In a possible implementation manner, Figure 2 shows a schematic diagram of a commentary data generation process provided by the embodiments of the present disclosure. As Figure 2 shown, generating the commentary data for the target object based on the target association data may include:
[0116] S201: Extract the structured data related to the target object based on the target association data.
[0117] In the embodiments of the present disclosure, data related to the actions of the target object can be identified from the target associated data. For example, at least one of the data such as the expression type, action type, action completion accuracy, action completion time, and action success rate of the target object. Structured processing is performed on the data such as the action type, action completion accuracy, action completion time, and action success rate to obtain structured data related to the target object.
[0118] For example, if the target object is basketball player C, the actions of basketball player C that can be identified from the target associated data include dunking actions and passing actions. According to the target associated data, the completion time, completion accuracy, and dunk success rate of basketball player C's dunking actions, as well as the completion time, completion accuracy, and dunk success rate of basketball player C's passing actions can be extracted.
[0119] If the target object is actor A, the actions and expressions of actor A that can be identified from the target associated data include crying actions and sad expressions. According to the target associated data, the completion time, completion accuracy of actor A's crying actions, and the matching degree between the crying actions and sad expressions can be extracted.
[0120] S202, analyze the structured data based on the video information of the target video stream and the attribute configuration information of the target object.
[0121] In the embodiments of the present disclosure, the video information may include at least one of information such as the video stream type, video stream process, and video stream duration. The video stream type may include a sports event type, a game event type, a film and television variety video type, or a live video stream type, etc. The process of the video stream may include each chapter or each stage of the video stream. For example, a basketball game video can be divided into the first half video and the second half video, and a variety video can be divided into multiple episodes, etc. The video stream duration includes the standard duration and the current playback duration. The standard duration includes the standard duration of various event videos, the total length of a movie, or the total duration of each episode / each season / each part of a TV drama, etc. The current stage of the video stream can be determined based on the standard duration and the current playback duration. For example, the current stage of a basketball game video is the last three minutes of the first half, the last quarter of the second half, the last half quarter of the second half, or the last three minutes of the event, etc.
[0122] The attribute configuration information of the target object includes the role positioning of the target object in the video stream. For example, the positioning of athletes in a sports event video and information such as the abilities and competitive levels that athletes should possess. The role positioning of actors in a film and television drama video and the acting skills level that actors should possess when performing roles.
[0123] In the embodiments of the present disclosure, structured data can be analyzed based on the video information of the target video stream and the attribute configuration information of the target object, and combined with data such as the video process, the expression, position, and actions of the target object to evaluate the technical level of the target object. For example, for a sports event video, the technical and tactical level demonstrated by the athletes in the sports event video can be evaluated based on data such as the positioning of the athletes, the configured competitive level that the athletes should possess, the progress of the game, the position and actions of the athletes, etc. In the embodiments of the present disclosure, if the target object is an athlete, structured data related to the actions of the target object can be extracted from the target video stream, such as data on action types, the accuracy and success rate of action completion, etc. Through the tactical analysis module integrated with the multi-modal large model, the structured data related to the actions of the target object is analyzed to generate analysis data for evaluating the tactical strategy of the target object. The tactical analysis module can also combine the position, actions, and game process of the target object on the field to evaluate the technical and tactical skills of the team or athlete and generate analysis data for the tactical strategy.
[0124] In the embodiments of the present disclosure, the generated structured data can be stored in the temporary database of the electronic device for quick invocation when generating commentary data and for interactive query of video information and target object information.
[0125] S203, call the target processing model to perform a generation process on the analysis result of the structured data to obtain commentary data for the target object.
[0126] In the embodiments of the present disclosure, the target processing model can be an AI model capable of generating voice commentary. The target processing model can generate text description information corresponding to the analysis result of the structured data, and then convert the text description information into voice commentary. In the embodiments of the present disclosure, the multi-modal large model can integrate a multi-language conversion module to provide multi-language support for the commentary. Viewers can switch the preferred commentary language according to their own needs.
[0127] In a possible implementation manner, generating commentary data for the target object based on the target association data may include at least one of the following implementation manners E1 - E4:
[0128] Implementation manner E1: Obtain the user portrait data of the target user, and input the user portrait data and the analysis result obtained based on the target association data into the target processing model to generate commentary data for the target object.
[0129] In the embodiments of the present disclosure, the user portrait can be obtained by mining and analyzing the behavior data of the user. The user portrait can reflect the user's preferences or habits.
[0130] The target processing model can obtain user preferences based on user profile data, and then generate analysis result description information that conforms to user preferences based on the analysis results obtained from the target association data, and generate commentary information based on the analysis result description information. For example, if the user prefers the tactical analysis of sports events, commentary data for detailed commentary on the tactics in sports event videos can be generated. If the user prefers the plot, commentary data for detailed commentary on the plot in movies and TV shows can be generated.
[0131] Implementation method E2: Obtain target knowledge base data, and control the target processing model to generate commentary data for the target object based on the target knowledge base data and the target association data, where the target knowledge base data is located at the electronic device end or in the cloud.
[0132] In the embodiments of the present disclosure, the target knowledge base pre-integrates content such as athlete profiles, team historical data, rule explanations of various sports events, plot materials of various movies and TV shows, actor materials, and basic commodity materials of various commodities. The target knowledge base data can be integrated into the electronic device, or can be deployed in the cloud. During the process of generating commentary data, the target processing model can query data such as the basic information, historical records, and personal honors of the target object from the target knowledge base for generating commentary data.
[0133] In the embodiments of the present disclosure, during the process of generating commentary data, the target processing model can also query the latest data about the target object on the network, such as instant comments and news reports about the target object, for generating commentary data.
[0134] When generating commentary data, the target processing model can generate richer and more comprehensive commentary data for the target object by integrating the target video stream data, target object data involved in the target knowledge base, and data about the target object on the network, in combination with the target association data.
[0135] Implementation method E3: Obtain feedback data from the target user, and generate commentary data for the target object based on the feedback data and the target association data, where the feedback data includes the feedback data of the target user for the target video stream and / or the feedback data for the commentary data.
[0136] In the embodiments of the present disclosure, the feedback data of the target user may include emotional feedback data and / or behavioral feedback data. The emotional feedback data may include the attitude of the target user towards the performance of the target object in the target video stream, such as positive, negative, and neutral. The behavioral feedback data includes the satisfaction comments of the target user on the performance of the target object in the target video stream or on the commentary data. The commentary content may be adjusted according to the emotional feedback or behavioral feedback of the user, so that the commentary is real-time and attractive. For example, if the user comments that the commentary on the technical and tactical cooperation of the athlete in the commentary content of the sports event video is not clear enough, the target processing model can enhance the commentary content on the technical and tactical cooperation of the athlete after receiving the user's feedback comment. If the user comments that the emotion of the commentary on the athlete at the key event node in the commentary content of the sports event video is not rich enough, the target processing model can enhance the emotion of the commentary on the athlete at the key event node after receiving the user's feedback comment, so as to improve the immersion of the audience.
[0137] In a possible implementation manner, the data generation method provided by the embodiments of the present disclosure may further include at least one of the following implementation manners F1 - F3:
[0138] Implementation manner F1: Obtain question data from the target user, and call the target processing model to generate commentary data for the question data and its response result.
[0139] In the embodiments of the present disclosure, the electronic device may obtain the question data sent by the target user through the interaction interface with the user. The interaction interface may include a text interaction interface and / or a voice interaction interface, etc. The user can access the interaction interface through various devices such as a mobile phone, a tablet, and a computer, and send questions about the target video stream, the performance of the target object, and the commentary content through the interaction interface. For example, the electronic device may also provide a text interaction interface (such as a real-time questionnaire interface). If the target video stream is a tennis match and the target object is a tennis player in the tennis match, the electronic device can collect the opinions or questions of the audience on the tennis match, the tennis player, the performance of the tennis player in the tennis match, the commentary content, and the commentary style through the real-time questionnaire. The electronic device may also provide a voice interaction interface, and the user can feedback questions in real time through the voice interaction interface. After receiving the questions feedback by the user, the electronic device can perform semantic analysis on the questions through the target processing model, understand the user's intention, and generate commentary data for the question data and its response result according to the user's intention. Or, adjust the current commentary content and commentary style according to the user's intention. The electronic device can also query the data that meets the user's intention from the target knowledge base or the network according to the user's intention, and feedback the queried data to the user in text or voice form in real time.
[0140] Implementation method F2: Optimize the model parameters and / or knowledge base data of the target processing model based on the feedback data of the target user.
[0141] In the embodiments of the present disclosure, the target processing model has a feedback mechanism. The target processing model can collect the feedback data of the user and the interaction data of the electronic device through the feedback mechanism, and optimize the accuracy of the data analysis of the target processing model according to the feedback data and the interaction data.
[0142] Implementation method F3: Generate target recommendation data for the target user based on the interaction data with the target user.
[0143] In the embodiments of the present disclosure, the interaction data between the electronic device and the target user may include at least one of question-and-answer interaction data, feedback interaction data, like interaction data, comment interaction data, and barrage interaction data. For each user, the preference information of the user can be analyzed according to the interaction data between the electronic device and the target user, such as the athletes, actors, video types, commentary styles, commentary contents, etc. that the user prefers. The personalized recommended content can be provided to the user according to the preference information of the user. For example, if it is analyzed that the user likes tennis players, information such as tennis events, relevant news and videos of tennis players can be provided to the user.
[0144] In a possible implementation manner, the data generation method provided by the embodiments of the present disclosure may further include at least one of the following implementation methods G1 - implementation method G4:
[0145] Implementation method G1: Display and output the commentary data to the target display area of the electronic device and / or play it to the target space environment based on the output device configuration information of the electronic device.
[0146] In the embodiments of the present disclosure, the output device configuration information of the electronic device may include speaker configuration information, display screen configuration information, etc. The target display area may include a barrage display area, a subtitle display area, a chart display area, etc. The electronic device can output the commentary content through voice, barrage, subtitle, animated picture, or text live broadcast, etc.
[0147] In the embodiments of the present disclosure, the electronic device can also use a sound-focusing screen to achieve directional output of the commentary content to the target space environment. The target space environment includes the environment where the user is located.
[0148] In the embodiments of the present disclosure, the electronic device can only display the key content in the commentary content. For example, in the event commentary content, the historical record data table of the athlete and the event score, etc.
[0149] Implementation method G2: Display and output the analysis result corresponding to the target association data and / or the commentary data to the target display area of the electronic device.
[0150] In the embodiments of the present disclosure, key content in the analysis results corresponding to the target associated data and / or explanatory data may also be extracted. For example, for content related to statistical data, the statistical data is displayed in a target display area. The target display area may be set in a specified area of the electronic device, such as the lower left corner area or the upper right corner area, etc. The content related to statistical data may include historical battle record statistical data of the target object and data such as the performance success rate of the target object.
[0151] Implementation method G3: Synchronously output the explanatory data to at least one terminal device connected to the electronic device.
[0152] In the embodiments of the present disclosure, the terminal devices connected to the electronic device may include terminal devices such as mobile phones, computers, and tablets.
[0153] In the embodiments of the present disclosure, the explanatory content generated by voice output can be played to the audience in real time to ensure that the explanation is synchronized with the video progress. The explanatory content can also be displayed on the screen in text form for the audience to view, facilitating the audience to quickly understand the video content. The structured data generated in the video can also be displayed in a visual manner such as charts and key data indicators to help the audience better understand the video process.
[0154] Figure 3 Another schematic flowchart of the data generation method provided by the embodiments of the present disclosure is shown.
[0155] As Figure 3 shown, the data generation method is applied to an electronic device, and the method includes:
[0156] S301, Obtain the event data of the event participants in the event video stream.
[0157] S302, Generate explanatory data for a target participant among the event participants based on the event data, where the target participant includes at least one of the event participants, and the explanatory data at least includes data describing and / or evaluating the event performance of the target participant.
[0158] An example implementation example 1 is described. Example 1: The participants in the event include the players of the two competing teams in the basketball finals. The event data includes "The event video stream is that the basketball finals are in progress, and star Z of team A made a crucial three-pointer in the last minute." The electronic device can generate commentary data that describes and / or evaluates the performance of the target participant in the event based on the event data: "Star Z made a three-pointer in the last minute! His three-point shooting percentage has reached 50% in this game. Before this shot, he had made 4 key shots outside the three-point line! This brought his personal score to 30 points." As the commentary progresses, the electronic device can also display a shooting heat map of star Z during the game on the screen, and the shooting heat map can show the shooting positions and shooting percentages of star Z in this game. The audience can ask questions through voice or text during the commentary: "What is star Z's three-point shooting percentage for the entire season?" The electronic device can query in the target knowledge base or via the network and output the instant answer content: "In this season, star Z's three-point shooting percentage is 38.5%, and his shooting percentage in critical moments is as high as 42%."
[0159] An example implementation example 2 is described. Example 2 shows the content of the electronic device's tactical analysis of a football game: The participants in the event include all the players of team A and team B in the football game. The event data includes "Team A used a high-pressure pressing tactic to successfully suppress the midfield of team B." The electronic device can generate commentary data that describes and / or evaluates the performance of the target participant in the event based on the event data: "Team A used a high-pressure pressing tactic in this game, especially in the midfield area, exerting great pressure on the offensive organization of team B. The number of steals by team A has reached 15 times, and 10 of them occurred in the midfield area, making it difficult for team B to complete effective passes." A dynamic tactical board can be displayed on the screen, and the dynamic tactical board shows the running trajectories and pressing routes of team A on the field, and uses arrows to mark the positions where the passes were intercepted. The audience can ask questions through the interaction interface: "What is the passing success rate of the midfield players of team A?" The electronic device can query in the target knowledge base or via the network and output the instant answer content: "The passing success rate of the midfield players of team A is 82%, higher than 70% of team B."
[0160] Exemplary embodiment 3 is described. Embodiment 3 shows a case of analyzing the technical movements in a tennis match: The participants in the event include the athletes in the tennis match. The event data includes "In the tennis match, player C obtained a crucial point through a precise serve." The electronic device can generate commentary data for describing and / or evaluating the performance of the target participant in the event based on the event data: "The serving speed of player C reached 200 km / h this time, and the angle was tricky, directly forcing the opponent to return the ball out of bounds. The serving scoring rate of player C in this match has reached 85%." During the commentary, the system synchronously displays a slow-motion replay of the serve and marks the flight trajectory of the ball and the landing point of the serve on the screen. The audience can ask questions through the interaction interface: "How did player C's serving scoring rate perform throughout the Grand Slam season?" The electronic device can query in the target knowledge base or the network and output the instant answer content: "In the Grand Slam matches of this season, the average serving scoring rate of player C was 78%, but in the crucial matches, this value increased to 83%."
[0161] Exemplary embodiment 4 is described. Embodiment 4 shows a case of real-time commentary in an e-sports match: The participants in the event include the e-sports athletes in the e-sports match. The event data includes "In the e-sports match, the character used by player A made a wonderful operation at a critical moment." The electronic device can generate commentary data for describing and / or evaluating the performance of the target participant in the event based on the event data, including: "The character selected by player A used skills to counter-kill two opposing heroes at a critical moment, gaining a huge advantage for the team. Currently, his KDA (kills / deaths / assists) in this match has reached 5 / 1 / 8, and his performance is very excellent." The skill usage frequency and kill records of player A can also be shown on the screen, as well as the win rate curve of the character selected by player A in the entire professional league. The audience can ask questions through the interaction interface: "What is the win rate of the character selected by player A in the current version?" The electronic device can query in the target knowledge base or the network and output the instant answer content: "In the current version, the win rate of this character is 53.2%, and it is particularly favored by players who are good at the roaming style in the professional arena."
[0162] Example embodiment 5 will be described. Embodiment 5 shows a case of biomechanical analysis in track and field competitions: The event participants include athletes in track and field competitions. The event data includes "In the 100-meter sprint of the track and field competition, athlete T broke his season best result." The electronic device can generate commentary data for describing and / or evaluating the event performance of the target participant based on the event data, including "The stride frequency of athlete T reached 4.5 steps per second during the acceleration phase, and the average speed in the last 10 meters before the finish line was 10.2 meters per second, breaking his season best result." A speed curve graph of athlete T can be displayed on the screen, comparing his previous results and analyzing the improvements in his acceleration and sprint phases in this competition. The audience can interactively ask questions such as "How did athlete T perform in the past 5 competitions?" The electronic device can query through the target knowledge base or the network and output an instant response with a chart content "In the past 5 competitions, the average result of this athlete was 9.85 seconds, and today's performance is his fastest."
[0163] Example embodiment 6 will be described. Embodiment 6 shows a case of pitching analysis in baseball games: The event participants include players in baseball games. The event data includes "In the baseball game, the pitcher struck out the opponent with a fast straight ball." The electronic device can generate commentary data for describing and / or evaluating the event performance of the target participant based on the event data, including "The fast straight ball just thrown by pitcher A reached a speed of 150 km / h, and the rotation speed of the ball was 2500 revolutions per minute. The hitting success rate of the opponent for this type of pitch is only 15%." A trajectory graph of the pitch can be shown on the screen, marking the rotation axis and the landing position of the ball. The audience can interactively ask questions such as "What is pitcher A's strikeout rate ranking in the league?" The electronic device can query through the target knowledge base or the network and output an instant response content "Pitcher A's strikeout rate in this season is 30%, ranking third in the league."
[0164] By using the data generation method provided in the embodiments of the present disclosure, the generated commentary content reduces the interference of human subjective opinions, improves the objectivity and accuracy of the commentary content, and provides a more fair interpretation of the event for the audience. Moreover, the coverage range of the commentary content is more comprehensive, capable of meticulously analyzing complex game scenarios, enhancing the comprehensiveness of the event commentary, and enabling the audience to obtain a more comprehensive understanding of the game. Additionally, the target model can automatically generate commentary content in combination with the most up-to-date information, enabling the audience to keep abreast of the game dynamics and the latest movements of the athletes at any time, enhancing the timeliness and freshness of the viewing experience. Also, the audience can interact in real time, raise their questions and get real-time answers, enhancing the audience's sense of participation and improving the personalization and interactivity of the audience's viewing experience. Furthermore, the electronic device can generate commentary content that conforms to the audience's preferences for different audiences, meeting the personalized needs of the audience and enhancing the relevance of the commentary content and the satisfaction of the user experience. Moreover, the real-time generated commentary content can synchronize with the progress of the game, and the emotional tendency of the commentary can be adjusted according to user feedback, making the commentary content enhance the audience's sense of immersion and making the audience's viewing experience more dynamic and entertaining. Also, by visually displaying the commentary content through data visualization, the audience can more intuitively understand the game process, and supporting multi-terminal synchronous output of the commentary content also facilitates the audience to obtain commentary information on different devices, enhancing the flexibility and convenience of viewing. Additionally, automatically generating commentary content reduces the cost of manual commentary and the pressure on human resources, improves efficiency, and can flexibly switch between multiple languages for commentary to meet the global sports viewing needs. Moreover, the commentary content can include rich technical and tactical analyses, which can help the audience deeply understand the strategic intentions in the game and enhance the professionalism and depth of viewing.
[0165] According to an embodiment of the present disclosure, the present disclosure provides an electronic device, including at least one processor and at least one processing model capable of running on the processor, and the processing model can be called by a target application to perform at least one of the following: identifying a target object and its associated data in a target video stream; generating commentary data for the target object based on the associated data, where the commentary data at least includes data describing and / or evaluating the performance data of the target object in the target video stream.
[0166] In the embodiments of the present disclosure, the target application may include a video playback application or an intelligent agent application installed in the electronic device.
[0167] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.
[0168] Figure 4FIG. 0 shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0169] As Figure 4 shown, the device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0170] A plurality of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as, for example, a keyboard, a mouse, etc.; an output unit 407, such as, for example, various types of displays, speakers, etc.; a storage unit 408, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 409, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0171] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the data generation method. For example, in some embodiments, the data generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the data generation method described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the data generation method in any other suitable manner (e.g., by means of firmware). The electronic device 400 may also include at least one processing model capable of running on the processor, and the processing model can be called by a target application to perform at least one of the following: identifying a target object and its associated data in a target video stream; generating commentary data for the target object based on the associated data, where the commentary data at least includes data describing and / or evaluating the performance data of the target object in the target video stream.
[0172] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0173] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0174] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0175] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0176] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0177] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0178] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0179] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" can explicitly or implicitly include at least one such feature. In the description of this disclosure, "a plurality" means two or more, unless otherwise specifically defined.
[0180] As described above, the above are only specific embodiments of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by this disclosure can easily think of changes or substitutions, which should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.
Claims
1. A data generation method, comprising: Obtaining target associated data of a target object in a target video stream; Based on the target-related data, explanation data for the target object is generated, and the explanation data at least includes data describing and / or evaluating performance data of the target object in the target video stream.
2. The method according to claim 1, wherein obtaining target association data of a target object in a target video stream comprises at least one of the following: Determine a target object to be identified based on the target reference information, and obtain target video frame data associated with the target object from the target video stream, wherein the target object includes at least one of the objects in the target video stream; An artificial intelligence processing model is called to obtain target associated data of a target object in the target video stream, and the artificial intelligence model can respond to a call from a target application to execute an operation of obtaining the associated data, and the target application is an application that outputs the target video stream.
3. The method according to claim 2, wherein: Determining the object to be identified based on the target reference information includes at least one of the following: Obtaining first attribute information of the target video stream, determining a target object to be identified based on the first attribute information, so as to obtain target video frame data associated with the target object from the target video stream; Obtaining user information of a target user, and determining a target object to be identified based on the user information, so as to obtain target video frame data associated with the target object from the target video stream; Obtaining first attribute information and audience information of the target video stream, and determining a target object to be identified based on the first attribute information and the audience information, so as to obtain target video frame data associated with the target object from the target video stream; The news hot spot information for the target video stream is obtained, and the target object to be identified is determined based on the news hot spot information, so as to obtain the target video frame data associated with the target object from the target video stream.
4. The method according to claim 3, wherein: Determine the target object to be identified, including at least one of the following: In the case where the target video stream is a match video stream, all participants who participate in the corresponding match are determined as the target objects; In the case where the target video stream is a film / variety show video stream, at least one character in the film / variety show video is determined as the target object; When the target video stream is a live video stream, at least a commodity in the live video stream is determined as the target object; Determine at least one object in the target video stream that matches the user portrait information of the target user as the target object; and / or, Acquiring target video frame data associated with the target object from the target video stream, comprising: Segmenting the target video stream into independent video frame images; Using an artificial intelligence processing model to identify a target object in the video frame image and analyze the behavior of the target object; A video frame image having a target behavior action is acquired as the target video frame data.
5. The method according to claim 1, wherein generating explanation data for the target object based on the target-related data comprises: Extracting structured data related to the target object based on the target-related data; Analyzing the structured data based on the video information of the target video stream and the attribute configuration information of the target object; The target processing model is called to generate and process the analysis result of the structured data to obtain the interpretation data for the target object.
6. The method according to claim 1 or 5, wherein generating explanation data for the target object based on the target association data comprises at least one of the following: Obtaining user portrait data of a target user, and inputting the user portrait data and an analysis result obtained based on the target association data into a target processing model to generate explanation data for the target object; Obtain target knowledge base data, and control a target processing model to generate explanation data for the target object based on the target knowledge base data and the target association data, wherein the target knowledge base data is located at an electronic device end or in a cloud; Feedback data from a target user is obtained, and explanation data for the target object is generated based on the feedback data and the target association data, wherein the feedback data includes feedback data of the target user for the target video stream and / or feedback data for the explanation data.
7. The method according to claim 1, further comprising at least one of the following: Obtaining question data from a target user, and calling a target processing model to generate explanation data for the question data and its answer result; Optimizing model parameters and / or knowledge base data of a target processing model based on feedback data from target users; Generate target recommendation data for the target user based on the interaction data with the target user.
8. The method according to claim 1, further comprising at least one of the following: Outputting the explanation data to a target display area of the electronic device and / or playing it to a target space environment based on the output device configuration information of the electronic device; Outputting the analysis result corresponding to the target-related data and / or the explanation data to a target display area of the electronic device; The commentary data is synchronously output to at least one terminal device connected to the electronic device.
9. A data generation method, comprising: Obtaining event data of event participants in the event video stream; Based on the event data, commentary data for a target participant among the event participants is generated, wherein the target participant includes at least one of the event participants, and the commentary data at least includes data describing and / or evaluating the event performance of the target participant.
10. An electronic device comprising at least one processor and at least one processing model capable of running on the processor, wherein the processing model can be called by a target application to perform at least one of the following: Identify target objects and their associated data in a target video stream; Based on the association data, explanation data for the target object is generated, and the explanation data at least includes data describing and / or evaluating performance data of the target object in the target video stream.
Citation Information
Cited By
Direct current transmission project stray current training sample generation method
CN121350608A