Interactive Record Generation from Multimedia Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating interactive records in real-time interactions, such as video conferences, face inefficiencies due to the inability to accurately determine speech information and operation details, leading to low interactive efficiency and poor user experience.

Innovation Solution

A method and apparatus for generating interactive records by collecting behavior data from multimedia streams, including speech and operation information, to create interactive record data that allows users to intuitively understand core ideas and improve interaction efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If speech information is processed and played back from screen recording video, then speech content can be obtained, but interactive efficiency is low and user experience is poor

Engineering Contradiction:
Improvespeech content determinationVSAvoidinteractive efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs speech recognition and operation recognition in advance during the interaction process, generating interactive record data that includes speech content, operation details, and timestamps. This preliminary processing eliminates the need for manual review or playback during subsequent interactions, allowing users to quickly search and reference past interaction content without re-watching videos or replaying recordings.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If speech information is collected and processed, then speech content can be determined, but system complexity increases

Engineering Contradiction:
Improvespeech content determinationVSAvoidprocessing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system introduces an intermediary processing layer that includes speech recognition modules, operation recognition modules, and data generation modules. These intermediaries translate raw speech and operation data into structured interactive record data with clear semantics, making the information easily accessible and actionable without requiring complex manual analysis or playback systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If manual review of screen recording video is performed, then core ideas can be determined, but time consumption increases

Engineering Contradiction:
Improvecore idea determinationVSAvoidtime for interaction review
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system replaces the mechanical process of manual video playback and review with automated recognition systems. Speech recognition technology converts spoken words into text, operation recognition technology identifies user actions, and these are automatically compiled into structured interactive record data. This substitution eliminates the need for users to manually watch and analyze video recordings, dramatically reducing the time required to determine core ideas and interaction details.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12087285B2Method and apparatus for generating interaction record, and device and medium
Publication Date: 2024.09.10 DOUYIN VISION CO LTD
  • US12087285B2 patent drawing
  • US12087285B2 patent drawing
  • US12087285B2 patent drawing

AI summary

A method and apparatus for generating an interaction record, and a device and a medium are provided. The method includes: firstly, from a multimedia data stream, collecting behavior data, represented by the multimedia data stream, of a user, wherein the behavior data includes voice information and/or operation information; and then, on the basis of the behavior data, generating interaction record data corresponding to the behavior data. According to the technical solution, by means of collecting voice information and/or operation information from a multimedia data stream, and generating interaction record data on the basis of the voice information and the operation information, an interacting user can determine interaction information by using the interaction record data, and the interaction efficiency of the interacting user is improved, thereby also improving the user experience.