Methods and systems for personalized indexing and searching online media by spoken word content are disclosed. Some embodiments may include: receiving, at one or more servers, media files and corresponding transcripts, indexing, via the one or more servers, the transcript text in correlation with the associated media files, hosted locations, and aligned timecodes for textual transcript occurrences, accepting, via search interfaces communicatively coupled with the one or more servers, user text search queries to search the indexed transcript text, matching the user text search queries with specific media files and timestamps where matching spoken words and phrases are located, based on the indexed transcript text and returning search results to users, the search results including links to media files where matches occur and direct playback links, the direct playback links embedded with timestamps to commence playbacking from times where search term instances being spoken in the media files.