Text Analytics With Rapidminer Part 5 of 6 - Automatic Document Categorization
Published on 11/12/2010
This is the final second-to-last installment of a six-part series on text mining in RapidMiner. This video describes how to automatically categorize documents. This could be useful for a research project, or say finance.
You could use it to classify documents as “positive” or “negative”, thus doing sentiment analysis. You could do it with financial news text, and classify documents as “stock went up” or “stock went down” after the release, and make (short-term) predictions of future stock movements. You can also see which words are important discriminants. Once you’ve trained a learning algorithm, you can use it on unseen data.
Topics covered:
- Cross-validation
- The nearest neighbor learning algorithm
- The naive bayes learning algorithm
Here is part 6
If you’re not familiar with RapidMiner, see my other videos on my Youtube Channel.
Thanks for watching. Leave a comment for what you’d like to see next!
Also, check out the awesome RapidMiner finance videos on Neural Market Trends.
... views