Machine learning classification of entrepreneurs in British historical census data

Montebruno, Piero and Bennett, Robert and Smith, Harry and van Lieshout, Carry (2019): Machine learning classification of entrepreneurs in British historical census data. Published in: Information Processing & Management , Vol. 57, No. 3 (May 2020): p. 102210.

There is a more recent version of this item available.

Preview

PDF
MPRA_paper_100469.pdf
Download (8MB) | Preview

Abstract

This paper presents a binary classification of entrepreneurs in British historical data based on the recent availability of big data from the I-CeM dataset. The main task of the paper is to attribute an employment status to individuals that did not fully report entrepreneur status in earlier censuses (1851-1881). The paper assesses the accuracy of different classifiers and machine learning algorithms, including Deep Learning, for this classification problem. We first adopt a ground-truth dataset from the later censuses to train the computer with a Logistic Regression (which is standard in the literature for this kind of binary classification) to recognize entrepreneurs distinct from non-entrepreneurs (i.e. workers). Our initial accuracy for this base-line method is 0.74. We compare the Logistic Regression with ten optimized machine learning algorithms: Nearest Neighbors, Linear and Radial Support Vector Machine, Gaussian Process, Decision Tree, Random Forest, Neural Network, AdaBoost, Naive Bayes, and Quadratic Discriminant Analysis. The best results are boosting and ensemble methods. AdaBoost achieves an accuracy of 0.95. Deep-Learning, as a standalone category of algorithms, further improves accuracy to 0.96 without using the rich text-data that characterizes the OccString feature, a string of up to 500 characters with the full occupational statement of each individual collected in the earlier censuses. Finally, and now using this OccString feature, we implement both shallow (bag-of-words algorithm) learning and Deep Learning (Recurrent Neural Network with a Long Short-Term Memory layer) algorithms. These methods all achieve accuracies above 0.99 with Deep Learning Recurrent Neural Network as the best model with an accuracy of 0.9978. The results show that standard algorithms for classification can be outperformed by machine learning algorithms. This confirms the value of extending the techniques traditionally used in the literature for this type of classification problem.

Item Type:	MPRA Paper
Original Title:	Machine learning classification of entrepreneurs in British historical census data
Language:	English
Keywords:	machine learning; deep learning; logistic regression; classification; big data; census
Subjects:	M - Business Administration and Business Economics ; Marketing ; Accounting ; Personnel Economics > M1 - Business Administration > M13 - New Firms ; Startups N - Economic History > N8 - Micro-Business History > N83 - Europe: Pre-1913
Item ID:	100469
Depositing User:	Dr Piero Montebruno
Date Deposited:	28 Jun 2020 12:30
Last Modified:	28 Jun 2020 12:30
References:	1. Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., Zheng, X (2015). TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. 2. Abdi, A., Shamsuddin, S. M., Hasan, S., & Piran, J. (2019). Deep learning-based sentiment classification of evaluative text based on Multi-feature fusion. Information Processing and Management, 56(4), 1245–1259. 3. Al-Salemi, B., Ayob, M., Kendall, G., & Noah, S. A. M. (2019). Multi-label Arabic text categorization: A benchmark and baseline comparison of multi-label learning algorithms. Information Processing and Management, 56(1), 212–227. 4. Alvarez-Galvez, J. (2016). Discovering complex interrelationships between socioeconomic status and health in Europe: A case study applying Bayesian Networks. Social Science Research, 56, 133–143. https://doi.org/https://doi.org/10.1016/j.ssresearch.2015.12.011 5. Bennett, R. J., Montebruno, P., Smith, H., & van Lieshout, C. (2018). Reconstructing entrepreneur and business numbers for censuses 1851-81. Working paper 9. Cambridge, UK. 6. Bennett, R. J., Montebruno, P., Smith, H., & van Lieshout, C. (2019). Entrepreneurial discrete choice: Modelling decisions between selfemployment, employer and worker status. Working paper 15. Cambridge, UK. 7. Bennett, R. J., Montebruno, P., Smith, H., & van Lieshout, C. (2019). Reconstructing proprietor numbers for censuses 1851-81: Extension and alternative Working paper 9.2. Cambridge, UK. 8. Bennett, R., Smith, H., van Lieshout, C., Montebruno, P., Newton, G. (2020). British Business Census of Entrepreneurs, 1851-1911. [data collection]. UK Data Service. SN: 8600 9. Bennett, R. J., Smith, H., van Lieshout, C., Montebruno, P., & Newton, G. (2019). The Age of Entrepreneurship: Business Proprietors, Self-employment and Corporations Since 1851. 10. Blanchflower, D. G., & Oswald, A. J. (1998). What Makes an Entrepreneur? Journal of Labor Economics, 16(1), 26–60. 11. Boutell, M. R., Luo, J., Shen, X., & Brown, C. M. (2004). Learning multi-label scene classification. Pattern Recognition, 37(9), 1757–1771. 12. Cameron, A. C., & Trivedi, P. K. (2005). Microeconometrics: methods and applications. Cambridge: Cambridge University Press. 13. Capobianco, S., & Marinai, S. (2019). Deep neural networks for record counting in historical handwritten documents. Pattern Recognition Letters, 119, 103–111. 14. Cheng, W., & Hüllermeier, E. (2009). Combining instance-based learning and logistic regression for multilabel classification. Machine Learning, 76(2–3), 211–225. 15. Chollet, F. (2018). Deep learning with Python. Deep learning with Python. Shelter Island, NY: Manning. 16. Dawe, N. (2018). Python Code for Two-class AdaBoost (3-clause BSD License). 17. Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. 18. Freund, Y., & Schapire, R. R. E. (1996). Experiments with a New Boosting Algorithm. International Conference on Machine Learning, 148–156. 19. Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. Ann. Statist., 29(5), 1189–1232. 20. Fürnkranz, J., Hüllermeier, E., Loza Mencía, E., & Brinker, K. (2008). Multilabel classification via calibrated label ranking. Machine Learning, 73(2), 133–153. 21. Géron, A. (2017). Hands-on machine learning with Scikit-Learn and TensorFlow: concepts, tools, and techniques to build intelligent systems. Hands-on machine learning with Scikit-Learn and TensorFlow: concepts, tools, and techniques to build intelligent systems (First edit). Sebastopol, CA: O’Reilly. 22. Goldberger, A. S. (1991). A course in econometrics. A course in econometrics. Cambridge, Mass.; London: Harvard University Press. 23. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. Deep learning. Cambridge, Mass.; London, England: The MIT Press. 24. Head, T. (2018). Python Code for ROC curve (3-clause BSD License). 25. Higgs, E. (2004). Life, death and statistics: civil registration, censuses and the work of the General Register Office, 1836-1952. Life, death and statistics: civil registration, censuses and the work of the General Register Office, 1836-1952. Hatfield: Local Population Studies. 26. Higgs, E., & Schürer, K. (2014). Integrated Census Microdata (I-CeM), 1851-1911, UK Data Archive data deposit SN-7481. UK Data Service. 27. Hindman, M. (2015). Building Better Models: Prediction, Replication, and Machine Learning in the Social Sciences. The ANNALS of the American Academy of Political and Social Science, 659(1), 48–62. 28. Hu, Y., Li, H., Cao, Y., & Li, T. (2006). Automatic extraction of titles from general documents using machine learning. Information Processing & Management, 42(5), 1276–1293. 29. Hüllermeier, E., Fürnkranz, J., Cheng, W., & Brinker, K. (2008). Label ranking by learning pairwise preferences. Artificial Intelligence, 172(16–17), 1897–1916. 30. James, G., Witten, D., Hastie, T., & Tibshirani, R. (2013). An introduction to statistical learning: with applications in R. An introduction to statistical learning: with applications in R. 31. Kastrati, Z., Imran, A. S., & Yayilgan, S. Y. (2019). The impact of deep learning on document classification using semantically rich representations. Information Processing and Management, 56(5), 1618–1632. 32. Kucukyilmaz, T., Cambazoglu, B. B., Aykanat, C., & Baeza-Yates, R. (2017). A machine learning approach for result caching in web search engines. Information Processing and Management, 53(4), 834–850. 33. van Lieshout, C., Bennett, R. J., & Smith, H. (2019). Extracted data on employers and farmers compared with published tables in the Census General Reports, 1851-1881. Working paper 13. 34. Liu, Y., Jin, L., & Lai, S. (2019). Automatic labeling of large amounts of handwritten characters with gate-guided dynamic deep learning. Pattern Recognition Letters, 119, 94–102. 35. Matloff, N. S. (2011). The art of R programming. The art of R programming. San Francisco, Calif.: No Starch Press. 36. McCulloch, W., & Pitts, W. (1943). A logical calculus of the ideas immanent in nervous activity. The Bulletin of Mathematical Biophysics, 5(4), 115–133. 37. Montebruno, P., Bennett, R., van Lieshout, C., Smith, H., & Satchell, A. (2019). Shifts in agrarian entrepreneurship in mid-Victorian England and Wales. The Agricultural History Review, 67(1), 71–108. 38. Montebruno, P., Bennett, R. J., van Lieshout, C., & Smith, H. (2019). A tale of two tails: Do Power Law and Lognormal models fit firm-size distributions in the mid-Victorian era? Physica A: Statistical Mechanics and Its Applications, 523, 858–875. 39. Montebruno, P., Bennett, R. J., van Lieshout, C., & Smith, H. (2019). Research data supporting “A tale of two tails: Do Power Law and Lognormal models fit firm-size distributions in the mid-Victorian era?” Mendeley Data. 40. Murphy, K. P. (2012). Machine learning a probabilistic perspective. Machine learning a probabilistic perspective. Cambridge, MA: MIT Press. 41. Murthy, D., & Gross, A. J. (2017). Social media processes in disasters: Implications of emergent technology use. Social Science Research, 63, 356–370. 42. Parker, S. C. (2004). The Economics of Self-Employment and Entrepreneurship. The Economics of Self-Employment and Entrepreneurship. Cambridge: Cambridge University Press. 43. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., … Duchesnay, E. (2011). Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2825–2830. 44. Rabe-Hesketh, S., & Skrondal, A. (2012). Multilevel and longitudinal modelling using Stata. Volume 2, Categorical responses, counts, and survival. Multilevel and longitudinal modeling using Stata. Volume 2, Categorical responses, counts, and survival (3rd ed.). College Station, Tex.: Stata Press. 45. Read, J., Pfahringer, B., Holmes, G., & Frank, E. (2011). Classifier chains for multi-label classification. Machine Learning, 85(3), 333–359. 46. Reichenberg, O., & Berglund, T. (2019). “Stepping up or stepping down?“: The earnings differences associated with Swedish temporary workers’ employment sequences. Social Science Research. 47. Rosenblatt, F. (1958). The perceptron: A probabilistic model for information storage and organization in the brain. Psychological Review, 65(6), 386–408. 48. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1985). Learning Internal Representations by Error Propagation. La Jolla, California. 49. Schapire, R., & Singer, Y. (1999). Improved Boosting Algorithms Using Confidence-rated Predictions. Machine Learning, 37(3), 297–336. 50. Schürer, K., Penkova, T., & Shi, Y. (2015). Standardising and Coding Birthplace Strings and Occupational Titles in the British Censuses of 1851 to 1911. Historical Methods: A Journal of Quantitative and Interdisciplinary History, 48(4), 195–213. 51. Schwartz, R. L., Phoenix, T., & Foy, brian d. (2008). Learning Perl (5th ed.). Beijing; Farnham: O’Reilly. 52. Su, Z., & Meng, T. (2016). Selective responsiveness: Online public demands and government responsiveness in authoritarian China. Social Science Research, 59, 52–67. 53. Tang, B., Kay, S., & He, H. (2016). Toward Optimal Feature Selection in Naive Bayes for Text Categorization. IEEE Transactions on Knowledge and Data Engineering, 28(9), 2508–2521. 54. Tang, X., Chen, L., Cui, J., & Wei, B. (2019). Knowledge representation learning with entity descriptions, hierarchical types, and textual relations. Information Processing and Management, 56(3), 809–822. 55. The Editors of the American Heritage Dictionaries. (2011). The American Heritage dictionary of the English language. Boston: Houghton Mifflin Harcourt. 56. Tong, S., & Chang, E. (2001). Support vector machine active learning for image retrieval. In Proceedings of the ninth ACM international conference on multimedia (Vol. 9, pp. 107–118). ACM. 57. Treasury Committee. (1890). Report of the Committee appointed by the Treasury to inquire into certain questions connected with the taking of the Census, presented to both Houses of Parliament by Command of Her Majesty. Minutes of evidence, appendices BPP 1890 LVIII. London. 58. Tsoumakas, G., Katakis, I., & Vlahavas, I. (2011). Random k-Labelsets for Multilabel Classification. IEEE Transactions on Knowledge and Data Engineering, 23(7), 1079–1089. 59. Varoquaux, G., & Müller, A. (2018). Python Code for Classifier Comparison (3-clause BSD License). Modified for documentation by Jaques Grobler. 60. Wolpert, D. H. (1996). The Lack of A Priori Distinctions Between Learning Algorithms. Neural Computation, 8(7), 1341–1390. 61. Wu, Q., Ye, Y., Zhang, H., Ng, M. K., & Ho, S.-S. (2014). ForesTexter: An efficient random forest algorithm for imbalanced text categorization. Knowledge-Based Systems, 67(C), 105–116. 62. Zhang, M.-L., & Zhou, Z.-H. (2007). ML-KNN: A lazy learning approach to multi-label learning. Pattern Recognition, 40(7), 2038–2048.
URI:	https://mpra.ub.uni-muenchen.de/id/eprint/100469

Available Versions of this Item

Machine learning classification of entrepreneurs in British historical census data. (deposited 28 Jun 2020 12:30) [Currently Displayed]
- Machine learning classification of entrepreneurs in British historical census data. (deposited 06 Apr 2021 01:43)

All papers reproduced by permission. Reproduction and distribution subject to the approval of the copyright owners.

View Item